Ashby Unavailable

Incident Report for Ashby

Postmortem

Summary

On Sunday, May 31st, at 10pm PST, Ashby was slow and had a high error rate for approximately 1 hour, from 10pm to 11:10pm (PST).

During this time, the Ashby main application would have been slow, and approximately 50% of user actions would have resulted in Ashby displaying an error. Some candidates may have noticed a slower response to their job applications, but all applications were successful and none were lost.

Why did this happen?

On Sunday, May 31st, at 4pm PST, an account with one of our AI providers became unavailable. As a result, calls to that provider started to error. These calls are configured to automatically retry periodically until they succeed, but the retries caused a steady increase in the number of such calls per minute. In turn, this drained the capacity of a shared subsystem that manages many different processes at Ashby.

In particular, that shared subsystem is used for storing user session data. As the shared subsystem became overwhelmed, user sessions became slower to respond.

This caused requests to our web servers to take significantly longer to process. Which then caused our web servers to begin running out of capacity.

At 6am, capacity reached a critical low. At this point, our system began rejecting user actions. This is done to prevent complete system failure.

How did we resolve the situation?

At 10:04pm PST, our on-call engineers were notified. By 10:16pm PST, our on-call engineers were beginning an investigation.

At 11:03pm PST, we identified the subsystem that had reached capacity.

At 11:07pm PST, we stopped processing certain types of automated actions.

At 11:09pm PST, we identified the issue with the external provider and rectified it immediately.

At 11:10pm PST, error rates were down to 10%.

At 11:20pm PST, error rates were down to normal levels (~0%).

At 11:54pm PST, normal service was resumed, and no data was lost.

What have we put in place to prevent it from happening in the future?

We’ve done an internal postmortem on this incident and have implemented or plan to implement a variety of changes that achieve three things:

  1. Reduced the likelihood of future failure
  2. Faster incident response times
  3. Lower impact on customers in the event of failure

Specifically, we have done the following:

  • We were already in the process of replacing the subsystem at the center of the outage. We have accelerated that work. The replacement is not vulnerable to this particular kind of failure.
  • We’ve added additional monitoring around our AI providers to alert us to failures like this sooner.
  • We are prioritizing a fix for the retry mechanism that led to this incident.
  • We are migrating our user session storage to a new subsystem that is not vulnerable to this particular kind of failure.
Posted Jun 19, 2026 - 18:22 UTC

Resolved

This incident has been resolved.
Posted Jun 01, 2026 - 09:04 UTC

Update

All systems are now operational, we are continuing to monitor
Posted Jun 01, 2026 - 08:29 UTC

Update

AI features are enabled again, and we're processing the backlog of accumulated work.
Posted Jun 01, 2026 - 08:07 UTC

Monitoring

The application has recovered. AI features are still affected, and we are actively working to mitigate the issues.
Posted Jun 01, 2026 - 07:31 UTC

Identified

We are seeing signs of recovery, though there may be residual instability. AI features will be unavailable as we mitigate the remaining issues.
Posted Jun 01, 2026 - 07:15 UTC

Update

We are continuing to investigate this issue.
Posted Jun 01, 2026 - 06:49 UTC

Update

We are currently investigating the issue
Posted Jun 01, 2026 - 06:34 UTC

Update

We are continuing to investigate this issue.
Posted Jun 01, 2026 - 06:33 UTC

Investigating

We are currently investigating this issue.
Posted Jun 01, 2026 - 06:33 UTC
This incident affected: Third-Party Integrations (Job Feed, Google Gmail, Google Calendar, Google Meet, Zoom, ATS Sync, HRIS Integration, Dropbox Sign, Microsoft 365, Help Center AI Chat, Assessments), Authentication Services (Google, Office 365, Magic Link, Single Sign On), Ashby APIs (Ashby API, Reports API, Job Post API, Documentation), Ashby Products (Recruiting, Analytics, Hosted Job Boards, Scheduling, Chrome Extension, Mobile, AI Notetaker), and Infrastructure Providers (SendGrid API v3, SendGrid Parse API, AWS).