Roark Status · Incident history
Past incidents.
Every incident we've posted, newest first. Active updates live on the status page.
September 2026
- ResolvedMajor·Affected Agent service
Simulated callers not responding in some simulation runs
Impact window: 2026-09-29, 00:50–02:25 UTC (~95 minutes) During this window, some simulation runs were affected. In those calls the simulated caller could not generate responses. It gave generic replies such as "Sorry, could you say that another way?" or stopped responding, and the call ended early. What happened Requests from the agent service to an upstream language model provider used by our simulated callers were rejected for the duration of the window, so the simulated caller had nothing to say. Resolution Service was restored and simulated callers have been responding normally since the end of the window. Data impact No data was lost. Results from simulation runs started in this window do not reflect your agent's behaviour and can be safely rerun. Next steps We're adding safeguards so an issue with an upstream model provider does not stop simulations, and changing how these failures are reported so affected runs are marked as a platform error rather than as agent behaviour.
Sep 29, 2026 · lasted 1h 35m
- ResolvedMajor·Affected Agent service
Simulated callers not responding in some simulation runs
Impact window: 2026-09-27, 23:45–00:50 UTC (~65 minutes) During this window, some simulation runs were affected. In those calls the simulated caller could not generate responses. It gave generic replies such as "Sorry, could you say that another way?" or stopped responding, and the call ended early. What happened Requests from the agent service to an upstream language model provider used by our simulated callers were rejected for the duration of the window, so the simulated caller had nothing to say. Resolution Service was restored and simulated callers have been responding normally since the end of the window. Data impact No data was lost. Results from simulation runs started in this window do not reflect your agent's behaviour and can be safely rerun. Next steps We're adding safeguards so an issue with an upstream model provider does not stop simulations, and changing how these failures are reported so affected runs are marked as a platform error rather than as agent behaviour.
Sep 27, 2026 · lasted 1h 5m
July 2026
June 2026
- ResolvedMinor·Affected Dashboard
Investigating slower metric collection
We're fully caught up. Metric collection is processing in real time again and dashboards are showing current data, so the customer-facing impact from this incident is over. As a final follow-up, we'll reprocess the small number of jobs that timed out earlier during the backlog — no data will be lost in the meantime. Thanks for your patience.
Jun 19, 2026 · lasted 3d 15h
- ResolvedMinor·Affected Customer API
Elevated error rates on Customer API
Impact window: 2026-06-05, 15:00–15:04 UTC (~3 minutes) On June 5, 2026, the Customer API experienced a brief disruption of approximately three minutes. During this window, a portion of API requests may have failed or timed out. The service recovered automatically and has been fully operational since 15:04 UTC. What happened A backend service process encountered an unexpected fault in an underlying runtime component and automatically restarted. While the replacement instance was starting up and passing health checks, incoming requests during that short window were not served successfully. Resolution The service self-healed once the new instance became healthy. We confirmed full recovery and normal request success rates immediately afterward. Data impact None. No data was lost or affected. Requests that failed during the window were not partially processed and can be safely retried. Next steps To reduce the likelihood and impact of similar events, we're adding redundancy to this service so a single process restart cannot affect availability, and updating the underlying runtime component involved.
Jun 5, 2026 · lasted 16m