Is CircleCI Down?
No. CircleCI is up right now.
Every monitored CircleCI system is operational as of October 11, 2026. PulsAPI did not find an outage in its latest check of the official status feed.
Know the moment CircleCI breaks
CircleCI is healthy right now. We keep checking (every 60 seconds once it starts moving) and email you the moment that changes.
Free, no password, no card. Unsubscribe in one click.
CircleCI Component Status
The main CircleCI systems on this page. All 39 are operational right now.
36 more components tracked
Their status is in the bar above. Free accounts see every component by name.
Also tracking CircleCI Dependencies » Auth0 User Authentication, CircleCI Dependencies » AWS, CircleCI Dependencies » Google Cloud Platform Google Cloud DNS, CircleCI Dependencies » Google Cloud Platform Google Cloud Networking, CircleCI Dependencies » Google Cloud Platform Google Cloud Storage, CircleCI Dependencies » Google Cloud Platform Google Compute Engine, CircleCI Dependencies » mailgun API, CircleCI Dependencies » mailgun Outbound Delivery, CircleCI Dependencies » mailgun SMTP, CircleCI Dependencies » OpenAI, CircleCI Insights, CircleCI Releases, and 24 more.
Component statuses may change independently during partial outages.
Recent CircleCI Outages & Incidents
Incident history from the official CircleCI status page, including resolution updates.
Delay in data appearing in UI and APIResolved Oct 9, 2026 · 14h 35mResolved
Between 11:15 UTC and 23:47 UTC on October 9, customers experienced delays and failures across pipelines, workflows and jobs, along with delays in the CircleCI UI, API and notifications. From 11:15 to about 13:00 UTC, pipelines were created but not processed, and outbound webhooks and Slack and email notifications were not sent. From about 13:00 to 14:40 UTC, workflows and jobs started with delays while we worked through the backlog. From about 15:15 to 20:48 UTC, workflows and jobs started with increasing delays, of up to about an hour, and the CircleCI UI and API fell behind by up to about 2 hours. During this time, some jobs failed to start and some in-progress workflows failed. From 20:48 to about 22:10 UTC, we temporarily paused starting new workflows and jobs so our systems could recover. New pipelines were accepted and queued during this time. From about 22:10 to 23:47 UTC, queued workflows and jobs resumed and were processed. Some of them had waited up to about 2 hours. The issue has been resolved and all affected functionality has returned to normal. Usage charge data on the Project Usage page is not yet up to date and will be updated once our data backfills complete. Customers whose jobs failed during this window may rerun affected jobs. We thank you for your patience while our engineers worked on implementing a fix. --- [2026-10-10T01:23:13.373Z] (monitoring) Workflows and jobs are starting and running normally, and the CircleCI UI and API are up to date. We are continuing to monitor our systems closely to make sure they remain stable. We will provide another update by 02:00 UTC. --- [2026-10-10T00:45:09.426Z] (identified) Workflows and jobs have continued to start normally, within seconds, and the CircleCI UI and API remain up to date. Slack and email notifications delayed earlier in the incident are still being delivered. Usage charge data on the Project Usage page is stale. We are continuing to monitor our systems closely and will provide another update by 01:15 UTC. --- [2026-10-09T23:59:55.337Z] (identified) Since our last update, the backlog of queued workflows and jobs has cleared, as of about 23:47 UTC. Workflows and jobs are now starting within seconds, and the CircleCI UI and API are up to date. We are monitoring our systems closely to make sure they remain stable. We will provide another update by 00:30 UTC. --- [2026-10-09T23:38:12.044Z] (identified) We are still working through a significant backlog of queued workflows and jobs. Wait times are improving but they remain longer than normal and some jobs have waited over an hour. The CircleCI UI and API remain up to date. Please avoid retriggering pipelines or workflows while the backlog clears, as this may create duplicates. We will provide another update by 00:00 UTC. --- [2026-10-09T23:00:06.282Z] (identified) A large number of queued workflows and jobs have started, and the CircleCI UI and API have caught up and are showing updates normally with few delays. We are continuing to work through the remaining queued work. Jobs from the queue may have waited about an hour or more to start, and some Linux jobs are waiting up to about 10 minutes for capacity. Please continue to avoid retriggering while queued work continues to process. We will provide another update by 23:30 UTC. --- [2026-10-09T22:28:06.582Z] (identified) Since our last update, we have begun gradually resuming work, and some queued workflows and jobs are starting again. Most queued work is still waiting, and jobs that do start may have waited over an hour. We are increasing the rate carefully to avoid overloading our systems. Delayed updates continue to appear in the CircleCI UI and API. Please continue to avoid retriggering pipelines or workflows, as this may create duplicates as queued work resumes. We will provide another update by 23:00 UTC. --- [2026-10-09T22:09:04.560Z] (identified) Since our last update, most of the pipelines that were delayed earlier in the incident have now been processed, and their workflows are queued to start. New workflows and jobs remain paused while we prepare to resume work safely. Delayed updates continue to appear in the CircleCI UI and API. Please continue to avoid retriggering pipelines or workflows, as this may create duplicates once work resumes. We will provide another update by 22:30 UTC. --- [2026-10-09T21:22:37.747Z] (identified) Since about 20:48 UTC, new workflows and jobs are not starting. We have temporarily paused starting new work so that our systems can catch up on work already in progress, including delayed updates to the CircleCI UI and API. Older updates are now appearing in the UI and API as this backlog clears. New pipelines are still being accepted and queued Please avoid retriggering pipelines or workflows during this time, as this may create duplicates once work resumes. We will provide another update by 22:00 UTC. --- [2026-10-09T20:06:36.341Z] (identified) Jobs are starting faster and fewer are failing: most jobs are now starting within about 4 minutes, down from about 6 minutes, and the longest waits have dropped from about 30 minutes to about 20 minutes. Pipelines from earlier in the incident are now being worked through. The CircleCI UI is still behind, and some older updates may take longer to appear. Jobs may already have run even if the UI does not show them yet, so please avoid retriggering workflows that don't appear in the UI. We will provide another update by 20:30 UTC. --- [2026-10-09T19:29:54.594Z] (identified) Since our last update, jobs are starting faster: most are now starting within about 6 minutes, down from about 8 minutes. The CircleCI UI is catching up but still about an hour behind. New pipelines are being processed as they arrive. Jobs may already have run even if the UI does not show them yet, so please avoid retriggering workflows that don't appear in the UI. We will provide another update by 20:00 UTC. --- [2026-10-09T18:57:01.746Z] (identified) What's changed since our last update * Jobs are starting faster. Most are now starting within about 8 minutes, down from about 15 minutes, and the longest waits have dropped from over an hour to about 45 minutes. * The CircleCI UI is beginning to catch up, but it is still behind. Workflows and jobs may already have run, or may be running, even if the UI shows them as not started or does not show them at all. * Pipelines are being processed faster, but new pipelines are still delayed. Still ongoing * Some workflows and jobs are still failing. * Please avoid retriggering workflows that don't yet appear in the UI, as this may create duplicates. * Usage charge data on the Project Usage page is stale. Next update We will provide another update by 19:30 UTC. --- [2026-10-09T18:32:52.602Z] (identified) What's changed since our last update * The CircleCI UI is about an hour behind, and up to about 90 minutes for some updates. Workflows and jobs may already have run, or may be running, even if the UI shows them as not started or does not show them at all. * Jobs that are running are starting faster. Most are now starting within about 15 minutes, down from about 28 minutes, though some are still waiting over an hour. * New pipelines are increasingly delayed. * Some workflows are still failing. Workflows that could not be processed in time are being marked as failed. * Please avoid retriggering workflows that don't yet appear in the UI, as this may create duplicates. Still ongoing: Usage charge data on the Project Usage page is stale. Next update We will provide another update by 19:00 UTC. --- [2026-10-09T18:08:03.166Z] (identified) Since 11:15 UTC, customers have experienced disruption to pipelines and notifications. Between 11:15 UTC and about 13:00 UTC, pipelines were created but not processed, and outbound webhooks and transactional notifications (Slack and email) were not sent. Between about 13:00 UTC and 14:40 UTC, workflows and jobs started with delays while we worked through the backlog. Since about 15:15 UTC, workflows and jobs have again been starting with delays. What’s impacted Customers triggering pipelines and running workflows and jobs, customers viewing pipelines, workflows and jobs in the CircleCI UI, and customers viewing usage data on the Project Usage page. What to expect - Since about 17:00 UTC, we are seeing an increase in delays. Most jobs are now taking about 25 minutes to start, and some are waiting up to about an hour. - Since about 17:15 UTC, some in-progress workflows have failed and a small number of jobs are failing to start. - New pipelines are being queued and are taking longer to start. - The CircleCI UI is more than an hour behind. New workflows and status changes may take that long to appear. - Please avoid retriggering workflows that don't yet appear in the UI, as this may create duplicates. - Usage charge data on the Project Usage page is currently stale. Our engineers continue to work to resolve the issue at hand. Next update We will provide another update by 18:30 UTC. --- [2026-10-09T17:34:01.223Z] (identified) Since 11:15 UTC, customers have experienced disruption to pipelines and notifications. Between 11:15 UTC and about 13:00 UTC, pipelines were created but not processed, and outbound webhooks and transactional notifications (Slack and email) were not sent. Between about 13:00 UTC and 14:40 UTC, workflows and jobs started with delays while we worked through the backlog. Since about 15:15 UTC, workflows and jobs have again been starting with delays. What’s impacted Customers triggering pipelines and running workflows and jobs, customers viewing pipelines, workflows and jobs in the CircleCI UI, and customers viewing usage data on the Project Usage page. What can you expect - Since about 17:00 UTC, delays have increased significantly. Most jobs are now taking about 12 minutes to start, and some are waiting up to about 30 minutes. - Since about 17:15 UTC, some jobs are failing to start. - Pipelines are being processed more slowly, so some may take longer to start. - The CircleCI UI is more than an hour behind. New workflows and status changes may take that long to appear. - Please avoid retriggering workflows that don’t yet appear in the UI, as this may create duplicates. - Usage charge data on the Project Usage page is currently stale. Thank you for your patience while our engineers work to resolve this. Next update We will provide another update by 18:00 UTC. --- [2026-10-09T17:06:30.170Z] (monitoring) Since 11:15 UTC, customers have experienced disruption to pipelines and notifications. Between 11:15 UTC and about 13:00 UTC, pipelines were created but not processed, and outbound webhooks and transactional notifications (Slack and email) were not sent. Between about 13:00 UTC and 14:40 UTC, workflows and jobs started with delays while we worked through the backlog. Since about 15:15 UTC, workflows and jobs have again been starting with delays. What’s impacted Customers running workflows and jobs, customers viewing pipelines, workflows and jobs in the CircleCI UI, and customers viewing usage data on the Project Usage page. What can you expect - Delays starting workflows and jobs are holding steady. Most jobs are starting within about 4 minutes, and some are waiting up to about 8 minutes, down from about 25 minutes. - The CircleCI UI remains about 20 minutes behind on average, and some updates are taking longer. New workflows and status changes may take that long to appear. - Please avoid retriggering workflows that don’t yet appear in the UI, as this may create duplicates. - Usage charge data on the Project Usage page is currently stale. Thank you for your patience while our engineers work to reduce these delays. Next update We will provide another update by 17:30 UTC. --- [2026-10-09T16:32:53.245Z] (monitoring) Since 11:15 UTC, customers have experienced disruption to pipelines and notifications. Between 11:15 UTC and about 13:00 UTC, pipelines were created but not processed, and outbound webhooks and transactional notifications (Slack and email) were not sent. Between about 13:00 UTC and 14:40 UTC, workflows and jobs started with delays while we worked through the backlog. Since about 15:15 UTC, workflows and jobs have again been starting with delays. What's impacted Customers running workflows and jobs, customers viewing pipelines, workflows and jobs in the CircleCI UI, and customers viewing usage data on the Project Usage page. What can you expect - Delays starting workflows and jobs have stopped increasing and are beginning to ease. Most jobs are starting within about 4 minutes, and some are waiting up to about 25 minutes. - The CircleCI UI is now about 20 minutes behind, down from about 30 minutes. New workflows and status changes may take that long to appear. - Please avoid retriggering workflows that don't yet appear in the UI, as this may create duplicates. - Usage charge data on the Project Usage page is currently stale. Thank you for your patience while our engineers work to reduce these delays. Next update We will provide another update by 17:00 UTC. --- [2026-10-09T16:02:42.495Z] (monitoring) Since 11:15 UTC, customers have experienced disruption to pipelines and notifications. Between 11:15 UTC and about 13:00 UTC, pipelines were created but not processed, and outbound webhooks and transactional notifications (Slack and email) were not sent. Between about 13:00 UTC and 14:40 UTC, workflows and jobs started with delays while we worked through the backlog. What's impacted Customers running workflows and jobs, customers viewing pipelines, workflows and jobs in the CircleCI UI, and customers viewing usage data on the Project Usage page. What can you expect - Since about 15:15 UTC, workflows and jobs have been starting with increasing delays. Most jobs are currently starting within about 4 minutes, and some are waiting up to about 20 minutes. - The CircleCI UI remains about 30 minutes behind on average, and some updates are taking longer. New workflows and status changes may take that long to appear. - Please avoid retriggering workflows that don't yet appear in the UI, as this may create duplicates. - Usage charge data on the Project Usage page is currently stale. Thank you for your patience while our engineers work to reduce these delays. Next update We will provide another update by 16:30 UTC. --- [2026-10-09T15:32:03.279Z] (monitoring) Since 11:15 UTC, customers have experienced disruption to pipelines and notifications. Between 11:15 UTC and about 13:00 UTC, pipelines were created but not processed, and outbound webhooks and transactional notifications (Slack and email) were not sent. Between about 13:00 UTC and 14:40 UTC, workflows and jobs started with delays while we worked through the backlog. What's impacted Customers viewing pipelines, workflows and jobs in the CircleCI UI, and customers viewing usage data on the Project Usage page. What can you expect - Workflows and jobs are running on time. - The CircleCI UI is about 30 minutes behind. New workflows and status changes may take that long to appear. - Our engineers are actively working to speed up processing so the UI can catch up. - Please avoid retriggering workflows that don't yet appear in the UI, as this may create duplicates. - Usage charge data on the Project Usage page is currently stale. Thank you for your patience while our engineers work to reduce this delay. Next update We will provide another update by 16:00 UTC. --- [2026-10-09T14:54:02.940Z] (monitoring) Since 11:15 UTC, customers have experienced disruption to pipelines and notifications. Between 11:15 UTC and about 13:00 UTC, pipelines were created but not processed, and outbound webhooks and transactional notifications (Slack and email) were not sent. Between about 13:00 UTC and 14:40 UTC, workflows and jobs started with delays while we worked through the backlog. What’s impacted Customers running workflows and jobs, and customers viewing usage data on the Project Usage page. What can you expect - The backlog of pipelines and workflows from the incident has cleared. - Since about 14:45 UTC, most workflows and jobs are starting within about 15 seconds, and start times are continuing to return to normal. - Usage charge data on the Project Usage page is currently stale. Thank you for your patience while our engineers monitor the recovery. Next update We will provide another update by 15:20 UTC. --- [2026-10-09T14:43:52.603Z] (identified) Since 11:15 UTC, customers have experienced disruption to pipelines and notifications. Between 11:15 UTC and about 13:00 UTC, pipelines were created but not processed, and outbound webhooks and transactional notifications (Slack and email) were not sent. Between about 13:00 UTC and 14:40 UTC, workflows and jobs started with delays while we worked through the backlog. What’s impacted Customers running workflows and jobs, and customers viewing usage data on the Project Usage page. What can you expect - The backlog of workflows from pipelines triggered during the incident cleared at about 14:40 UTC. - Most jobs are now starting within about 30 seconds, and some are waiting up to about 3 minutes. - Usage charge data on the Project Usage page is currently stale. Thank you for your patience while our engineers continue to work on the underlying issue. Next update We will provide another update by 15:10 UTC. --- [2026-10-09T14:29:58.944Z] (identified) Since 11:15 UTC, customers have experienced disruption to pipelines and notifications. Between 11:15 UTC and about 13:00 UTC, pipelines were created but not processed, and outbound webhooks and transactional notifications (Slack and email) were not sent. What's impacted Customers running workflows and jobs, and customers viewing usage data on the Project Usage page. What can you expect - Delays starting workflows and jobs are improving. Most jobs are now starting within about 3 minutes, down from about 10 minutes at 13:50 UTC. - Some workflows from pipelines triggered during the incident are still waiting up to about 40 minutes to start. - Working through the backlog is taking slightly longer than expected. We now expect to clear it by about 14:45 UTC. - Workflows from pipelines triggered during the incident are still waiting to start, so retriggering may result in duplicate workflows. - Usage charge data on the Project Usage page is currently stale. Thank you for your patience while our engineers work to resolve this. Next update We will provide another update by 15:00 UTC. --- [2026-10-09T14:00:23.628Z] (identified) Since 11:15 UTC, customers have experienced disruption to pipelines and notifications. Between 11:15 UTC and about 13:00 UTC, pipelines were created but not processed, and outbound webhooks and transactional notifications (Slack and email) were not sent. What's impacted Customers running workflows and jobs, and customers viewing usage data on the Project Usage page. What can you expect - Pipelines are being processed again, including those triggered during the incident. - Since about 13:25 UTC, workflows and jobs have been starting with increasing delays as we work through the backlog. Most jobs are currently starting within about 10 minutes, and some are waiting up to about 25 minutes. - We expect to work through this backlog by about 14:35 UTC. - Workflows from pipelines triggered during the incident are still waiting to start, so retriggering may result in duplicate workflows. - Usage charge data on the Project Usage page is currently stale. Thank you for your patience while our engineers work to resolve this. Next update We will provide another update by 14:30 UTC. --- [2026-10-09T13:30:31.975Z] (identified) Between 11:15 UTC and about 13:00 UTC, pipelines were created but not processed, and outbound webhooks and transactional notifications (Slack and email) were not sent. What's impacted: Customers triggering pipelines, and customers relying on outbound webhooks and email and Slack notifications. What can you expect: - Since about 13:00 UTC, pipelines are being processed again and jobs are starting. - Most jobs are currently starting within about 4 minutes, and some are waiting up to about 10 minutes. Pipelines triggered during the incident are still being processed and may take longer to start. - Outbound webhooks and notifications have resumed. - Pipelines are still being processed from the backlog, so retriggering may result in duplicate pipelines. Thank you for your patience while our engineers work to resolve this. Next update: We will provide another update by 14:00 UTC. --- [2026-10-09T13:00:40.161Z] (identified) What's impacted: Customers triggering pipelines, and customers relying on outbound webhooks and email and Slack notifications. What can you expect Since 11:15 UTC: - Pipelines are still being created while this is ongoing but not yet processed, so retriggering may result in duplicate pipelines. - Outbound webhooks and transactional notifications (Slack and email) are not being sent. - Pipeline statuses may not be visible. - Usage charge data on the Project Usage page will be stale. Thank you for your patience while our engineers work to resolve this. Next update: We will provide another update by 13:30 UTC. --- [2026-10-09T11:33:00.542Z] (identified) There is a delay in our data infrastructure causing delays before build data and some notifications become visible. We are working on a mitigation.
Elevated wait times for machine jobsResolved Oct 8, 2026 · 3h 35mResolved
## **Summary** Between October 1 and October 8, 2026, CircleCI customers experienced five separate incidents in which jobs waited longer than normal to start, or failed to start. Three of the incidents were driven by limited instance availability from our cloud provider. One was caused by a configuration issue in our scheduling infrastructure, and one by memory limits on an internal service, which we raised during the incident. * **October 1, 07:57 to 10:24 UTC:** Docker Gen 2 jobs waited up to 30 minutes to start. A configuration issue prevented the Gen 2 scheduling servers from using their high-performance local disks. During a burst in traffic, the fallback disks saturated and job placement slowed. * **October 7, 13:27 to 14:06 UTC:** Machine jobs waited longer than normal because we could not acquire enough virtual machines from our cloud provider. * **October 7, 14:24 to 15:21 UTC:** Some jobs failed with an infrastructure error, and workflows were delayed. The service that gives each job its configuration at startup ran out of memory and restarted. * **October 7, 18:49 to 20:05 UTC:** Docker Gen 2 jobs waited longer than normal because we could not acquire enough instances from our cloud provider. * **October 8, 12:48 to 16:00 UTC:** Linux machine jobs and remote Docker jobs waited up to 50 minutes to start. Our cloud provider did not have enough of the instance type we request, and a defect in the logic that picks a different instance type kept us from using alternatives. Jobs that waited during these incidents were delayed, not lost. Customers whose jobs failed with an infrastructure error during the October 7 incident from 14:24 to 15:21 UTC can rerun them. The original status pages can be found below: * [October 1 - Delays starting Gen 2 Docker Jobs](https://status.circleci.com/incidents/22mrv46g6n7g) * [October 7 - Delay on starting Machine Job Tasks](https://status.circleci.com/incidents/tj6wdnfjwr64) * [October 7 - Elevated level of infra fails on customer jobs](https://status.circleci.com/incidents/0z31ldjd8m5v) * [October 7 - Increased task wait times for Docker Gen2](https://status.circleci.com/incidents/39rfrjmk5pvd) * [October 8 - Elevated wait times for machine jobs](https://status.circleci.com/incidents/jn756xtsy4xg) ## **Background** Every job that runs on CircleCI needs compute. Where that compute comes from depends on the executor. Docker jobs run on clusters of servers that schedule work onto a fleet of instances. The scheduler is the software on those servers that decides which instance runs each job. Gen 2 Docker jobs ran on one such cluster at the time of the October 1 incident. Gen 1 Docker jobs are spread across several clusters. Machine jobs, including remote Docker jobs, get a dedicated virtual machine. A remote Docker job is handled as a Linux machine job from submission through provisioning. Our machine provisioning service requests each machine from our cloud provider in one of the two regions we normally use for machine jobs. When the provider cannot supply the instance type we request, the service falls back: it requests a different instance type, or a different region, from a list we maintain. A job waits until a machine is available. When any job starts, it asks a job configuration service for its configuration. That service runs as a set of pods, which are small copies of the same program. If too many pods stop, the remaining pods carry more load, and jobs cannot start until the service recovers. These paths depend on three kinds of capacity: the throughput of our scheduling systems, the supply of instances from our cloud provider, and the memory and number of copies of the services that start jobs. A shortfall in any of them makes jobs wait or fail. ## **What Happened** _\(All times UTC\)_ ### **October 1: Docker Gen 2 jobs delayed up to 30 minutes** At 07:44, disk activity on the servers that schedule Gen 2 Docker jobs rose to 90% of capacity and stayed there. At 07:57, the number of Docker jobs waiting to start began to grow, and wait times rose with it. An automated alert paged the on-call engineer at 08:08. We declared an incident at 08:42. At 08:43, engineers identified high disk latency on the scheduling servers. They first believed the disks were too slow for the load, so they prepared a change to a faster disk type. At 09:01, the team agreed on two workstreams: move the servers to faster disks, and add a second Gen 2 cluster so that jobs would split across two clusters. At 09:41, engineers found the root cause. Each scheduling server has fast disks attached directly to it, and the scheduler is meant to keep its working data there. These servers run on a security-hardened operating system image, and on that image the fast disks were not being used. The scheduler wrote its data to slower network-attached disks instead. The script that sets up the fast disks failed without an error, and a gap in logging on these servers hid the failure. Gen 2 traffic had also become increasingly bursty. The slower disks saturated during those bursts, and the scheduler fell behind. Wait times peaked at 09:54, when jobs on the Gen 2 medium resource class waited up to 30 minutes. At 09:55, engineers applied a fix that set up the fast disks correctly. At 10:11, the scheduler began placing jobs faster, which confirmed the servers could keep up with new jobs. Wait times for Gen 2 medium fell to 13 minutes at 10:14, 7 minutes at 10:17, and 1 minute 30 seconds at 10:20. We moved the status page to monitoring at 10:24 and resolved the incident at 10:45. At 11:20, the second Gen 2 cluster began running jobs. ### **October 7, 13:27 to 14:06: Elevated wait times for machine jobs** We declared an incident at 13:27 after seeing elevated queueing and wait times. We posted to the status page at 13:31. At 13:31, engineers saw that we were struggling to acquire virtual machines for machine jobs, and that the number of jobs waiting to start was high mainly for machine jobs. At 13:34, wait times were elevated across executors. At 13:48, wait times began to recover as more machines became available, and we resolved the incident at 14:06. ### **October 7, 14:24 to 15:21: Job failures and workflow delays** Shortly after the previous incident was marked as resolved, we declared a new incident at 14:24 due to a sharp drop in the number of jobs submitted to our execution systems. At 14:26, engineers saw that pods of the job configuration service were restarting. Pods that restart leave fewer copies of the service to answer requests, and jobs that could not retrieve their configuration after repeated attempts failed with an infrastructure error. Workflows that were already running waited for updates from jobs that had not started. At 14:35, engineers increased the amount of available memory to the pods and increased the number of pods, which gave the service more copies to share the load. We also increased the amount of memory available Job submissions began recovering, and the number of messages flowing through our workflow system returned to normal levels. At 14:44, we moved the status page to monitoring. Pods in the job configuration service reached their memory limits and restarted. We increased memory allocations and pod counts, and the service recovered. At 15:13, we increased the number of instances in our Gen 2 Docker clusters so we could work through the backlog of waiting jobs quickly. We resolved the incident at 15:21. ### **October 7, 18:49 to 20:05: Elevated wait times for Docker Gen 2 jobs** We declared a new incident at 18:49 after Docker Gen 2 wait times rose. At 18:51, engineers confirmed that a shortage of available instances kept us from acquiring enough of them, which raised wait times. Wait times differed by resource class size. At 19:52, average wait times were lower than an hour earlier as more capacity became available, though brief spikes continued, and we moved the status page to monitoring. We resolved the incident at 20:05. ### **October 8: Linux machine and remote Docker jobs wait up to 50 minutes** We declared an incident at 12:48 after seeing a large number of jobs waiting to start and long wait times for machine and remote Docker jobs. At 12:49, engineers saw that our cloud provider could not supply enough instances. Delays continued to climb as we investigated. At 13:13, we began examining the logic that chooses a different instance type or region when our first choice is unavailable. At 13:24, we suspected that a defect in that logic was stopping us from using other instance types in our primary region, so we disabled our fallback region and confirmed the cause of the defect. Engineers began working on a fix and deployed it at 14:49. The service began requesting other instance types and started machines in large numbers. At 14:58, the queue of jobs waiting for a machine dropped steadily. Average wait times across resource classes continued to drop over the next hour as we worked through the backlog of waiting jobs. At 16:00, the queue of jobs waiting for a machine cleared and the number of waiting jobs returned to normal. We moved the status page to monitoring at 16:09 and resolved the incident at 16:35. ## **Future Prevention and Process Improvement** We are taking the following steps to strengthen the resilience of job start times, including supply diversification, fallback logic, and earlier detection. **We are expanding our compute regions.** In the October 7 and October 8 incidents, our cloud provider could not supply enough instances in the regions we used. We are adding new regions to our fallback logic in order handle spikes for demand that outpace our preferred regions. **We are fixing the logic that picks a different instance type.** On October 8, a defect kept the machine provisioning service from using other instance types in the same region. We deployed a fix during the incident and are now reviewing and refactoring this logic as a whole, with tests, so that a shortage of one instance type moves work to another. **We are expanding the instance types we offer and securing additional capacity.** We are working with our cloud provider to reserve additional capacity in our existing regions, and we are adding instance types that jobs can run on. Industry demand for high-performance compute is growing. We are diversifying where and how we source capacity so our customers are insulated from supply variation. **We added a second Gen 2 Docker cluster so traffic bursts are distributed across two clusters.** We are also adding alerts on scheduling server disk saturation and on how fast the scheduler places jobs, and an alert that fires when a server starts without its fast disks in use. We made the disk setup script fail with an error, and we restored log shipping from these servers so that a similar problem is visible sooner. Customer experience is our top priority, and we commit to continually improving the reliability of our systems to match the trust that our customers place in us. We thank our customers for their patience while our team worked to resolve these incidents. Please reach out to our support team with any questions or concerns. --- [2026-10-08T16:35:33.396Z] (resolved) Between 12:40 UTC and 16:00 UTC on October 8, customers using Linux machine jobs and remote Docker experienced elevated wait times. The issue has been resolved and wait times have returned to normal. We thank you for your patience while our team worked on implementing a fix. --- [2026-10-08T16:09:48.199Z] (monitoring) Wait times for customers using Linux machine jobs and remote Docker have returned to normal. We are monitoring to confirm wait times remain stable while we continue to add capacity. We will provide another update by 16:30 UTC. --- [2026-10-08T15:59:19.201Z] (identified) Wait times continue to decrease, but customers using Linux machine jobs and remote Docker are still experiencing delays. Wait times currently average about 90 seconds, with the longest waits up to about 6 minutes on some resource classes. We are working to add capacity as quickly as possible. We will provide another update by 16:30 UTC. --- [2026-10-08T15:31:10.070Z] (identified) Wait times continue to decrease, but customers using Linux machine jobs and remote Docker are still experiencing delays. Wait times currently average about 6 minutes. The longest waits, up to about 20 minutes, are on the 2xlarge, arm.2xlarge and gpu.nvidia.small resource classes. We are working to add capacity as quickly as possible. We will provide another update by 16:00 UTC. --- [2026-10-08T15:01:32.048Z] (identified) A fix has been deployed and wait times are decreasing, but customers using Linux machine jobs and Remote Docker are still experiencing delays. Wait times currently average about 11 minutes, with the longest waits exceeding 35 minutes on some resource classes. We are working to add capacity as quickly as possible. We will provide another update by 15:30 UTC. --- [2026-10-08T14:37:22.456Z] (identified) Customers using Linux machine jobs and remote Docker are experiencing elevated wait times. Wait times have started to decrease and now average about 20 minutes, with the longest waits exceeding 40 minutes on some resource classes. Our engineers have identified the issue and are working on a fix. We will provide another update by 15:00 UTC. --- [2026-10-08T14:04:11.147Z] (identified) Customers using Linux machine jobs are experiencing elevated wait times, averaging about 40 minutes, with the longest waits exceeding 50 minutes. Most Linux machine resource classes are affected, including medium, large, xlarge, 2xlarge and their Arm equivalents. Our engineers have identified the issue and are working on a fix. We will provide another update by 14:30 UTC. --- [2026-10-08T13:34:39.803Z] (identified) Customers using Linux machine jobs are experiencing elevated wait times, averaging about 11 minutes, with the longest waits exceeding 30 minutes on the medium, arm.medium and arm.large resource classes. Our engineers have identified the issue and are working on a fix. We will provide another update by 14:00 UTC. --- [2026-10-08T13:00:31.109Z] (identified) Customers may be experiencing elevated wait times for machine jobs. We are working to resolve this.
Increased task wait times for Docker Gen2Resolved Oct 7, 2026 · 1h 10mResolved
## **Summary** Between October 1 and October 8, 2026, CircleCI customers experienced five separate incidents in which jobs waited longer than normal to start, or failed to start. Three of the incidents were driven by limited instance availability from our cloud provider. One was caused by a configuration issue in our scheduling infrastructure, and one by memory limits on an internal service, which we raised during the incident. * **October 1, 07:57 to 10:24 UTC:** Docker Gen 2 jobs waited up to 30 minutes to start. A configuration issue prevented the Gen 2 scheduling servers from using their high-performance local disks. During a burst in traffic, the fallback disks saturated and job placement slowed. * **October 7, 13:27 to 14:06 UTC:** Machine jobs waited longer than normal because we could not acquire enough virtual machines from our cloud provider. * **October 7, 14:24 to 15:21 UTC:** Some jobs failed with an infrastructure error, and workflows were delayed. The service that gives each job its configuration at startup ran out of memory and restarted. * **October 7, 18:49 to 20:05 UTC:** Docker Gen 2 jobs waited longer than normal because we could not acquire enough instances from our cloud provider. * **October 8, 12:48 to 16:00 UTC:** Linux machine jobs and remote Docker jobs waited up to 50 minutes to start. Our cloud provider did not have enough of the instance type we request, and a defect in the logic that picks a different instance type kept us from using alternatives. Jobs that waited during these incidents were delayed, not lost. Customers whose jobs failed with an infrastructure error during the October 7 incident from 14:24 to 15:21 UTC can rerun them. The original status pages can be found below: * [October 1 - Delays starting Gen 2 Docker Jobs](https://status.circleci.com/incidents/22mrv46g6n7g) * [October 7 - Delay on starting Machine Job Tasks](https://status.circleci.com/incidents/tj6wdnfjwr64) * [October 7 - Elevated level of infra fails on customer jobs](https://status.circleci.com/incidents/0z31ldjd8m5v) * [October 7 - Increased task wait times for Docker Gen2](https://status.circleci.com/incidents/39rfrjmk5pvd) * [October 8 - Elevated wait times for machine jobs](https://status.circleci.com/incidents/jn756xtsy4xg) ## **Background** Every job that runs on CircleCI needs compute. Where that compute comes from depends on the executor. Docker jobs run on clusters of servers that schedule work onto a fleet of instances. The scheduler is the software on those servers that decides which instance runs each job. Gen 2 Docker jobs ran on one such cluster at the time of the October 1 incident. Gen 1 Docker jobs are spread across several clusters. Machine jobs, including remote Docker jobs, get a dedicated virtual machine. A remote Docker job is handled as a Linux machine job from submission through provisioning. Our machine provisioning service requests each machine from our cloud provider in one of the two regions we normally use for machine jobs. When the provider cannot supply the instance type we request, the service falls back: it requests a different instance type, or a different region, from a list we maintain. A job waits until a machine is available. When any job starts, it asks a job configuration service for its configuration. That service runs as a set of pods, which are small copies of the same program. If too many pods stop, the remaining pods carry more load, and jobs cannot start until the service recovers. These paths depend on three kinds of capacity: the throughput of our scheduling systems, the supply of instances from our cloud provider, and the memory and number of copies of the services that start jobs. A shortfall in any of them makes jobs wait or fail. ## **What Happened** _\(All times UTC\)_ ### **October 1: Docker Gen 2 jobs delayed up to 30 minutes** At 07:44, disk activity on the servers that schedule Gen 2 Docker jobs rose to 90% of capacity and stayed there. At 07:57, the number of Docker jobs waiting to start began to grow, and wait times rose with it. An automated alert paged the on-call engineer at 08:08. We declared an incident at 08:42. At 08:43, engineers identified high disk latency on the scheduling servers. They first believed the disks were too slow for the load, so they prepared a change to a faster disk type. At 09:01, the team agreed on two workstreams: move the servers to faster disks, and add a second Gen 2 cluster so that jobs would split across two clusters. At 09:41, engineers found the root cause. Each scheduling server has fast disks attached directly to it, and the scheduler is meant to keep its working data there. These servers run on a security-hardened operating system image, and on that image the fast disks were not being used. The scheduler wrote its data to slower network-attached disks instead. The script that sets up the fast disks failed without an error, and a gap in logging on these servers hid the failure. Gen 2 traffic had also become increasingly bursty. The slower disks saturated during those bursts, and the scheduler fell behind. Wait times peaked at 09:54, when jobs on the Gen 2 medium resource class waited up to 30 minutes. At 09:55, engineers applied a fix that set up the fast disks correctly. At 10:11, the scheduler began placing jobs faster, which confirmed the servers could keep up with new jobs. Wait times for Gen 2 medium fell to 13 minutes at 10:14, 7 minutes at 10:17, and 1 minute 30 seconds at 10:20. We moved the status page to monitoring at 10:24 and resolved the incident at 10:45. At 11:20, the second Gen 2 cluster began running jobs. ### **October 7, 13:27 to 14:06: Elevated wait times for machine jobs** We declared an incident at 13:27 after seeing elevated queueing and wait times. We posted to the status page at 13:31. At 13:31, engineers saw that we were struggling to acquire virtual machines for machine jobs, and that the number of jobs waiting to start was high mainly for machine jobs. At 13:34, wait times were elevated across executors. At 13:48, wait times began to recover as more machines became available, and we resolved the incident at 14:06. ### **October 7, 14:24 to 15:21: Job failures and workflow delays** Shortly after the previous incident was marked as resolved, we declared a new incident at 14:24 due to a sharp drop in the number of jobs submitted to our execution systems. At 14:26, engineers saw that pods of the job configuration service were restarting. Pods that restart leave fewer copies of the service to answer requests, and jobs that could not retrieve their configuration after repeated attempts failed with an infrastructure error. Workflows that were already running waited for updates from jobs that had not started. At 14:35, engineers increased the amount of available memory to the pods and increased the number of pods, which gave the service more copies to share the load. We also increased the amount of memory available Job submissions began recovering, and the number of messages flowing through our workflow system returned to normal levels. At 14:44, we moved the status page to monitoring. Pods in the job configuration service reached their memory limits and restarted. We increased memory allocations and pod counts, and the service recovered. At 15:13, we increased the number of instances in our Gen 2 Docker clusters so we could work through the backlog of waiting jobs quickly. We resolved the incident at 15:21. ### **October 7, 18:49 to 20:05: Elevated wait times for Docker Gen 2 jobs** We declared a new incident at 18:49 after Docker Gen 2 wait times rose. At 18:51, engineers confirmed that a shortage of available instances kept us from acquiring enough of them, which raised wait times. Wait times differed by resource class size. At 19:52, average wait times were lower than an hour earlier as more capacity became available, though brief spikes continued, and we moved the status page to monitoring. We resolved the incident at 20:05. ### **October 8: Linux machine and remote Docker jobs wait up to 50 minutes** We declared an incident at 12:48 after seeing a large number of jobs waiting to start and long wait times for machine and remote Docker jobs. At 12:49, engineers saw that our cloud provider could not supply enough instances. Delays continued to climb as we investigated. At 13:13, we began examining the logic that chooses a different instance type or region when our first choice is unavailable. At 13:24, we suspected that a defect in that logic was stopping us from using other instance types in our primary region, so we disabled our fallback region and confirmed the cause of the defect. Engineers began working on a fix and deployed it at 14:49. The service began requesting other instance types and started machines in large numbers. At 14:58, the queue of jobs waiting for a machine dropped steadily. Average wait times across resource classes continued to drop over the next hour as we worked through the backlog of waiting jobs. At 16:00, the queue of jobs waiting for a machine cleared and the number of waiting jobs returned to normal. We moved the status page to monitoring at 16:09 and resolved the incident at 16:35. ## **Future Prevention and Process Improvement** We are taking the following steps to strengthen the resilience of job start times, including supply diversification, fallback logic, and earlier detection. **We are expanding our compute regions.** In the October 7 and October 8 incidents, our cloud provider could not supply enough instances in the regions we used. We are adding new regions to our fallback logic in order handle spikes for demand that outpace our preferred regions. **We are fixing the logic that picks a different instance type.** On October 8, a defect kept the machine provisioning service from using other instance types in the same region. We deployed a fix during the incident and are now reviewing and refactoring this logic as a whole, with tests, so that a shortage of one instance type moves work to another. **We are expanding the instance types we offer and securing additional capacity.** We are working with our cloud provider to reserve additional capacity in our existing regions, and we are adding instance types that jobs can run on. Industry demand for high-performance compute is growing. We are diversifying where and how we source capacity so our customers are insulated from supply variation. **We added a second Gen 2 Docker cluster so traffic bursts are distributed across two clusters.** We are also adding alerts on scheduling server disk saturation and on how fast the scheduler places jobs, and an alert that fires when a server starts without its fast disks in use. We made the disk setup script fail with an error, and we restored log shipping from these servers so that a similar problem is visible sooner. Customer experience is our top priority, and we commit to continually improving the reliability of our systems to match the trust that our customers place in us. We thank our customers for their patience while our team worked to resolve these incidents. Please reach out to our support team with any questions or concerns. --- [2026-10-07T20:05:32.007Z] (resolved) This incident has been resolved. --- [2026-10-07T19:52:41.159Z] (monitoring) We're still seeing spikes in task wait times but on average wait times are looking better for all Docker Gen2. --- [2026-10-07T19:00:45.704Z] (identified) Note: delays will be more noticeable for people running on Docker `xlarge` and above. --- [2026-10-07T18:54:58.747Z] (identified) There is an increased wait time for jobs on Docker Gen2.
CircleCI Uptime History
Daily uptime recorded by PulsAPI over the last 90 days, with the days that carried a CircleCI incident marked. Each day is measured from our own checks, so the record continues even when the vendor publishes nothing.
Going further back: CircleCI uptime by month and every recorded CircleCI outage.
Live CircleCI Outage Reports
Problems reported by engineers using CircleCI. Reports often appear 10 to 20 minutes before the official status page moves, and stay on this page for 90 days.
No CircleCI problems reported yet.
Something wrong with CircleCI on your end?
Reports are anonymous, take about ten seconds, and help the next engineer who checks.
About CircleCI
CircleCI is a continuous integration and continuous deployment (CI/CD) platform running over one million builds per day for engineering teams. CircleCI outages stall software delivery pipelines. Track CircleCI status and outages here.
Over the past 30 days, PulsAPI has measured CircleCI uptime at 99.87%. Response times, incident history, and SLA tracking for CircleCI are available on the PulsAPI dashboard.
CircleCI Status FAQ
Is CircleCI down right now?
The verdict at the top of this page comes from the latest poll of the official CircleCI status feed and refreshes automatically, so you are not looking at a stale cached answer. PulsAPI re-checks a vendor every 60 seconds whenever it is not fully operational, or has changed status in the last 30 minutes, so an incident is always on the fastest cadence. Vendors that are stable and rarely watched are re-checked every 5 to 20 minutes instead.
What should I do if CircleCI is not working?
Check the component list on this page first to see if your failure matches a known outage. If it does, you can report the issue for other engineers, follow the incident timeline for vendor updates, and run any fallback your app supports until CircleCI recovers.
How do I check CircleCI uptime history?
PulsAPI tracks CircleCI uptime over 30, 60, and 90-day rolling windows. The current 30-day uptime is 99.87%. The 90-day uptime strip on this page is public: no account needed. A free account adds alerts whenever CircleCI changes state, plus 15 days of incident detail; Pro extends that to 90 days and Business to 13 months.
How fast does PulsAPI detect CircleCI outages?
PulsAPI reads the official CircleCI status page. PulsAPI re-checks a vendor every 60 seconds whenever it is not fully operational, or has changed status in the last 30 minutes, so an incident is always on the fastest cadence. Vendors that are stable and rarely watched are re-checked every 5 to 20 minutes instead. Alerts reach Slack, Discord, Microsoft Teams, PagerDuty, or email within seconds of detection.
Does PulsAPI track individual CircleCI components?
Yes. PulsAPI monitors 39 individual CircleCI components, so you can see which service or region is affected instead of a generic "service degraded" banner.
Can I get alerts when CircleCI goes down?
Yes. Create a free account and subscribe to CircleCI: the free plan covers three services with unlimited email alerts and no monthly notification cap. Paid plans lift the monitor limit entirely (Pro and above do not meter how many vendors you watch) and route the same alerts into Slack, Discord, Microsoft Teams, PagerDuty and webhooks.
How do I report an issue with CircleCI to the community?
Click "Report an issue" at the top of this page, pick the issue type and severity, and optionally add your region and a short note. Reports are anonymous and help other engineers spot CircleCI problems early. Official status pages often lag real incidents by 10 to 20 minutes.
Is there a free way to monitor CircleCI?
Yes. The free plan covers 3 monitored services with unlimited email alerts, permanently, so CircleCI plus two more costs nothing. Every new account also opens with a 30-day Business trial (unlimited monitors, Slack alerts, SLA reporting) with no credit card.
Track CircleCI with the rest of your stack
2463+ vendors in one dashboard, and an email the moment any of them changes state. Three free, and no monitor limit at all from $29/mo, unlike the alternatives, which keep charging by the monitor.
Start free, 3 services