Replicate Outage History
39 incidents recorded over the last 365 days, 39 of them resolved. Durations are measured from when PulsAPI first saw the problem to when it cleared, which is usually longer than the vendor's own figure.
Every Recorded Replicate Incident
| Date | Incident | Duration | Status |
|---|---|---|---|
| Aug 10, 2026 | Delayed scaling due to node failureAll affected models have been scaling correctly for more than 30 minutes at this point, and we see no residual prediction queues. Thank you for your patience! | 40 min | Resolved |
| Aug 5, 2026 | Significant degradationWe have been monitoring for a while and services seem to be back. Thank you for your patience | 2h 3m | Resolved |
| Aug 2, 2026 | API degraded for A100sWe have restarted the affected component and service should be back to normal. Thank you for your patience | 8 min | Resolved |
| Jul 31, 2026 | Degraded scale-out due to failed setups pulling from huggingfaceWe're looking stable now, and with caching back on. Thanks again for your patience! | 2h 14m | Resolved |
| Jul 22, 2026 | Hitting GPU Capacity for H100s creating large queue times for some modelsThis issue is resolved | 26h 28m | Resolved |
| Jul 16, 2026 | HuggingFace download issuesWe have not seen elevated rates of HuggingFace/model setup errors for a couple hours now, so we believe this incident is cleared | 7h 32m | Resolved |
| Jul 13, 2026 | H100 GPU shortage resulting in high queue timesH100 capacity has returned to normal levels | 25h 1m | Resolved |
| Jul 10, 2026 | High contention on H100 hardwareWe are back under max capacity for H100 hardware. Thank you for your patience! | 8h 13m | Resolved |
| Jun 30, 2026 | Limited H100 capacityThe H100 capacity is back in good health. Thanks for your patience! | 5h 34m | Resolved |
| Jun 3, 2026 | We're seeing long setup times and high contention for models on some L40S and H200 clusters.System is back to operating normally | 1h 6m | Resolved |
| May 28, 2026 | Degraded performance on flux-2-klein-4bThis issue has been resolved and queue times are back to normal | 1h 47m | Resolved |
| May 21, 2026 | Prediction and Training status updates delayedMessage flows are healthy. | 1h 50m | Resolved |
| May 21, 2026 | Constrained H100 capacityH100 hardware contention has resolved. Thank you for your patience! | 6h 56m | Resolved |
| May 12, 2026 | Constrained capacity for H100 hardwareThere is no more contention for H100 hardware. Thank you for your patience! | 4h 16m | Resolved |
| Apr 19, 2026 | Degraded A100 hardwareAll A100 capacity is back. Thanks for your patience! | 26 min | Resolved |
| Apr 9, 2026 | A100 capacity unavailable during storage maintenanceThe maintenance is complete and all systems are reporting healthy. Thank you for your patience! | 47 min | Resolved |
| Mar 23, 2026 | Downstream errors for Black Forest Labs modelsBlack Forest Labs has resolved the issue. | 9h 37m | Resolved |
| Mar 10, 2026 | Degraded performance on Flux SchnellFlux Schnell requests are being served normally. | 5h 21m | Resolved |
| Feb 20, 2026 | Model Predictions Stuck at "Starting"Models are once again operational | 59 min | Resolved |
| Jan 26, 2026 | Increased setup failures for T4 modelsModel setup failures have dropped to normal levels since 10:40 UTC. Thank you for your patience! | 7h 43m | Resolved |
| Jan 20, 2026 | Predictions and training unavailable for multiple modelsAll infrastructure problems have been resolved and model predictions and training are available again for the relevant hardware types. | 1h 29m | Resolved |
| Jan 20, 2026 | Flux Schnell unavailableThe `black-forest-labs/flux-schnell` model is fully operational now. Thank you for your patience! | 3h 9m | Resolved |
| Jan 15, 2026 | Prediction ErrorsIncident is resolved. | 2h 55m | Resolved |
| Dec 18, 2025 | High demand for H100 hardware typeWe are all clear now. Thank you for your patience! | 34h 4m | Resolved |
| Dec 11, 2025 | Limited availability of L40S hardwareWe are back under capacity again. Thank you for your patience! | 1h 11m | Resolved |
| Nov 18, 2025 | Global network outageWe have seen stable behavior for the past 90 minutes. Thanks for your patience! | 6h 44m | Resolved |
| Nov 13, 2025 | sora-2-pro currently unavailableThis has been resolved. | 106h 51m | Resolved |
| Oct 29, 2025 | Downstream Service DisruptionThe downstream issues have been mitigated/abated. | 21h 0m | Resolved |
| Oct 22, 2025 | Luma models not runningAll luma models are operational again | 1h 59m | Resolved |
| Oct 21, 2025 | Intermittent issues with `cog push` with large imagesWe've confirmed that the issue is resolved. | 2h 10m | Resolved |
| Oct 20, 2025 | Replicate Platform OutageWe are seeing consistently healthy behavior across all systems now. Thank you for your patience! | 7h 17m | Resolved |
| Oct 20, 2025 | Widespread service degradationWe have applied all mitigations as recommended by external service providers and we are now seeing consistently healthy behavior. Thanks for your patience! | 12h 6m | Resolved |
| Sep 30, 2025 | Heygen models outageOur heygen models are back to fully operational | 47h 17m | Resolved |
| Sep 29, 2025 | Google Models are downGoogle models such as nano banana, imagen, veo-3 are currently back to full operation | 2h 21m | Resolved |
| Sep 29, 2025 | Low inbound network transfer speed affecting multiple systems | 1h 54m | Resolved |
| Sep 26, 2025 | Users unable to purchase prepaid credit | 32 min | Resolved |
| Sep 25, 2025 | Some degraded performance on official models | 18h 13m | Resolved |
| Sep 1, 2025 | Topaz official models are unavailable | 2h 58m | Resolved |
| Aug 24, 2025 | Issues with instances booting | 1h 1m | Resolved |
How to Read This Replicate Incident Log
Each row is an incident PulsAPI observed, not a summary written afterwards. The duration is wall-clock time between the first failing check and the first clean one, so it includes the window before Replicate acknowledged anything. Vendor post-mortems typically measure from acknowledgement, which is why their numbers are usually shorter.
Incidents still open have no duration yet and are listed as ongoing rather than being given a running total. The archive covers the last 365 days; anything older has aged out of the window rather than never having happened.
The longest single Replicate outage in this window ran 4d 10h, against a mean recovery of 10h 6m. If you depend on Replicate in a customer-facing path, the longest figure is the one to design around. The mean is what happens on a normal bad day; the maximum is what happens on the worst one.