Replicate Outage History

39 incidents recorded over the last 365 days, 39 of them resolved. Durations are measured from when PulsAPI first saw the problem to when it cleared, which is usually longer than the vendor's own figure.

Every Recorded Replicate Incident

Replicate incidents, newest first, with date, duration and status.
DateIncidentDurationStatus
Aug 10, 2026Delayed scaling due to node failureAll affected models have been scaling correctly for more than 30 minutes at this point, and we see no residual prediction queues. Thank you for your patience!40 minResolved
Aug 5, 2026Significant degradationWe have been monitoring for a while and services seem to be back. Thank you for your patience2h 3mResolved
Aug 2, 2026API degraded for A100sWe have restarted the affected component and service should be back to normal. Thank you for your patience8 minResolved
Jul 31, 2026Degraded scale-out due to failed setups pulling from huggingfaceWe're looking stable now, and with caching back on. Thanks again for your patience!2h 14mResolved
Jul 22, 2026Hitting GPU Capacity for H100s creating large queue times for some modelsThis issue is resolved26h 28mResolved
Jul 16, 2026HuggingFace download issuesWe have not seen elevated rates of HuggingFace/model setup errors for a couple hours now, so we believe this incident is cleared7h 32mResolved
Jul 13, 2026H100 GPU shortage resulting in high queue timesH100 capacity has returned to normal levels25h 1mResolved
Jul 10, 2026High contention on H100 hardwareWe are back under max capacity for H100 hardware. Thank you for your patience!8h 13mResolved
Jun 30, 2026Limited H100 capacityThe H100 capacity is back in good health. Thanks for your patience!5h 34mResolved
Jun 3, 2026We're seeing long setup times and high contention for models on some L40S and H200 clusters.System is back to operating normally1h 6mResolved
May 28, 2026Degraded performance on flux-2-klein-4bThis issue has been resolved and queue times are back to normal1h 47mResolved
May 21, 2026Prediction and Training status updates delayedMessage flows are healthy.1h 50mResolved
May 21, 2026Constrained H100 capacityH100 hardware contention has resolved. Thank you for your patience!6h 56mResolved
May 12, 2026Constrained capacity for H100 hardwareThere is no more contention for H100 hardware. Thank you for your patience!4h 16mResolved
Apr 19, 2026Degraded A100 hardwareAll A100 capacity is back. Thanks for your patience!26 minResolved
Apr 9, 2026A100 capacity unavailable during storage maintenanceThe maintenance is complete and all systems are reporting healthy. Thank you for your patience!47 minResolved
Mar 23, 2026Downstream errors for Black Forest Labs modelsBlack Forest Labs has resolved the issue.9h 37mResolved
Mar 10, 2026Degraded performance on Flux SchnellFlux Schnell requests are being served normally.5h 21mResolved
Feb 20, 2026Model Predictions Stuck at "Starting"Models are once again operational59 minResolved
Jan 26, 2026Increased setup failures for T4 modelsModel setup failures have dropped to normal levels since 10:40 UTC. Thank you for your patience!7h 43mResolved
Jan 20, 2026Predictions and training unavailable for multiple modelsAll infrastructure problems have been resolved and model predictions and training are available again for the relevant hardware types.1h 29mResolved
Jan 20, 2026Flux Schnell unavailableThe `black-forest-labs/flux-schnell` model is fully operational now. Thank you for your patience!3h 9mResolved
Jan 15, 2026Prediction ErrorsIncident is resolved.2h 55mResolved
Dec 18, 2025High demand for H100 hardware typeWe are all clear now. Thank you for your patience!34h 4mResolved
Dec 11, 2025Limited availability of L40S hardwareWe are back under capacity again. Thank you for your patience!1h 11mResolved
Nov 18, 2025Global network outageWe have seen stable behavior for the past 90 minutes. Thanks for your patience!6h 44mResolved
Nov 13, 2025sora-2-pro currently unavailableThis has been resolved.106h 51mResolved
Oct 29, 2025Downstream Service DisruptionThe downstream issues have been mitigated/abated.21h 0mResolved
Oct 22, 2025Luma models not runningAll luma models are operational again1h 59mResolved
Oct 21, 2025Intermittent issues with `cog push` with large imagesWe've confirmed that the issue is resolved.2h 10mResolved
Oct 20, 2025Replicate Platform OutageWe are seeing consistently healthy behavior across all systems now. Thank you for your patience!7h 17mResolved
Oct 20, 2025Widespread service degradationWe have applied all mitigations as recommended by external service providers and we are now seeing consistently healthy behavior. Thanks for your patience!12h 6mResolved
Sep 30, 2025Heygen models outageOur heygen models are back to fully operational47h 17mResolved
Sep 29, 2025Google Models are downGoogle models such as nano banana, imagen, veo-3 are currently back to full operation2h 21mResolved
Sep 29, 2025Low inbound network transfer speed affecting multiple systems1h 54mResolved
Sep 26, 2025Users unable to purchase prepaid credit32 minResolved
Sep 25, 2025Some degraded performance on official models18h 13mResolved
Sep 1, 2025Topaz official models are unavailable2h 58mResolved
Aug 24, 2025Issues with instances booting1h 1mResolved

How to Read This Replicate Incident Log

Each row is an incident PulsAPI observed, not a summary written afterwards. The duration is wall-clock time between the first failing check and the first clean one, so it includes the window before Replicate acknowledged anything. Vendor post-mortems typically measure from acknowledgement, which is why their numbers are usually shorter.

Incidents still open have no duration yet and are listed as ongoing rather than being given a running total. The archive covers the last 365 days; anything older has aged out of the window rather than never having happened.

The longest single Replicate outage in this window ran 4d 10h, against a mean recovery of 10h 6m. If you depend on Replicate in a customer-facing path, the longest figure is the one to design around. The mean is what happens on a normal bad day; the maximum is what happens on the worst one.