DataRobot Outage History
36 incidents recorded over the last 365 days, 31 of them resolved. Durations are measured from when PulsAPI first saw the problem to when it cleared, which is usually longer than the vendor's own figure.
Every Recorded DataRobot Incident
| Date | Incident | Duration | Status |
|---|---|---|---|
| Aug 18, 2026 | Codespace Session Validation Failure in Generative AI PlaygroundThe fix is now deployed to all the Prod environments and Generative AI Playground is now working as expected. | 2h 24m | Resolved |
| Jul 24, 2026 | Issues with Vector Database CreationEngineering team observed that the Vector Database creation is operational across all DataRobot MTS environments. The incident is Resolved. | 1h 10m | Resolved |
| Jul 2, 2026 | Issues with prediction requests against custom model deploymentsThe issue is contained and the team has verified that predictions against Custom Models in all MTSaaS environments are healthy. | 34 min | Resolved |
| Jun 24, 2026 | Managed US AI Cloud Model Monitoring Planned MaintenanceThe scheduled maintenance has been completed. | 16h 0m | Scheduled |
| Jun 17, 2026 | Managed EU AI Cloud Model Monitoring Planned MaintenanceThe scheduled maintenance has been completed. | 2h 0m | Scheduled |
| Jun 3, 2026 | Managed Japan AI Cloud Model Monitoring Planned MaintenanceDataRobot is performing an infrastructure maintenance on Managed Japan AI Cloud (app.jp.datarobot.com) on June 3rd, 2026, between 5:00 PM and 9:00 PM UTC. During this window, users may experience intermittent… | 4h 0m | Scheduled |
| Jun 3, 2026 | Managed Japan AI Cloud Model Monitoring Planned MaintenanceThe scheduled maintenance has been completed. | 4h 0m | Scheduled |
| May 29, 2026 | Batch Jobs queued up in US MTSA fix has been deployed to production and the issue is now resolved. All systems are operating normally. | 120h 52m | Resolved |
| May 28, 2026 | Delay in Feature Drift Statistics ProcessingFeature drift statistics processing has been restored to normal. All metrics are now up to date. No data was lost. | 1h 17m | Resolved |
| May 26, 2026 | Customers Experiencing Errors with New Custom Model Creation.Custom Models, Custom Applications & Data upload services are back to Operational state in US SAAS Environment. Issue is Resolved. | 1h 9m | Resolved |
| May 8, 2026 | Widespread intermittent service issues for new workloads in US ProductionEngineering confirmed the issue is resolved and all services are restored. | 20h 5m | Resolved |
| Apr 10, 2026 | Delay in processing actual messagesEngineering has applied the required infrastructure configuration changes. The service is operating normally and no further user impact is observed. Engineering will continue monitoring cluster health to ensure… | 36 min | Resolved |
| Apr 9, 2026 | Elevated Errors on Managed AI CloudEngineering has implemented the required fixes to resolve the elevated error rates. Services are now operating normally, and no further user impact has been observed. The team will continue to monitor the system to… | 15h 7m | Resolved |
| Mar 30, 2026 | Degraded Performance on DataRobot MTS due to Quay outageQuay.io functionality has been restored and DataRobot environments are fully stabilized. | 12h 10m | Resolved |
| Mar 13, 2026 | Performance Degradation on Managed AI CloudThis incident has been resolved. | 54 min | Resolved |
| Mar 11, 2026 | Intermittent UI disruptions on Managed AI CloudThis incident has been resolved. | 142h 26m | Resolved |
| Mar 11, 2026 | Network issue related to Kubernetes in US clusterThe mitigation implemented by Engineering has resolved the Kubernetes network issue, and the incident is now contained. | 2h 5m | Resolved |
| Feb 18, 2026 | Degraded Performance on the DataRobot MTS due to Quay outageThis incident is now resolved. | 15 min | Resolved |
| Feb 17, 2026 | LLM blueprints deployments cannot be createdRollback of the JP cluster to the previous version is complete and the problem has been mitigated. | 7 min | Resolved |
| Feb 16, 2026 | Agent Application Template Impacted After Moderations Library Upgrade.New version of Agentic application template is released, the issue is resolved | 1h 13m | Resolved |
| Feb 13, 2026 | Degraded Performance on the DataRobot US MTSThe incident has now been resolved. All services are now operational. | 52 min | Resolved |
| Jan 8, 2026 | Problem connecting to DataRobot due to client browser cachingThis incident has been resolved. | 332h 3m | Resolved |
| Dec 29, 2025 | DataRobot LLM Gateway OpenAI models. Planned MaintenanceThe scheduled maintenance has been completed. | 2h 0m | Scheduled |
| Dec 17, 2025 | Degraded Platform Performance Due to Docker Hub OutageBetween 16:50 UTC and 17:09 UTC, one of external providers(DockerHub) had an outage that might have caused temporary delays in starting platform workloads. The issue has been resolved and normal operations have resumed… | -41 min | Resolved |
| Dec 5, 2025 | Degraded Platform Performance Due to Docker Hub OutageThe engineering team has confirmed recovery across all MTS and STS environments. No further updates are expected. | 25 min | Resolved |
| Nov 26, 2025 | Some Users Facing Issues In Accessing Notebooks/Codespaces in US/EU/JP MTS.This incident has been resolved. | 25h 8m | Resolved |
| Nov 20, 2025 | Some Users Are Unable To Login Into US,EU & JP MTS clusters.The issue is fixed. Access to the US, EU, and JP clusters is restored. | 30 min | Resolved |
| Nov 7, 2025 | The AutoML experiment creation workflow and visibility of certain tabs impacted in MTS environmentsThe issue has been resolved. | 4h 35m | Resolved |
| Nov 6, 2025 | Notebooks are timing out occasionally on US Production.The Engineering team hasn't observed any new occurrences during monitoring. All services remain fully operational. | 23h 59m | Resolved |
| Nov 4, 2025 | Data Connectors Hanging on CreationThe Engineering team has successfully resolved the connection and processing timeouts affecting data operations in MTS Production. Services are now running stable. | 35 min | Resolved |
| Oct 22, 2025 | Temporary Access Issue for Org Admins to Custom ApplicationsThis incident has been resolved. | 1h 24m | Resolved |
| Oct 20, 2025 | AWS services outage affects STS and MTS DataRobot environmentsThis incident has been resolved. AWS is reporting systems as operational. Engineering will continue to monitor the situation. | 12h 22m | Resolved |
| Oct 1, 2025 | UI Issue Affecting Model Visibility and Project ManagementEngineering team has applied the fix to address UI issue affecting Model Visibility and Project Management. This problem has now been resolved. | 8h 24m | Resolved |
| Sep 25, 2025 | Issue With DockerHub Services Impacting MTS & STS DataRobot Clusters.Issue is resolved and DataRobot MTS & STS services are back to normal. | 37 min | Resolved |
| Sep 9, 2025 | Users in APAC region are experiencing network issues when accessing application endpoints in US MTSThis incident has been resolved. | 36h 24m | Resolved |
| Sep 2, 2025 | Serverless Prediction Servers are impacted when an a new deployment is created or an existing deployment is modified.Engineering has applied the fix to mitigate the issue impacting serverless deployments when a new deployment is created or an existing deployment is modified. The problem has now been contained. | 2h 30m | Resolved |
How to Read This DataRobot Incident Log
Each row is an incident PulsAPI observed, not a summary written afterwards. The duration is wall-clock time between the first failing check and the first clean one, so it includes the window before DataRobot acknowledged anything. Vendor post-mortems typically measure from acknowledgement, which is why their numbers are usually shorter.
Incidents still open have no duration yet and are listed as ongoing rather than being given a running total. The archive covers the last 365 days; anything older has aged out of the window rather than never having happened.
The longest single DataRobot outage in this window ran 13d 20h, against a mean recovery of 22h 52m. If you depend on DataRobot in a customer-facing path, the longest figure is the one to design around. The mean is what happens on a normal bad day; the maximum is what happens on the worst one.