DataRobot Outage History

36 incidents recorded over the last 365 days, 31 of them resolved. Durations are measured from when PulsAPI first saw the problem to when it cleared, which is usually longer than the vendor's own figure.

Every Recorded DataRobot Incident

DataRobot incidents, newest first, with date, duration and status.
DateIncidentDurationStatus
Aug 18, 2026Codespace Session Validation Failure in Generative AI PlaygroundThe fix is now deployed to all the Prod environments and Generative AI Playground is now working as expected.2h 24mResolved
Jul 24, 2026Issues with Vector Database CreationEngineering team observed that the Vector Database creation is operational across all DataRobot MTS environments. The incident is Resolved.1h 10mResolved
Jul 2, 2026Issues with prediction requests against custom model deploymentsThe issue is contained and the team has verified that predictions against Custom Models in all MTSaaS environments are healthy.34 minResolved
Jun 24, 2026Managed US AI Cloud Model Monitoring Planned MaintenanceThe scheduled maintenance has been completed.16h 0mScheduled
Jun 17, 2026Managed EU AI Cloud Model Monitoring Planned MaintenanceThe scheduled maintenance has been completed.2h 0mScheduled
Jun 3, 2026Managed Japan AI Cloud Model Monitoring Planned MaintenanceDataRobot is performing an infrastructure maintenance on Managed Japan AI Cloud (app.jp.datarobot.com) on June 3rd, 2026, between 5:00 PM and 9:00 PM UTC. During this window, users may experience intermittent…4h 0mScheduled
Jun 3, 2026Managed Japan AI Cloud Model Monitoring Planned MaintenanceThe scheduled maintenance has been completed.4h 0mScheduled
May 29, 2026Batch Jobs queued up in US MTSA fix has been deployed to production and the issue is now resolved. All systems are operating normally.120h 52mResolved
May 28, 2026Delay in Feature Drift Statistics ProcessingFeature drift statistics processing has been restored to normal. All metrics are now up to date. No data was lost.1h 17mResolved
May 26, 2026Customers Experiencing Errors with New Custom Model Creation.Custom Models, Custom Applications & Data upload services are back to Operational state in US SAAS Environment. Issue is Resolved.1h 9mResolved
May 8, 2026Widespread intermittent service issues for new workloads in US ProductionEngineering confirmed the issue is resolved and all services are restored.20h 5mResolved
Apr 10, 2026Delay in processing actual messagesEngineering has applied the required infrastructure configuration changes. The service is operating normally and no further user impact is observed. Engineering will continue monitoring cluster health to ensure…36 minResolved
Apr 9, 2026Elevated Errors on Managed AI CloudEngineering has implemented the required fixes to resolve the elevated error rates. Services are now operating normally, and no further user impact has been observed. The team will continue to monitor the system to…15h 7mResolved
Mar 30, 2026Degraded Performance on DataRobot MTS due to Quay outageQuay.io functionality has been restored and DataRobot environments are fully stabilized.12h 10mResolved
Mar 13, 2026Performance Degradation on Managed AI CloudThis incident has been resolved.54 minResolved
Mar 11, 2026Intermittent UI disruptions on Managed AI CloudThis incident has been resolved.142h 26mResolved
Mar 11, 2026Network issue related to Kubernetes in US clusterThe mitigation implemented by Engineering has resolved the Kubernetes network issue, and the incident is now contained.2h 5mResolved
Feb 18, 2026Degraded Performance on the DataRobot MTS due to Quay outageThis incident is now resolved.15 minResolved
Feb 17, 2026LLM blueprints deployments cannot be createdRollback of the JP cluster to the previous version is complete and the problem has been mitigated.7 minResolved
Feb 16, 2026Agent Application Template Impacted After Moderations Library Upgrade.New version of Agentic application template is released, the issue is resolved1h 13mResolved
Feb 13, 2026Degraded Performance on the DataRobot US MTSThe incident has now been resolved. All services are now operational.52 minResolved
Jan 8, 2026Problem connecting to DataRobot due to client browser cachingThis incident has been resolved.332h 3mResolved
Dec 29, 2025DataRobot LLM Gateway OpenAI models. Planned MaintenanceThe scheduled maintenance has been completed.2h 0mScheduled
Dec 17, 2025Degraded Platform Performance Due to Docker Hub OutageBetween 16:50 UTC and 17:09 UTC, one of external providers(DockerHub) had an outage that might have caused temporary delays in starting platform workloads. The issue has been resolved and normal operations have resumed…-41 minResolved
Dec 5, 2025Degraded Platform Performance Due to Docker Hub OutageThe engineering team has confirmed recovery across all MTS and STS environments. No further updates are expected.25 minResolved
Nov 26, 2025Some Users Facing Issues In Accessing Notebooks/Codespaces in US/EU/JP MTS.This incident has been resolved.25h 8mResolved
Nov 20, 2025Some Users Are Unable To Login Into US,EU & JP MTS clusters.The issue is fixed. Access to the US, EU, and JP clusters is restored.30 minResolved
Nov 7, 2025The AutoML experiment creation workflow and visibility of certain tabs impacted in MTS environmentsThe issue has been resolved.4h 35mResolved
Nov 6, 2025Notebooks are timing out occasionally on US Production.The Engineering team hasn't observed any new occurrences during monitoring. All services remain fully operational.23h 59mResolved
Nov 4, 2025Data Connectors Hanging on CreationThe Engineering team has successfully resolved the connection and processing timeouts affecting data operations in MTS Production. Services are now running stable.35 minResolved
Oct 22, 2025Temporary Access Issue for Org Admins to Custom ApplicationsThis incident has been resolved.1h 24mResolved
Oct 20, 2025AWS services outage affects STS and MTS DataRobot environmentsThis incident has been resolved. AWS is reporting systems as operational. Engineering will continue to monitor the situation.12h 22mResolved
Oct 1, 2025UI Issue Affecting Model Visibility and Project ManagementEngineering team has applied the fix to address UI issue affecting Model Visibility and Project Management. This problem has now been resolved.8h 24mResolved
Sep 25, 2025Issue With DockerHub Services Impacting MTS & STS DataRobot Clusters.Issue is resolved and DataRobot MTS & STS services are back to normal.37 minResolved
Sep 9, 2025Users in APAC region are experiencing network issues when accessing application endpoints in US MTSThis incident has been resolved.36h 24mResolved
Sep 2, 2025Serverless Prediction Servers are impacted when an a new deployment is created or an existing deployment is modified.Engineering has applied the fix to mitigate the issue impacting serverless deployments when a new deployment is created or an existing deployment is modified. The problem has now been contained.2h 30mResolved

How to Read This DataRobot Incident Log

Each row is an incident PulsAPI observed, not a summary written afterwards. The duration is wall-clock time between the first failing check and the first clean one, so it includes the window before DataRobot acknowledged anything. Vendor post-mortems typically measure from acknowledgement, which is why their numbers are usually shorter.

Incidents still open have no duration yet and are listed as ongoing rather than being given a running total. The archive covers the last 365 days; anything older has aged out of the window rather than never having happened.

The longest single DataRobot outage in this window ran 13d 20h, against a mean recovery of 22h 52m. If you depend on DataRobot in a customer-facing path, the longest figure is the one to design around. The mean is what happens on a normal bad day; the maximum is what happens on the worst one.