Honeycomb Outage History
48 incidents recorded over the last 365 days, 44 of them resolved. Durations are measured from when PulsAPI first saw the problem to when it cleared, which is usually longer than the vendor's own figure.
Every Recorded Honeycomb Incident
| Date | Incident | Duration | Status |
|---|---|---|---|
| Aug 21, 2026 | Canvas not loadingThis incident has been resolved. | 37 min | Resolved |
| Aug 10, 2026 | US1 cold query outageThe fix is in place and we've confirmed resolution. Queries and SLOs are processing normally. | 7h 33m | Resolved |
| Jul 31, 2026 | Querying issuesStarting at 9:50 AM Central Time on July 31st we received alerts around querying health and latency. The cause was eventually tracked down to a clash of two different factors. First, we have been introducing Lambda… | -72 min | Resolved |
| Jul 27, 2026 | Triggers intermittently failing in EU regionThe degradation in querying and triggers has been resolved and we're back to full functionality. | 29h 16m | Resolved |
| Jul 2, 2026 | honeycomb.io marketing website not workingThis incident has been resolved. | 6h 20m | Resolved |
| Jun 29, 2026 | Activity Log delayed in USBackfill of the Activity Log events has completed. All events during the affected period should be present. | 2h 39m | Resolved |
| Jun 25, 2026 | Ingest outage in EUThis incident has been resolved. | 2h 31m | Resolved |
| Jun 24, 2026 | Ingest Outage in US and EUOn June 24, we experienced approximately 30 minutes of severe data ingestion degradation in Honeycomb’s US and EU instances. During this degradation, Honeycomb rejected a significant percentage of inbound telemetry, and… | 50 min | Resolved |
| Jun 19, 2026 | MCP tool access degradedMCP Write Tools have been restored. MCP users may need to reconnect to see all available tools. | 41 min | Resolved |
| Jun 18, 2026 | Enhance IssuesBeginning June 17th, enhance-related requests were failing to complete. We've since identified and deployed a fix, and enhance is now fully operational. Thank you for your patience. | 0 min | Resolved |
| Jun 17, 2026 | Querying IssuesWe have identified the root cause of the intermittent slowness and query failures affecting our Production EU Region. The issue was triggered by a large volume of dataset deletions that caused elevated database load… | 7h 13m | Resolved |
| Jun 17, 2026 | Querying Issues in EUThe issue is now resolved and querying is back to fully functional. | 31 min | Resolved |
| Jun 11, 2026 | Elevated API errors in production-eu1Requests to Honeycomb's Query Data and Management APIs in the production-eu1 environment encountered increased error rates between approximately 15:35 and 22:24 UTC. Event ingestion was not affected. The issue has been… | -467 min | Resolved |
| Jun 10, 2026 | Delayed ingest in EUAll services are now healthy. | 1h 3m | Resolved |
| Jun 8, 2026 | Canvas UI issuesThis incident has been resolved. | 1h 3m | Resolved |
| Jun 3, 2026 | Canvas tools unavailableThis incident has been resolved. | 0 min | Resolved |
| May 21, 2026 | Investigating querying issuesThis is an administrative update, marking the incident as Resolved. There has been no further impact since 23:20UTC on May 21 2026. | 17h 20m | Resolved |
| May 8, 2026 | activity log for US instance is delayedThis incident has been resolved. | 80h 37m | Resolved |
| May 7, 2026 | Scheduled Maintenance - US RegionScheduled maintenance is complete. All services are healthy and stable. Ingest, querying SLO processing, anomaly detection, Service Maps, and Activity Log are all confirmed operating normally. Backfill and SLO repair… | 3h 0m | Scheduled |
| May 4, 2026 | Scheduled Maintenance – EU RegionScheduled maintenance is complete. All services are healthy and stable. Ingest, Querying, SLO processing, Anomaly Detection, Service Maps, Triggers and Activity Log are all confirmed operating normally. Backfill and SLA… | 3h 0m | Scheduled |
| May 1, 2026 | Gap in Activity Log data in the EU region.The database failover that caused query issues earlier in the EU (https://status.honeycomb.io/incidents/n855d8kzp32y) has also had knock-on effects on our Activity Log beta feature. Due to recovery issues on the data… | -411 min | Resolved |
| May 1, 2026 | Query interuption due to database failover in the EU regionAt 12:50, our main database underwent an automated failover. This failover led to issues with some of our query engine connection pools, and caused queries to fail for roughly 10 minutes before self resolving. | -52 min | Resolved |
| Apr 30, 2026 | Links in Slack triggers not workingWe are confident the issue is fully resolved. | 1h 33m | Resolved |
| Apr 20, 2026 | Activity Log events interrupted for EUActivity Log is back to operational, and most events should be recovered. | 6 min | Resolved |
| Apr 15, 2026 | Activity Log ingest is delayedActivity Log is ingesting events again. We had no loss in data. | 20 min | Resolved |
| Apr 14, 2026 | Activity Log delayed ingest in EUAll activity log data is now live and up to date. | 1h 0m | Resolved |
| Apr 10, 2026 | Ingest issues for some customersWe can confirm that the fix was applied uniformly to all impacted and potentially impacted environments; no teams should see ingest errors related to this issue anymore. | 3h 42m | Resolved |
| Apr 9, 2026 | SLO processing delay in USWe have detected and remediated an issue around our SLO processing pipeline. From 17:00 to 17:15 UTC, events were being delayed, and alerts may likewise have been a few minutes late. Processing has caught up and is… | 1h 16m | Resolved |
| Apr 9, 2026 | Activity log delay in production USAll activity log data is now live and up to date. | 31h 17m | Resolved |
| Apr 9, 2026 | Partial Querying OutageWe experienced a partial querying outage from approximately 8:58 - 9:18 UTC. This outage may have resulted in some Trigger and SLO alert failures for the duration of the outage. The cause of the outage has been… | 0 min | Resolved |
| Mar 18, 2026 | Honeycomb UI displaying an errorThe issue has been identified and we have shipped a fix. | 11 min | Resolved |
| Mar 12, 2026 | Elevated Querying ErrorsThis incident has been resolved. | 1h 43m | Resolved |
| Mar 10, 2026 | Degraded Ingest due to database connectivity issuesThis incident is now resolved. We are also noting that some other parts of honeycomb (the query UI) may have seen spurious failures as well during the impacted period (13:22—13:44 UTC). | 24 min | Resolved |
| Feb 18, 2026 | Partial query outage in EU1Between 22:10 and 22:40 UTC on Feb18, 2026 Honeycomb experienced a partial query outage in our EU environment. The incident has been resolved, and we are continuing to investigate the contributing factors. | 0 min | Resolved |
| Feb 10, 2026 | Query ErrorsThis incident has been resolved. | 11 min | Resolved |
| Jan 22, 2026 | Transient querying and trigger failuresThis appears to have been a minor transient network issue; we do not expect this to recur | 15 min | Resolved |
| Jan 15, 2026 | Increased Latency and Errors for Events IngestWe observed a period of increased latency accompanied by a higher-than-normal error rate in some event ingest requests in the US. The service has since stabilized and is operating normally. We are continuing to monitor… | 0 min | Resolved |
| Jan 6, 2026 | Degraded querying and delayed records sent to Activity LogThis incident has been resolved. | 21h 40m | Resolved |
| Dec 16, 2025 | EU ONLY - Production EU Kafka migrationThe scheduled maintenance has been completed. | 2h 0m | Scheduled |
| Dec 5, 2025 | Querying and Ingest issues in EUThis incident started on December 5th, and is one of the longest in Honeycomb history, having been actively worked on and closed only on December 17th. Due to its impact and duration, we wanted to offer a partial and… | 282h 58m | Resolved |
| Dec 5, 2025 | System MaintenanceThe scheduled maintenance has been completed. | 15 min | Scheduled |
| Nov 25, 2025 | Activity log outage in production USAs of 23:00 UTC, all systems are back to synchronized. | 25h 22m | Resolved |
| Oct 20, 2025 | Delays in SLO, Service Maps processingThis incident has been resolved. At no point did we lose any customer data that hit our load balancers. Querying has been stable for hours, and all features that were degraded are functional. | 25h 24m | Resolved |
| Oct 20, 2025 | Investigating Querying IssuesAt 07:05 UTC on October 22, 2025, Honeycomb Platform On-Call received a page indicating that our [End to End tests](https://www.honeycomb.io/blog/how-honeycomb-uses-honeycomb-part-3-end-to-end-failures) that run every 5… | 2h 40m | Resolved |
| Oct 8, 2025 | Partially Degraded Honeycomb Intelligence FeaturesThe system has remained stable, we do believe everything to now be in order. | 57 min | Resolved |
| Sep 25, 2025 | ServiceMaps down due to issues with an upstream providerService Maps functionality has been restored. | 1h 15m | Resolved |
| Aug 28, 2025 | PagerDuty outage prevented Trigger and SLO notificationsA PagerDuty outage prevented some Trigger and SLO notifications from reaching your workspaces. The issue has since been resolved, and notification systems are now operating normally. Impact Window: 21:54 PT – 23:27 PT | -730 min | Resolved |
| Aug 27, 2025 | Partial degradation of data liveness for a subset of US usersWe have determined the cause and implemented a fix. Customers should be seeing normal operations again. | 17 min | Resolved |
How to Read This Honeycomb Incident Log
Each row is an incident PulsAPI observed, not a summary written afterwards. The duration is wall-clock time between the first failing check and the first clean one, so it includes the window before Honeycomb acknowledged anything. Vendor post-mortems typically measure from acknowledgement, which is why their numbers are usually shorter.
Incidents still open have no duration yet and are listed as ongoing rather than being given a running total. The archive covers the last 365 days; anything older has aged out of the window rather than never having happened.
The longest single Honeycomb outage in this window ran 11d 18h, against a mean recovery of 13h 14m. If you depend on Honeycomb in a customer-facing path, the longest figure is the one to design around. The mean is what happens on a normal bad day; the maximum is what happens on the worst one.