Incidents articles
Uptime checks, incident response, SLA tracking, and the parts of vendor monitoring that actually matter when something breaks.
8 articles in Incidents
Severity Levels for Third-Party Outages: A Scale That Holds Up
Standard SEV scales assume you can fix the thing that broke. When the broken thing is a vendor you cannot fix, severity has to key on user impact instead.
API Outage Communication Template: What to Say Before, During, and After Incidents
Use this API outage communication template to write faster customer updates, reduce confusion, and keep stakeholders aligned during third-party service incidents.
Incident Postmortem for a Third-Party Outage: Questions That Lead to Better Resilience
Run a better third-party outage postmortem with questions that focus on detection, dependency mapping, communication, fallback behavior, and resilience investment.
Incident Response Runbooks: A Template for Zero-Panic Outages
When the alert fires at 2 AM, you don't want to think, you want to follow a script. We've compiled battle-tested runbook templates from 50+ engineering teams, distilled into a single framework you can deploy today.
Root Cause Analysis Done Right: The 5 Whys & Beyond
The 5 Whys is a useful tool, but only if you avoid common traps: stopping too early, assigning blame, and ignoring systemic factors. Walk through three real post-mortems and learn what a blameless RCA looks like in practice.
On-Call Rotation Design: Sustainable Reliability Without Burnout
Pager burnout is a real crisis in reliability engineering. This guide covers on-call rotation design, escalation policies, alert fatigue reduction, and how auto-grouping cuts mean time-to-acknowledge by 42%.
How to Stand Up a War Room in 60 Seconds During a Vendor Outage
When a vendor goes down, the first 60 seconds determine how fast you recover. Here's the exact war room setup sequence that eliminates the 'who does what' confusion at the worst possible moment.
Postmortem Template for Third-Party Outages: When It Wasn't Your Fault
Most postmortem templates assume you caused the outage. Here's a template built specifically for third-party incidents, covering attribution, blast radius, resilience gaps, and vendor accountability.