Notes on API and vendor monitoring
Uptime checks, incident response, SLA tracking, and the parts of vendor monitoring that actually matter when something breaks.
91 articles · page 9 of 10
Monitoring Microsoft 365 Status: Teams, Outlook, SharePoint, and Azure AD
A practical guide to monitoring Microsoft 365 reliability, including Teams, Outlook, SharePoint, OneDrive, and Azure Active Directory. Component-level alerts, outage patterns, and an IT-ops-friendly setup.
The Blameless Postmortem Template: A Complete Guide with Real Examples
A ready-to-use blameless postmortem template with section-by-section guidance, real published postmortem examples, and the practices that separate useful postmortems from paperwork.
SLA Credits Explained: How to Calculate, Claim, and Negotiate Service Credits
A complete guide to SLA service credits, how they are calculated, how to claim them (most teams leave money on the table), and how to negotiate better SLA terms at contract renewal.
Active vs Passive Monitoring: When to Use Each (and Why You Need Both)
Active monitoring sends synthetic traffic on a schedule; passive monitoring observes real traffic flowing through your stack. Learn the tradeoffs, where each one breaks down, and how to combine them for end-to-end coverage.
Black-Box vs White-Box Monitoring: An SRE Guide to Choosing the Right Lens
Black-box monitoring tests your system from the outside; white-box exposes its internals. The Google SRE distinction explained, with concrete examples, instrumentation patterns, and the failure modes each one misses.
Push vs Pull Monitoring: Which Architecture Fits Your Stack?
Pull-based monitoring (Prometheus, ICMP probes) scrapes targets on a schedule; push-based monitoring (StatsD, OpenTelemetry OTLP) has agents send data outward. Compare the architectures, scaling profiles, and where each one breaks.
Heartbeat vs Health Check Endpoints: Designing Signals That Actually Mean Something
A heartbeat says 'I'm alive'; a health check says 'I'm ready to serve traffic.' Most teams conflate them and end up with green dashboards over broken services. Here's how to design each one correctly.
APM vs Infrastructure Monitoring vs Status Page Monitoring: What Each One Actually Sees
APM watches your code, infrastructure monitoring watches your servers, and status page monitoring watches everyone else's services. They overlap less than most teams assume, and the gaps are where outages live.
Multi-Cloud Reliability: Monitor AWS, Azure & GCP Together
Running workloads across AWS, Azure, and GCP? Here's how to get a unified view of reliability across all three, without building a custom aggregation layer.