Notes on API and vendor monitoring
Uptime checks, incident response, SLA tracking, and the parts of vendor monitoring that actually matter when something breaks.
91 articles · page 8 of 10
PulsAPI vs. Datadog: Why Engineering Teams Use Both
Datadog monitors your infrastructure. PulsAPI monitors what your infrastructure depends on. These tools complement each other, here's how to think about both and why using one doesn't replace the other.
How to Monitor Your Payment Stack: Stripe, Braintree, PayPal, and Adyen
Payment processor outages directly cost you revenue. Here's how to set up real-time monitoring for your entire payment stack, with component-level alerts, fallback strategies, and SLA tracking for each provider.
AWS Outage History: Every Major Incident from 2011 to 2026
Every major AWS outage from 2011 to 2026 in one sortable table: date, service, region, duration, trigger. Plus which region fails most and how often AWS goes down.
Cloudflare Outages Explained: A Complete History and How to Stay Resilient
A complete breakdown of major Cloudflare outages, including the 2019 regex incident, the 2022 BGP routing failure, and the 2024 control-plane outage, with practical strategies for surviving the next one.
SaaS Uptime Statistics 2026: Real Uptime Data Across 100+ Cloud Vendors
Actual uptime percentages, incident frequency, and SLA performance across 100+ SaaS vendors in 2026, measured by PulsAPI. See which vendors beat their SLA, which miss it, and how the market has changed year-over-year.
15 Best Status Page Examples in 2026 (And What Makes Them Actually Work)
A close look at the 15 best status pages of 2026, from GitHub and Stripe to smaller teams doing it right. What each page does well, where they fall short, and how to apply the lessons to your own status page.
MTTR, MTBF, MTTA, and MTTD Explained: The Complete Reliability Metrics Guide
A plain-English guide to the four reliability metrics every engineering team needs: mean time to repair, mean time between failures, mean time to acknowledge, and mean time to detect, with formulas, examples, and benchmarks.
How to Calculate Uptime Percentage: Formulas, Examples, and Common Pitfalls
A step-by-step guide to calculating uptime percentage correctly, with worked examples, the difference between raw and SLA-adjusted uptime, and the pitfalls that quietly inflate the numbers vendors publish.
Monitoring OpenAI API Outages: How to Build Resilient AI-Powered Features
OpenAI API has averaged 99.6% uptime, meaningfully lower than the SaaS average. Here's how to monitor OpenAI reliably, build fallbacks to Anthropic and Azure OpenAI, and design AI features that degrade gracefully.