100% Free Forever
AI-Powered Learning
Industry Expert Content
Certificates & Badges
Learn At Your Own Pace
Observability & Monitoring
35 minintermediate

Alerting and On-Call Routing

An alert fires when a metric crosses a threshold that signals user impact. Alertmanager — Prometheus's companion daemon — receives alert notifications from Prometheus and handles deduplication, grouping, inhibition, and routing to the right receiver: PagerDuty for critical outages, Slack for warnings, email for audit trails. Well-designed alerting wakes up the right person with the right context at the right time — nothing more.

Analogy🏏Cricket
🏏 Think of it like cricket: During India's 2023 World Cup final against Australia, the team management tracked three distinct data streams simultaneously to understand match performance. The scoreboard showed run rate, current score, and required run rate — aggregated numbers updated every over, equivalent to metrics. The commentary and match notes recorded each delivery's outcome — Rohit Sharma edged a yorker from Hazlewood at the 12th over, first ball — equivalent to logs. The ball-tracking DRS system traced the exact path of each delivery from Bumrah's hand through the air to the stumps, showing the full journey of that dismissal — equivalent to traces. Just as the scoreboard alone cannot explain why the run rate collapsed (you need the logs to see specific wicket events and traces to follow the pressure chain from bowler to batter to fielder), metrics alone cannot explain why your API latency spiked — you need logs for individual request events and traces to follow the request across services. The insight is that each pillar answers a different question: metrics give magnitude, logs give events, and traces give causality — and you need all three to diagnose a complex failure.
Lesson 21 of 24
0% complete