100% Free Forever
AI-Powered Learning
Industry Expert Content
Certificates & Badges
Learn At Your Own Pace
Observability & Monitoring
30 minintermediate

The Three Pillars — Metrics, Logs, Traces

Modern distributed systems can fail in thousands of ways — a memory leak in one microservice, a slow database query in another, a network partition between two pods. Without structured visibility into what is happening inside your system, debugging these failures becomes a process of educated guesswork. The three pillars of observability — metrics, logs, and traces — were developed precisely to eliminate that guesswork. Metrics tell you that something is wrong. Logs tell you what happened. Traces tell you where the problem propagated across service boundaries. Together, they form a complete picture of system health, allowing engineers to move from 'the site is slow' to 'this specific database call on checkout-service is timing out under load' in minutes rather than hours.

Analogy🏏Cricket
🏏 Think of it like cricket: During India's 2023 World Cup final against Australia, the team management tracked three distinct data streams simultaneously to understand match performance. The scoreboard showed run rate, current score, and required run rate — aggregated numbers updated every over, equivalent to metrics. The commentary and match notes recorded each delivery's outcome — Rohit Sharma edged a yorker from Hazlewood at the 12th over, first ball — equivalent to logs. The ball-tracking DRS system traced the exact path of each delivery from Bumrah's hand through the air to the stumps, showing the full journey of that dismissal — equivalent to traces. Just as the scoreboard alone cannot explain why the run rate collapsed (you need the logs to see specific wicket events and traces to follow the pressure chain from bowler to batter to fielder), metrics alone cannot explain why your API latency spiked — you need logs for individual request events and traces to follow the request across services. The insight is that each pillar answers a different question: metrics give magnitude, logs give events, and traces give causality — and you need all three to diagnose a complex failure.
Lesson 1 of 24
0% complete