100% Free Forever
AI-Powered Learning
Industry Expert Content
Certificates & Badges
Learn At Your Own Pace
Architecture

Monitoring and Observability

The practices and tooling — metrics, logs, and traces — that let engineers understand a distributed system's internal state from its external outputs and detect or diagnose problems.

Reliability & ResilienceIntermediate10 min readJul 9, 2026
Analogies

Monitoring and Observability

Monitoring is the practice of collecting predefined signals about a system's health and alerting when they cross known thresholds — it answers questions you already thought to ask, like 'is CPU usage above 90%' or 'is the error rate above 1%.' Observability is a broader property of the system itself: how well its internal state can be inferred purely from the external outputs it produces (metrics, logs, and traces), enabling engineers to answer questions they did not anticipate in advance, such as diagnosing a novel failure mode at 3am by exploring data rather than checking a pre-built dashboard. A system can be heavily monitored yet poorly observable if its telemetry only covers known failure modes and offers no way to investigate the unknown ones.

🏏

Cricket analogy: Monitoring is checking whether the required run rate crossed a known danger threshold like 12 per over; observability is having enough ball-by-ball data on line, length, and shot type to diagnose an unusual collapse nobody predicted.

The Three Pillars: Metrics, Logs, and Traces

Metrics are numeric time-series measurements — request rate, error count, latency percentiles, queue depth — cheap to store and query at scale, ideal for dashboards and threshold-based alerting, but they aggregate away individual request detail. Logs are discrete, timestamped event records, often free-text or structured (JSON), that capture rich per-event detail including error messages and stack traces, but are comparatively expensive to store and search at high volume. Distributed traces follow a single request as it flows across multiple services, recording the timing of each hop (a span) and stitching them into a single trace via a shared trace ID, which is the only one of the three pillars that directly answers 'where exactly, across all these services, did this specific slow or failing request spend its time.'

🏏

Cricket analogy: Metrics are the scoreboard's running totals of runs, wickets, and overs; logs are the ball-by-ball commentary text capturing rich detail on each delivery; a trace follows one controversial run-out through fielder, keeper, and third umpire to see exactly where the delay happened.

The RED and USE Methods

Two widely used frameworks help decide what to actually measure. The RED method, aimed at request-driven services, tracks Rate (requests per second), Errors (failed requests per second), and Duration (latency distribution, typically p50/p95/p99) — three signals that together characterize the health of any service handling requests. The USE method, aimed at resources like CPU, disk, or a connection pool, tracks Utilization (percentage of time the resource is busy), Saturation (amount of queued work waiting for the resource), and Errors (error events on the resource) — useful for diagnosing whether a resource itself is the bottleneck behind a RED-method symptom.

🏏

Cricket analogy: The RED method for a bowler is Rate (deliveries per over), Errors (wides and no-balls), Duration (time between deliveries); the USE method for the pitch is Utilization (how much it's bowled on), Saturation (footmarks queuing up wear), and Errors (cracks forming).

text
A distributed trace across three services (single trace ID: 7f3a2c)

span: api-gateway         [====================================] 240ms
  span: auth-service       [==]                                    12ms
  span: order-service                [========================]  150ms
    span: inventory-db-query            [========]                 60ms
    span: payment-service                         [==========]     70ms

Reading the trace: 150ms of the 240ms total was spent inside
order-service, and within that, payment-service's 70ms call is the
single largest contributor  the trace pinpoints exactly which hop
to investigate, which aggregate latency metrics alone cannot show.

Google's SRE practice popularized Service Level Indicators (SLIs, e.g. 'proportion of requests served in under 300ms'), Service Level Objectives (SLOs, e.g. '99.9% of requests over 30 days'), and error budgets (the allowed amount of unreliability implied by the SLO) as a way to turn raw metrics into a shared, quantitative language between engineering and the business for deciding when to prioritize reliability work over new features.

A common mistake is alerting on symptoms that do not map to genuine user impact, such as paging on-call for a single CPU spike on one of many redundant instances. Over time this produces alert fatigue, where engineers start ignoring or muting pages, which is dangerous because it desensitizes the team right before an alert that does matter. Alerts should generally be tied to SLO burn rate or direct user-facing symptoms (elevated error rate, elevated latency) rather than every internal metric anomaly.

Cardinality and Cost

A frequent operational trap in observability systems is uncontrolled cardinality: attaching a high-cardinality label (like a raw user ID or request ID) to a metric multiplies the number of distinct time series the monitoring system must store and index, sometimes by orders of magnitude, causing storage costs and query latency to explode or the metrics backend to fall over entirely. High-cardinality data (like exact request IDs) belongs in logs or traces, which are designed to index and search per-event detail, not in metrics, which are designed for cheap aggregation across a bounded label space.

🏏

Cricket analogy: Tagging every ball-by-ball metric with the exact seat number of a spectator who caught a six would explode the scoreboard system's storage with useless dimensions, that detail belongs in a highlight log, not the aggregate run-rate metric.

  • Monitoring answers pre-defined questions via thresholds and alerts; observability enables answering unanticipated questions by exploring telemetry.
  • Metrics, logs, and traces are the three complementary pillars — cheap aggregates, rich per-event detail, and cross-service request timing, respectively.
  • The RED method (Rate, Errors, Duration) characterizes request-driven service health; the USE method (Utilization, Saturation, Errors) characterizes resource health.
  • Distributed traces use a shared trace ID across spans to show exactly where time was spent across multiple services for a single request.
  • SLIs, SLOs, and error budgets translate raw telemetry into a shared reliability language for prioritization decisions.
  • Attaching high-cardinality labels to metrics (e.g. raw user IDs) can explode storage/query cost; such detail belongs in logs or traces instead.

Practice what you learned

Was this page helpful?

Topics covered

#Architecture#SystemDesignStudyNotes#SoftwareEngineering#MonitoringAndObservability#Monitoring#Observability#Three#Pillars#StudyNotes#SkillVeris

Frequently Asked Questions

21 categories · pick one to explore

Where can I get free study notes for programming and tech subjects?
SkillVeris offers completely free study notes covering programming and tech subjects, with no signup fees or paywalls. The notes are structured by course and topic, written for quick understanding, and enriched with the Learn Through Hobbies analogy method, so you can revise concepts through cricket, music, gaming, cooking and more.
Are SkillVeris study notes good for exam revision?
Yes, the study notes are designed for efficient revision: each topic answers its heading immediately, keeps explanations concise, and links to related glossary terms and cheat sheets. Students preparing for university exams or certification tests use them as quick revision notes because they distil concepts without the padding of full textbooks.
What subjects do the free study notes cover?
The study notes span the platform's main domains, including AI and machine learning, Python and programming, web development, DevOps, cloud, security and databases. Coverage mirrors the 37 live courses, so notes exist for the topics you are actually studying, and new note sets are added as courses launch.
How are SkillVeris study notes different from regular textbooks?
The notes are answer-first, concise and free, whereas textbooks are long and often expensive. Each section explains one concept directly, then reinforces it through selectable hobby analogies like cricket or cooking. Notes also cross-link to the glossary, blog and cheat sheets, letting you jump to related material instantly instead of flipping pages.
Can I use the developer study material without creating an account?
The study notes are free to access, and SkillVeris does not charge anything for its developer study material at any point. Browsing notes is straightforward from the Study Notes section, and if you want progress tracking, certificates and AI Mentor conversations tied to your learning, a free account unlocks those extras.
Do the study notes explain concepts with analogies?
Yes, this is a signature SkillVeris feature. Study notes use the Learn Through Hobbies method, explaining technical concepts through analogies from twelve domains including cricket, music, gaming, photography, travel, movies, fitness, chess, cooking, finance, business and sports. You can switch the analogy domain instantly to whichever hobby makes the concept click.
Are the revision notes suitable for last-minute exam preparation?
Yes, revision notes on SkillVeris work well for last-minute preparation because every section states the answer in its first sentences, so skimming is genuinely effective. Pair them with the relevant cheat sheet for formulas and syntax, and use the glossary for any unfamiliar term you meet while cramming.
Is there free study material for AI and machine learning?
Yes, SkillVeris provides free study notes across its AI and ML catalogue, covering Python for AI, deep learning frameworks like PyTorch and TensorFlow, Hugging Face Transformers, Large Language Models, RAG, AI agents and MLOps. All of it is free, making it a strong resource for Indian students and global learners alike.
Can beginners understand the study notes, or are they for experts?
Beginners can absolutely use them. The notes are written in plain language, define terms as they appear, and lean on hobby analogies to make abstract ideas concrete. Difficulty scales with the underlying course level, so beginner-course notes stay gentle while advanced-course notes go deeper, and the glossary supports you throughout.
How do study notes connect with SkillVeris courses?
Study notes are organised by course and topic, so they map directly to the structured courses and their 24–40-lesson curriculum. Many learners study a lesson first, then use the matching notes for revision before module assessments and the final exam, where 80 percent is required to pass and earn the certificate.
Are there study notes for Python specifically?
Yes, Python is well covered through notes tied to the Python-focused courses, including Python for AI and ML. Topics span fundamentals through applied machine learning usage. You can reinforce the notes with Python practice in Code Lab, which runs code in your browser with no installation required.
Do the study notes include code examples?
Yes, study notes include code examples wherever a concept is best shown in code, alongside explanations, key points and analogies. Reading a snippet in the notes and then reproducing it yourself in Code Lab is an effective loop, since Code Lab lets you run code in the browser across six languages.
How often is new study material added to SkillVeris?
Study material grows alongside the course catalogue. Whenever new courses join the platform's 37 live courses, matching study notes, glossary entries and cheat sheets are added so the resources stay in sync. Existing notes are also refined over time, so it is worth revisiting topics you studied earlier.
Can I use SkillVeris notes to prepare for technical interviews?
Yes, the notes make excellent interview revision because they compress each concept into direct, answer-first explanations, which mirrors how you should answer interview questions. Combine them with the SkillVeris interview questions feature, which includes readiness scoring, to test whether your revision has actually made you interview-ready.
Are the study notes mobile-friendly for studying on the go?
Yes, the study notes are built to load fast and read comfortably on mobile devices, so you can revise during a commute or between classes. Sections are short and answer-first, which suits small screens, and analogy switching works on mobile too, letting you study anywhere without carrying books.
What is the difference between study notes and cheat sheets?
Study notes explain concepts in depth with context, examples and analogies, making them ideal for learning and revision. Cheat sheets are compact quick-reference summaries of syntax, commands and key facts, ideal once you already understand a topic. Most learners study the notes first, then keep the cheat sheet handy while coding.
Do study notes help if I am stuck on a course lesson?
Yes, reading the matching study notes often clarifies a lesson because the same concept is explained from a different angle, frequently with a different analogy. If you are still stuck, ask the AI Mentor, which answers 24/7 at Quick, Detailed or Deep-dive depth until the idea genuinely makes sense.
Is there free study material for DevOps and cloud topics?
Yes, SkillVeris carries free study notes for DevOps and cloud topics as part of its coverage across 37 live courses. The material suits learners following the DevOps Engineer or Cloud Engineer paths, and it links to related glossary terms and cheat sheets so you can revise the whole toolchain in one place.
Can school or college students in India use these notes for projects?
Yes, students across India and worldwide use SkillVeris notes for coursework, projects and exam preparation, and everything is free, which matters for student budgets. The notes explain concepts clearly enough to cite in project reports, and Code Lab lets you prototype the project code directly in your browser.
How should I combine study notes with other SkillVeris resources?
A proven loop: learn from a course lesson, revise with the matching study notes, look up unfamiliar terms in the glossary, keep the cheat sheet open while practising in Code Lab, and quiz yourself with interview questions. The AI Mentor fills any remaining gaps 24/7, at whatever depth you need.

What Learners Say

Real journeys from the SkillVeris community — swipe for more.

SkillVeris taught me Python through Cricket. Now I’m building real projects and feeling confident!
Arjun S. · B.Tech Student
The best platform for hobby-based learning. Concepts finally stick.
Priya R. · Data Analyst
I went from zero coding to a portfolio of projects — all by learning through my love for gaming. Landed my first internship!
Kabir M. · CS Undergraduate
Trending Topics50 popular tags — tap to explore
Trending CoursesAll 37 free courses — tap to browse