Prometheus & Monitoring Interview Questions
Metrics-based monitoring & observability, the pull model, metric types (counter, gauge, histogram, summary), PromQL, labels & cardinality, exporters, Alertmanager, recording & alerting rules, Grafana, SLIs/SLOs/SLAs, error budgets and long-term storage.
61 questions
Popular Searches
1. What is Prometheus and what problems does it solve in monitoring?
easyFundamentals 6 mins2. What is the difference between monitoring and observability?
mediumConcepts 6 mins3. How does Prometheus's pull-based metrics model work?
mediumArchitecture 6 mins4. What is the difference between push and pull metric collection?
mediumArchitecture 6 mins5. What are the four core Prometheus metric types (counter, gauge, histogram, summary)?
mediumMetric Types 6 mins6. What is the difference between a counter and a gauge in Prometheus?
easyMetric Types 5 mins7. What is the difference between a histogram and a summary in Prometheus?
mediumMetric Types 7 mins8. What is PromQL and how do you write basic queries?
easyQuerying 6 mins9. What is the rate() function in PromQL and when do you use it?
mediumFunctions 6 mins10. What are labels in Prometheus and why are they powerful?
easyData Model 5 mins11. What is cardinality in Prometheus and why is high cardinality a problem?
mediumPerformance 6 mins12. What are exporters in Prometheus and how do node_exporter and others work?
mediumExporters 6 mins13. How does service discovery work in Prometheus?
mediumConfiguration 7 mins14. What is the Prometheus scrape configuration and how do targets work?
mediumConfiguration 7 mins15. How does Prometheus store time-series data (the TSDB)?
hardStorage 8 mins16. What is Alertmanager and how does alerting work in Prometheus?
mediumAlerting 7 mins17. How do you write alerting rules and recording rules in Prometheus?
mediumRules 7 mins18. What is the difference between an alerting rule and a recording rule?
easyRules 6 mins19. How does Prometheus federation work for scaling?
hardScaling 7 mins20. What is remote write and remote read in Prometheus?
mediumStorage 7 mins21. What is Grafana and how does it integrate with Prometheus?
easyVisualization 6 mins22. What is the difference between Prometheus and Grafana?
easyMonitoring Stack 6 mins23. What are the RED and USE methods for monitoring?
mediumMonitoring Methodology 7 mins24. What are SLIs, SLOs, and SLAs and how do they relate to monitoring?
mediumReliability 7 mins25. What is an error budget and how does it drive reliability decisions?
mediumSRE & SLOs 7 mins26. How do you monitor Kubernetes with Prometheus?
mediumKubernetes Monitoring 8 mins27. What is the Prometheus Operator and how does it simplify deployment?
mediumKubernetes Operators 7 mins28. How do you handle long-term storage for Prometheus metrics (Thanos, Cortex, Mimir)?
hardStorage & Scaling 8 mins29. How do you reduce noise and alert fatigue in a monitoring system?
mediumAlerting 7 mins30. What are common Prometheus monitoring best practices and pitfalls?
mediumBest Practices 8 mins31. How does Prometheus staleness handling work, and why can a series disappear from a query result?
hardTSDB & Query Semantics 8 mins32. Why can increase() return a fractional value, and how do rate() extrapolation and counter resets affect accuracy?
hardPromQL Semantics 9 mins33. A target is showing as DOWN in Prometheus — how do you debug it end to end?
mediumOperations & Debugging 9 mins34. How do you run Prometheus and Alertmanager in high availability without getting duplicate pages?
hardArchitecture & Scaling 9 mins35. Prometheus is using far more memory than expected and restarts are slow — how do you diagnose and fix it?
hardOperations & Performance 10 mins36. A PromQL query is timing out or returning 'query processing would load too many samples' — how do you make it cheap?
hardPromQL Performance 10 mins37. How do you alert on the absence of data, and why do you need a dead man's switch?
mediumAlerting Design 9 mins38. What are exemplars in Prometheus, and how does Prometheus interoperate with OpenTelemetry metrics?
hardEcosystem & Integration 10 mins39. How do you secure a Prometheus deployment — scrape credentials, the web UI, and the admin API?
hardSecurity 10 mins40. How do you unit test Prometheus alerting and recording rules before they reach production?
mediumTesting & CI 9 mins41. What is the difference between relabel_configs and metric_relabel_configs, and where do you drop unwanted series?
hardConfiguration & Scraping 9 mins42. Why is averaging percentiles wrong, and how do you aggregate latency correctly with histogram_quantile?
hardPromQL 9 mins43. An alert is firing in Prometheus but nobody was paged. How do you debug the Alertmanager path?
hardAlerting 10 mins44. When is the Pushgateway the right tool, and what operational traps does it introduce?
hardArchitecture & Scaling 9 mins45. What causes 'out of order sample' and 'duplicate sample for timestamp' errors, and how do honor_timestamps and out-of-order ingestion relate?
hardStorage & Ingestion 9 mins46. Prometheus will not start after a disk filled up. How do you recover the TSDB without losing more data than necessary?
hardStorage & Ingestion 10 mins47. An alert keeps firing and resolving every few minutes. How do you stabilise it without hiding a real problem?
hardAlerting 9 mins48. A single Prometheus can no longer scrape every target. How do you shard the scrape load?
hardArchitecture & Scaling 10 mins49. How does PromQL vector matching work, and how do you enrich a metric with labels from an info metric using group_left?
hardPromQL 9 mins50. How do you monitor something that cannot expose /metrics — an external endpoint, a certificate, or a shell script's output?
hardExporters & Integration 10 mins51. How do you compare current behaviour with last week in PromQL, and what do offset and the @ modifier actually do?
hardPromQL 9 mins52. Recording rules are producing gaps and alerts are late. How do you diagnose rule evaluation health and ordering?
hardAlerting 10 mins53. What are the Prometheus metric naming conventions, and why do base units matter?
mediumInstrumentation 8 mins54. What is the OpenMetrics exposition format, and how does Prometheus decide which format to parse?
mediumConfiguration & Scraping 9 mins55. How do you choose a scrape interval, and how must the rate() window relate to it?
mediumPromQL 8 mins56. What is the query step in a range query, and how does lookback delta change what you see?
hardPromQL 10 mins57. How do you choose histogram buckets, and what do native histograms change?
hardMetric Types 10 mins58. Which scrape limits protect a Prometheus server, and what happens when one is exceeded?
mediumOperations 9 mins59. Why does a rolling deployment spike Prometheus memory, and how do you manage series churn?
hardOperations 10 mins60. How do you decide what to precompute as a recording rule, and how should you name it?
mediumRules & Alerting 9 mins61. How do topk, bottomk and the _over_time functions differ, and when does topk mislead you?
mediumPromQL 9 mins