Log Management Best Practices Cheat Sheet
Patterns for structured logging, centralized log aggregation, retention, and querying across distributed systems.
Standard Log Levels
Common severity levels and when to use them.
- DEBUG- Detailed diagnostic info, disabled in production by default
- INFO- Routine operational events, e.g. request handled, job started
- WARN- Unexpected but recoverable condition
- ERROR- Operation failed, requires attention but process continues
- FATAL/CRITICAL- Unrecoverable error, process is about to exit
Structured JSON Logging
Emit logs as JSON so they're machine-parseable by aggregators.
{ "timestamp": "2026-07-08T10:22:31Z", "level": "error", "service": "checkout-api", "trace_id": "a1b2c3d4", "message": "payment gateway timeout", "http_status": 504, "user_id": "u_9182"}
Elasticsearch / Kibana Query (KQL)
Query syntax for filtering logs in Kibana Discover.
# Find errors from a specific service in the last 15mservice: "checkout-api" and level: "error"# Range query on a numeric fieldhttp_status >= 500 and http_status < 600# Wildcard search on message textmessage: *timeout*
logrotate Config
Rotate and compress local log files to control disk usage.
# /etc/logrotate.d/myapp/var/log/myapp/*.log { daily rotate 14 compress delaycompress missingok notifempty create 0640 appuser appgroup}
Centralized Logging Pipeline
Typical components in a log aggregation pipeline.
- Shipper (Filebeat/Fluent Bit)- Lightweight agent that tails log files and forwards them
- Aggregator (Logstash/Fluentd)- Parses, enriches, and transforms log events before storage
- Storage (Elasticsearch/Loki)- Indexed store optimized for full-text or label-based search
- Visualization (Kibana/Grafana)- Dashboards and ad-hoc search UI over stored logs
Fluent Bit Parser & Filter Pipeline
Tail container logs, enrich with Kubernetes metadata, and drop noise before shipping to storage.
[INPUT] Name tail Path /var/log/containers/*.log Parser docker Tag kube.* Mem_Buf_Limit 5MB Skip_Long_Lines On[FILTER] Name kubernetes Match kube.* Merge_Log On Keep_Log Off K8S-Logging.Parser On[FILTER] Name grep Match kube.* Exclude log healthcheck[OUTPUT] Name es Match * Host elasticsearch.logging.svc Port 9200 Logstash_Format On Retry_Limit 5
Loki LogQL Query Patterns
Label-first log querying for Grafana Loki, cheaper at scale than full-text indexing every line.
# All error log lines from checkout-api in the last 15m{service="checkout-api"} |= "error"# Parse JSON and filter on a nested field{service="checkout-api"} | json | http_status >= 500# Rate of error lines per second over 5m, suitable for alertingsum(rate({service="checkout-api"} |= "error" [5m])) by (service)# Extract and average a numeric field embedded in log lines{service="checkout-api"} | json | unwrap latency_ms | avg_over_time(5m)
Log Sampling Strategies at Scale
Ways to control ingest volume and cost without losing debuggability during incidents.
- Head sampling- Decide to keep or drop a log line at emit time, e.g. keep 1 in N DEBUG lines
- Tail sampling- Buffer a full request/trace's logs and decide to keep based on outcome; always keep errors and high latency
- Dynamic sampling- Automatically raise the sample rate during incidents or when an anomaly detector fires
- Level-based- Always keep WARN and above, sample INFO, drop DEBUG in production by default
- Cost lever- Sampling trades searchability of the long tail for ingestion/storage cost — never sample ERROR or FATAL
OpenTelemetry Collector Log Pipeline
Vendor-neutral collection, PII redaction, and routing of logs before they reach durable storage.
receivers: filelog: include: [/var/log/app/*.log] operators: - type: json_parserprocessors: redaction: allow_all_keys: true blocked_values: - '\d{3}-\d{2}-\d{4}' - '\b4[0-9]{12}(?:[0-9]{3})?\b' batch: timeout: 5sexporters: otlphttp: endpoint: https://logs.example.com:4318service: pipelines: logs: receivers: [filelog] processors: [redaction, batch] exporters: [otlphttp]
PII & Compliance Considerations
Legal and security constraints that shape what you're allowed to log and for how long.
- Never log raw- Passwords, auth tokens, full card numbers, SSNs — mask or hash before the log line is emitted
- Right to erasure- GDPR/CCPA may require deleting a user's log data on request; key indices so this is feasible
- Retention limits- Set index lifecycle policies per data class; hot logs 7-30d, compliance archives longer in cold storage
- Access control- Restrict who can query raw vs redacted views, and audit access to the logs themselves
- Redact at source- Prefer redacting in the app or shipper over relying on a downstream filter that can be misconfigured
Always propagate a trace/correlation ID through every log line for a request — it turns scattered logs into a single searchable thread across microservices.