Nagios Cheat Sheet
Reference for Nagios Core object configuration, check_ plugin usage, and service/host definitions for infrastructure monitoring.
Host Definition
A host object in a Nagios .cfg configuration file.
define host { use linux-server host_name web1 alias Web Server 1 address 192.168.1.10 max_check_attempts 3 check_period 24x7 notification_interval 30 notification_period 24x7}
Service Definition
A service check monitoring HTTP and disk usage.
define service { use generic-service host_name web1 service_description HTTP check_command check_http max_check_attempts 3 check_interval 5 retry_interval 1}define service { use generic-service host_name web1 service_description Disk Usage check_command check_disk!20%!10%!/}
Plugin Usage
Running Nagios plugins directly from the command line.
/usr/lib/nagios/plugins/check_ping -H 8.8.8.8 -w 100.0,20% -c 500.0,60%/usr/lib/nagios/plugins/check_http -H example.com -p 443 -S/usr/lib/nagios/plugins/check_disk -w 20% -c 10% -p //usr/lib/nagios/plugins/check_nrpe -H web1 -c check_load
Key Concepts
Core Nagios terminology and status codes.
- host / service- Monitored objects; a host has one or more service checks attached
- check_command- Plugin invocation defining how a check is performed
- contact / contactgroup- Defines who is notified and via what method when checks fail
- exit codes- 0=OK, 1=WARNING, 2=CRITICAL, 3=UNKNOWN — the plugin contract Nagios relies on
- NRPE- Agent enabling Nagios to execute plugins on remote hosts
- soft/hard state- Soft states are unconfirmed failures during retries; hard state triggers notification
Custom Plugin Script
A minimal custom check plugin in Bash honoring the Nagios exit-code contract, ready to wire up as check_command.
#!/bin/bash# check_queue_depth.sh -H <host> -w <warn> -c <crit>WARN=100CRIT=500while getopts "H:w:c:" opt; do case $opt in H) HOST=$OPTARG ;; w) WARN=$OPTARG ;; c) CRIT=$OPTARG ;; esacdoneDEPTH=$(curl -s "http://${HOST}:15672/api/queues/%2f/jobs" | jq '.messages')if [ "$DEPTH" -ge "$CRIT" ]; then echo "CRITICAL - queue depth is ${DEPTH} | depth=${DEPTH};${WARN};${CRIT}" exit 2elif [ "$DEPTH" -ge "$WARN" ]; then echo "WARNING - queue depth is ${DEPTH} | depth=${DEPTH};${WARN};${CRIT}" exit 1else echo "OK - queue depth is ${DEPTH} | depth=${DEPTH};${WARN};${CRIT}" exit 0fi
Escalations & Host Dependencies
Object definitions that stop notification storms: escalation chains and dependency suppression.
define serviceescalation { host_name web1 service_description HTTP first_notification 3 last_notification 5 notification_interval 60 contact_groups oncall-secondary}# Suppress alerts for services on a host behind a failed routerdefine hostdependency { host_name web1,web2,web3 dependent_host_name web1,web2,web3 execution_failure_criteria n notification_failure_criteria d,u dependent_host_name web1}define servicedependency { host_name core-router service_description PING dependent_host_name web1 dependent_service_description HTTP notification_failure_criteria w,c,u}
Passive Checks via NSCA
Submitting passive check results from a remote host into Nagios using send_nsca, for checks that push rather than get polled.
# On the monitored host: submit a passive resultprintf "web1\tDisk Usage\t0\tOK - 42%% used\n" | \ /usr/sbin/send_nsca -H nagios.internal -c /etc/nagios/send_nsca.cfg# Corresponding service definition on the Nagios serverdefine service { use generic-service host_name web1 service_description Disk Usage check_command check_dummy!0 active_checks_enabled 0 passive_checks_enabled 1 check_freshness 1 freshness_threshold 3600}
Scaling & Distributed Monitoring
Concepts for running Nagios beyond a single small server.
- check_multi / MRPE- Bundles several sub-checks into one plugin invocation to reduce process-spawn overhead at scale
- distributed monitoring (NSCA)- Satellite Nagios/NSClient instances push passive results to a central server instead of it polling every host
- obsess_over_services- Hook that forwards every check result to an external command, used to feed data into a central event pipeline
- flap detection- Detects a service oscillating between states and suppresses repeat notifications until it stabilizes
- check_interval vs retry_interval- Normal polling cadence vs. the faster cadence used while in an unconfirmed (soft) state
- event handler- Script triggered automatically on a state change, commonly used for auto-remediation (e.g. restart a service)
Auto-Remediation Event Handler
An event handler that restarts a crashed service automatically on the first hard CRITICAL state.
define service { use generic-service host_name web1 service_description nginx check_command check_procs!1:!nginx event_handler restart-nginx event_handler_enabled 1}define command { command_name restart-nginx command_line $USER1$/eventhandlers/restart_nginx.sh $SERVICESTATE$ $SERVICESTATETYPE$ $SERVICEATTEMPT$}# restart_nginx.sh — only act on the confirmed (HARD) critical stateif [ "$2" = "HARD" ] && [ "$1" = "CRITICAL" ]; then /usr/bin/systemctl restart nginxfi
Write custom checks to always exit with the standard Nagios codes (0/1/2/3) and print a one-line status message to stdout — any script honoring this contract can be dropped in as a check_command with zero Nagios-side changes.