On-call rotations ensure 24/7 coverage for production incidents, but poorly designed on-call programs burn out engineers and degrade the quality of both incident response and day-to-day engineering work. Effective on-call design balances coverage with sustainability: low alert noise, clear escalation paths, compensated on-call time, and runbooks that enable any rotation member to handle common incidents — not just the person who wrote the system.
Analogy🏏Cricket
🏏 Think of it like cricket: A batting coach who only reviews a batter's technique after a tournament has already ended. The batter has played 10 matches with a flawed grip — every run scored with bad technique is harder to unlearn than if the coach had corrected it in the first net session. Shifting security left is bringing the coach into the net sessions, not the post-tournament review.