Site Reliability Engineering (SRE) is a discipline that applies software engineering principles to infrastructure and operations problems. Originated at Google in 2003 under Ben Treynor Sloss, SRE treats operations as a software problem: automate what is repetitive, measure what matters, and spend engineering time on work that compounds in value rather than work that must be repeated indefinitely. The foundational SRE insight is that reliability is a feature — one that must be engineered deliberately, not hoped for.
Analogy🏏Cricket
🏏 Think of it like cricket: A batting coach who only reviews a batter's technique after a tournament has already ended. The batter has played 10 matches with a flawed grip — every run scored with bad technique is harder to unlearn than if the coach had corrected it in the first net session. Shifting security left is bringing the coach into the net sessions, not the post-tournament review.