Paperphyte/Glossary/Site Reliability Engineering (SRE)

What is Site Reliability Engineering (SRE)?

SRE is a way of making operations systematic and measurable: reliability is expressed as objectives (SLOs) and error budgets, and the team actively works with alerting, automation and improvement instead of reactive firefighting.

The concept comes from Google, who describe SRE as ”what happens when you ask a software engineer to design an operations function”. The core idea is the error budget: while availability is better than target, the team can take risk and ship fast; when the budget is spent, stability takes priority.

A good on-call setup is more than someone carrying a pager: clear, actionable alerts, runbooks for recurring problems and a mandate to improve the system after incidents.

Say hello

Got a problem worth solving?

Chat with us on WhatsApp. We reply within 48 hours.