The concept comes from Google, who describe SRE as ”what happens when you ask a software engineer to design an operations function”. The core idea is the error budget: while availability is better than target, the team can take risk and ship fast; when the budget is spent, stability takes priority.
A good on-call setup is more than someone carrying a pager: clear, actionable alerts, runbooks for recurring problems and a mandate to improve the system after incidents.