The clock starts when users are affected and stops when the service works again — not when the root cause is fixed. A low MTTR is therefore less about never failing and more about detecting fast, understanding fast and rolling back fast: good observability, alerts you can act on and deploys that can be rolled back. In its recent reports DORA calls the metric failed deployment recovery time and focuses on failures caused by a change; see DORA’s guide to the metrics.
Together with the error budget, MTTR makes reliability measurable, and a blameless post-mortem after every incident is what brings the number down over time. Follow the trend rather than single incidents — one long night says little about the system.