A year of carrying the pager taught us less about outages than about the decisions that quietly made them likelier. Six changes, in the order they paid off.
Alert on symptoms, not causes

Cause-based alerts multiply as the system grows and go stale the moment it changes. Symptom alerts stay small: is the thing slow, is it wrong, is it down.
Write the runbook before the incident

A runbook written during an outage is a transcript, not a runbook. Draft it when the change ships, while the reasoning is still in someone’s head.
Make the rollback boring

Teams hesitate to roll back when rolling back is exciting. Practise it on an ordinary Tuesday until it is the least interesting option available.
Hand over out loud

Most repeat incidents are handover failures. Five minutes of narration between shifts beats a paragraph in a channel nobody reads.
Let the pager set the roadmap

If the same alert fires three weeks running, that is not noise to be tuned out. It is a feature request with a timestamp.
Count the quiet weeks

Reliability work is invisible when it succeeds. Track the weeks nothing happened, or nobody will fund the next round of it.
Conclusion
None of these were technology changes. All of them reduced the number of nights somebody lost.