
Camille Roux
My first page is arithmetic, because most disagreements about availability are disagreements about the formula rather than about the incidents. There are several definitions and they give different answers from the same data. A ratio of successful requests to total requests weights by traffic. A ratio of successful measurement intervals to total intervals weights by time. An availability target expressed as a maximum downtime over a period is a third thing again, and it excludes planned maintenance where the other two do not. A long tail of low-traffic endpoints behaves badly under one definition and fine under another. An error budget is the same number expressed as a quantity of tolerated failure, and its value is that it connects reliability to a decision. Spending it requires deciding to, and having spent it means stopping feature work, which only means something if the trigger is agreed in advance rather than negotiated during an incident. An incident timeline is written during the incident. The postmortem is written after it. The timeline records what was observed and when, including the times the first page went unread. That last detail is worth recording because it changes the reading of everything after it. I cover alerts on rate of change rather than threshold, since a threshold tuned on quiet traffic fires late during the busiest hour. Every page here states the formula next to the figure, so two numbers from two teams can be compared rather than merely reported. I would rather a report state one number with its definition than three numbers with none.
About ToolSura
ToolSura offers 80+ free, privacy-first online tools that run 100% in your browser — no uploads, no logins. Learn more about our mission →