SLA, Outages & Reliability
What a month of incidents actually delivered, what the contract pays for it, and how little a handful of failures tells you about how often the next one is due.
The incident log
Against the contract
How often, and how long
Where the downtime came from
Capacity trend
Why the convention matters more than the log
Two parties can read the same incident log and quote availabilities several nines apart, without either of them being dishonest. The differences come from choices that are rarely written down clearly: whether planned maintenance counts, whether a partial outage counts in full or in proportion, whether the clock starts at the alarm or at the ticket, and what the measurement period is.
Most contracts exclude planned work, which is reasonable — it was notified and scheduled. Customers experienced it anyway, so the number they measured will be worse than the number in the report, and neither is wrong. Showing both readings is more useful than defending one.
A credit is not compensation
Service credits are usually expressed as a percentage of a monthly fee, and they are capped. A full day of outage might earn 25% of a month — a few hundred on a typical circuit — against a day of lost trading, idle staff or missed deliveries that cost orders of magnitude more.
That is not a flaw in the contract; it is what the contract is. A credit is a remedy sized to the service, not to your business. Putting the two figures side by side makes the decision about resilience an honest one: the second circuit either pays for itself against the real cost of downtime, or it does not, and the credit barely enters the calculation.
Small samples say very little
Three failures in a year gives an MTBF of 2,920 hours. That number looks like a measurement, and it is not. For failures arriving at a roughly constant rate, the 90% confidence interval around three observations spans more than a factor of five — the truth might be well under 1,500 hours or well over 7,000.
This matters because MTBF figures get compared. A vendor quoting 100,000 hours from a large population and a site quoting 2,920 from three incidents are not making the same kind of statement, and the second cannot be used to contradict or confirm the first. The interval is shown here for that reason.
Repeated short outages make it worse. Three flaps on one link within an afternoon are usually one underlying fault, and logging them as three failures inflates the failure rate while flattering the average repair time — moving both numbers in the direction that looks better.
What this does not do
It does not diagnose anything. Clustering groups incidents that are close in time on the same element; whether they share a cause is a judgement for someone with the logs. Nothing here is a compliance determination, and the contract's own definitions govern what counts as downtime.
Trend extrapolation is a description of past points, not a forecast. Traffic growth is rarely linear, and one new customer moves a link more than a year of trend does.
Everything is calculated in your browser. Incident logs, customer names, fees and cost figures are not uploaded.