SLA, Outages & Reliability

What a month of incidents actually delivered, what the contract pays for it, and how little a handful of failures tells you about how often the next one is due.

Three things this is careful about The measurement convention decides the answer — whether planned work counts changes the figure from the same log, so both readings are shown rather than one. A service credit is a contractual remedy, not compensation, so it is put next to what the outage cost. And MTBF from three failures looks precise and is not, so the confidence interval is shown rather than the point estimate alone.

The incident log

ElementMinutesImpact %Start (h)Planned

Against the contract

How often, and how long

Where the downtime came from

Capacity trend

Why the convention matters more than the log

Two parties can read the same incident log and quote availabilities several nines apart, without either of them being dishonest. The differences come from choices that are rarely written down clearly: whether planned maintenance counts, whether a partial outage counts in full or in proportion, whether the clock starts at the alarm or at the ticket, and what the measurement period is.

Most contracts exclude planned work, which is reasonable — it was notified and scheduled. Customers experienced it anyway, so the number they measured will be worse than the number in the report, and neither is wrong. Showing both readings is more useful than defending one.

A credit is not compensation

Service credits are usually expressed as a percentage of a monthly fee, and they are capped. A full day of outage might earn 25% of a month — a few hundred on a typical circuit — against a day of lost trading, idle staff or missed deliveries that cost orders of magnitude more.

That is not a flaw in the contract; it is what the contract is. A credit is a remedy sized to the service, not to your business. Putting the two figures side by side makes the decision about resilience an honest one: the second circuit either pays for itself against the real cost of downtime, or it does not, and the credit barely enters the calculation.

Small samples say very little

Three failures in a year gives an MTBF of 2,920 hours. That number looks like a measurement, and it is not. For failures arriving at a roughly constant rate, the 90% confidence interval around three observations spans more than a factor of five — the truth might be well under 1,500 hours or well over 7,000.

This matters because MTBF figures get compared. A vendor quoting 100,000 hours from a large population and a site quoting 2,920 from three incidents are not making the same kind of statement, and the second cannot be used to contradict or confirm the first. The interval is shown here for that reason.

Repeated short outages make it worse. Three flaps on one link within an afternoon are usually one underlying fault, and logging them as three failures inflates the failure rate while flattering the average repair time — moving both numbers in the direction that looks better.

What this does not do

It does not diagnose anything. Clustering groups incidents that are close in time on the same element; whether they share a cause is a judgement for someone with the logs. Nothing here is a compliance determination, and the contract's own definitions govern what counts as downtime.

Trend extrapolation is a description of past points, not a forecast. Traffic growth is rarely linear, and one new customer moves a link more than a year of trend does.

Everything is calculated in your browser. Incident logs, customer names, fees and cost figures are not uploaded.