SLA arithmetic, and the numbers that look more certain than they are
Availability depends on conventions nobody wrote down, credits are sized to the service rather than to your losses, and an MTBF from three failures spans a factor of six.
The convention decides the answer
Two people can read one incident log and quote availabilities several nines apart without either being dishonest. The gap comes from choices that are rarely written down as clearly as the target itself:
- Does planned work count? Most contracts exclude it. Customers experienced it anyway.
- Does a partial outage count in full? Losing one of two links is not the same as losing the service, but counting it as nothing is not right either.
- When does the clock start? At the alarm, at the ticket, or at the customer's call. The three can differ by half an hour.
- What is the period? The same four-hour outage is 99.44% of a 30-day month and 99.95% of a year.
None of this is edge-case pedantry. It is most of the substance of an SLA dispute, and it is why a report and a customer's own measurement disagree so reliably.
What a target actually allows
Over a 30-day month: 99% allows 7.2 hours of downtime, 99.9% allows 43 minutes, 99.99% allows 4.3 minutes. Over a year those become 3.65 days, 8.8 hours and 53 minutes.
Turning the target into a downtime budget is worth doing before signing anything. Forty-three minutes a month sounds tight until you count how long it takes to get an engineer to a cabinet at three in the morning, at which point a single site visit has spent the year's budget.
A credit is a remedy, not compensation
Service credits are a percentage of a monthly fee, and they are capped. A full day of outage on a nine-hundred-a-month circuit might earn a 25% credit — two hundred and twenty-five pounds — against a day of lost trading, idle staff and missed deliveries that runs to tens of thousands.
That is not a flaw in the contract. A credit is sized to the service, not to your business, and no supplier is going to underwrite your revenue for the price of a circuit. But it does mean the credit should play almost no part in a resilience decision. The question is whether a second path, or a shorter repair time, pays for itself against the real cost of an hour down. Against that number the credit is a rounding error.
Three failures do not make a measurement
Three failures in a year gives an MTBF of 2,920 hours. It looks like a measurement. It is an estimate from a sample of three, and the uncertainty is enormous.
For failures arriving at a roughly constant rate, the confidence interval comes from the chi-squared distribution. At 90% confidence, three observations put the true MTBF somewhere between roughly 1,300 and 7,600 hours — a span of nearly six to one. Thirty failures over the same window narrow it to well under two to one.
This matters when figures get compared. A vendor's 100,000-hour MTBF comes from a large population under controlled conditions; a site's 2,920 hours comes from three incidents and a year. They are not the same kind of statement, and the second neither confirms nor contradicts the first.
Counting one fault three times
Repeated short outages on one element within a few hours are usually one underlying fault reported several times — a link flapping before it fails properly, or a card resetting.
Logging them as separate failures moves two numbers at once, and both in the flattering direction. The failure count goes up, so MTBF goes down; the short durations drag the average repair time down, so MTTR looks better. A month of flapping can produce a reliability report that reads worse on frequency and better on responsiveness than what actually happened.
Grouping incidents that fall close together on the same element makes the effect visible. It does not diagnose anything — whether they share a cause is a judgement for someone with the alarms in front of them.
Where the downtime came from
Downtime concentrates. In most months a small number of elements account for the great majority of it, and the ranking is more useful than the total, because it is the only part that suggests what to do.
It is worth grouping by cause as well as by element. The same element appearing repeatedly points at a device; the same cause appearing across different elements points at a process — a change window, a supplier, a spares policy.
Trends describe the past
A straight line through six months of utilisation and a date where it crosses 80% looks like a forecast. It is an extrapolation, and traffic growth is rarely linear.
The fit quality is worth reading first. If the points do not lie near a line, the date means nothing at all. Even with a good fit, one new customer, one codec change or one backup job moves a link further than a year of trend does. Treat the date as a prompt to look, not as a plan.
Nothing leaves your browser
Incident logs describe your network and often your customers. Element names, durations, fees and cost figures all stay on your device; none of it is uploaded, stored or logged.
Frequently asked questions
Why do my supplier and I calculate different availabilities?
Almost always because of conventions rather than data. Whether planned work counts, whether partial outages are weighted, when the clock starts and what the period is all change the figure from the same log. Most contracts exclude planned work, so a supplier's report will usually be better than a customer's own measurement, and neither is wrong.
How much downtime does each target allow?
Over a 30-day month: 99% allows 7.2 hours, 99.9% allows 43 minutes, 99.99% allows 4.3 minutes. Over a year: 3.65 days, 8.8 hours and 53 minutes. Turning a target into a budget is more useful than counting nines.
Will a service credit cover my losses?
No, and it is not designed to. Credits are a capped percentage of a monthly fee, sized to the service rather than to your business. A day of outage might return a few hundred against tens of thousands of cost. Resilience decisions should be made against the real cost of downtime; the credit barely enters the arithmetic.
How reliable is an MTBF from a handful of failures?
Not very. Three failures in a year gives a point estimate of 2,920 hours and a 90% confidence interval spanning roughly 1,300 to 7,600. The number looks like a measurement and is an estimate from a sample of three, which is why the interval should always be quoted with it.
Why does a flapping link distort the numbers?
Because several short outages from one underlying fault get logged as several failures. That raises the failure count, lowering MTBF, while the short durations pull the average repair time down. Both numbers move, and both in the direction that misrepresents what happened.
Should I group incidents by element or by cause?
Both. The same element recurring points at a device. The same cause recurring across different elements points at a process — a change window, a supplier, a spares policy. The second is usually the more valuable finding and the easier one to miss.
Can a utilisation trend tell me when a link will fill?
It can tell you when a straight line through past points would cross a threshold, and how well that line fits. Traffic growth is rarely linear, and one new customer moves a link more than a year of trend does. Read the fit quality first; a poor fit makes the date meaningless.
Open the SLA & Outage Analyser →