The ticket count points at the wrong element
Downtime is frequency times duration. Twelve ten-minute failures cost two hours; two four-hour failures cost eight — and everybody knows about the first one.
The ticket count is the wrong unit
Every operations review starts from the same place: a list of elements sorted by how many incidents each one generated. It is the number the ticket system produces, it is easy to sort, and it is the wrong measure of harm.
Downtime is frequency multiplied by duration, and those two are largely independent. Consider a perfectly ordinary year:
- An access switch fails twelve times, ten minutes each. Total: two hours.
- A core line card fails twice, four hours each. Total: eight hours.
The switch has six times the tickets and a quarter of the impact. Everyone in the building knows about the switch. The line card is quiet, and it is where the year actually went.
There is a mechanism behind this, not just bad luck. Short failures get noticed, logged, escalated and closed — each one generates paperwork. A four-hour outage on a quiet element is one ticket. The reporting system systematically over-represents the frequent and under-represents the severe.
Rank by minutes. When the two rankings disagree, that disagreement is the single most useful thing in the review.
Repair time is shared infrastructure
The second thing counting misses is an asymmetry in what fixes are available.
Fixing a specific element — replacing it, making it redundant — only ever competes against that element's own contribution. Improving repair time competes against everything, because spares on site, a working access arrangement and a rehearsed runbook apply to every incident in the log at once.
So a modest across-the-board improvement is quietly playing a much bigger game. On a log where one element dominates, attacking that element directly still wins. On a log with no dominant cause — which is most mature networks — the boring process work wins outright, and it is usually cheaper.
The useful question for a budget conversation is the inverted one: how much better would our repair process have to get to make this purchase unnecessary? If the answer is "12% faster", the purchase is hard to justify. If it is "60% faster", it is probably the right buy.
Redundancy is not free and does not remove the downtime
A model that treats redundancy as eliminating an element's downtime will recommend redundancy every time, for everything.
It does not eliminate it. Protection switching takes time, and that time is downtime. Some failures are common-mode and take both halves. The redundant pair still needs maintenance windows, and now there are two things to maintain instead of one. The residual is small, and it is not zero.
Making the residual an explicit input rather than an assumption is the difference between a comparison and an advertisement.
Two honest people, one log, two availabilities
Before any of this can be argued about, there is a prior argument that usually goes unnoticed: two conventions differ.
Does planned work count as downtime? The customer generally thinks so; the contract generally says not. Both positions are defensible and they produce different numbers from identical data.
Does a partial outage count in full? Counting a 50% capacity loss as a full outage overstates the harm. Ignoring it entirely understates it. Weighting by the fraction lost is the middle answer, and it is a choice.
An availability figure without its convention attached is not a measurement. State it, and most SLA disputes turn out to be about the convention rather than the network.
Three failures cannot measure an MTBF
Divide the observation window by the number of failures and you get a number that looks like a measurement to four significant figures. With a handful of failures it is barely a rumour.
For a time-terminated observation, the confidence interval on MTBF comes from the chi-squared distribution — and with three failures over a year the 90% interval spans roughly a factor of four. The true value could be a quarter of the estimate or twice it. Any improvement claim you like fits inside that range.
Around ten failures is where the estimate starts to carry information. Even then it assumes the failure rate was roughly constant across the window, which after a significant remediation programme it emphatically was not — the whole point of the programme was to change it.
Zero failures is its own trap, and the more tempting one. It produces a lower bound and nothing else. An absence of failures is not evidence of a long time between them; it is evidence that you have not been watching long enough.
What none of this tells you
Grouping incidents by element is bookkeeping, not diagnosis. Repeated short outages on one element within a few hours usually mean a single underlying fault reported several times — which inflates the failure count and flatters the repair time simultaneously, in opposite directions.
An intervention's effect is an input, not a prediction. Whether spares on site actually deliver a 30% cut in repair time is a question about your organisation, its shift patterns and its access arrangements, and no arithmetic answers it.
And none of it addresses whether the availability target was reasonable in the first place, or whether the money would do more good somewhere else entirely.
Frequently asked questions
Why is the element with the most incidents usually not the problem?
Because downtime is frequency times duration. Twelve ten-minute failures cost two hours; two four-hour failures cost eight. The frequent one has six times the tickets and a quarter of the impact. Reporting also over-represents frequency — short failures each generate paperwork, while a long outage is one ticket.
Why does an across-the-board repair-time cut compete so well?
Because repair time is shared. Spares, access and a rehearsed runbook apply to every incident at once, so a modest improvement competes against the whole log — while fixing one element competes only against that element. On a log with no dominant cause the process work wins outright, and is usually cheaper.
How do I compare a spares contract with a redundant line card?
In minutes returned per pound. It is the only unit in which unlike purchases are comparable. Ranking by anything else — ticket count, how recently something broke, engineer frustration — reliably picks the wrong one.
Why doesn't redundancy remove all the downtime?
Protection switching takes time and that time is downtime; some failures are common-mode and take both halves; and the pair still needs maintenance, with two things to maintain instead of one. A model that assumes redundancy is perfect recommends it every time, for everything.
Why do two people get different availability from the same log?
Two conventions differ: whether planned work counts, and whether a partial outage counts in full. Both positions are defensible and they give different numbers from identical data. An availability figure without its convention is not a measurement.
How many failures do I need to quote an MTBF?
About ten before the estimate carries information. With three failures over a year the 90% confidence interval spans roughly a factor of four — the true value could be a quarter of your estimate or twice it, and any improvement claim fits inside that. Quote the interval, not the point estimate.
What does a year with no failures prove?
That you have not watched long enough. It gives a lower bound on MTBF and nothing else. An absence of failures is not evidence of a long time between them.
Does this find root causes?
No. Grouping by element is bookkeeping. Repeated short outages within a few hours usually mean one underlying fault reported several times, which inflates the failure count and flatters the repair time at the same moment.
Open the Operations Lab →