HomeGuides › Availability & shared risk

Availability, downtime and why “diverse” often is not

Redundancy arithmetic is simple. The assumption underneath it is the part that fails, and it fails by a factor of hundreds rather than a few percent.

Nines are a bad unit for a conversation

“Four nines” sounds like a specification. “Fifty-three minutes of outage a year” is the same statement, and it is the one people can actually argue about. The two are interchangeable, so it is worth knowing the conversions by heart:

Each nine is a factor of ten. That is why the gap between four and five nines is a different kind of engineering problem from the gap between two and three, and why arguments about the last nine cost so much more than the ones before it.

Series is easy and pessimistic

Anything that must work for the service to work sits in series, and availabilities multiply. A chain is therefore always worse than its worst element, and adding anything to a path — however reliable — can only make it worse. A patch panel at 99.9999% still costs you thirty seconds a year.

The useful consequence is that one element usually dominates. If the access fibre is 99.95% and everything else is 99.999%, the access fibre is essentially the whole answer, and money spent anywhere else buys nothing. Working out which element contributes most of the downtime is normally more valuable than the total.

Parallel is easy and dangerous

Redundancy multiplies unavailabilities. Two paths that are each down 0.1% of the time are both down 0.1% of 0.1% of the time — one part in a million. Two three-nines paths become six nines, and eight hours of annual downtime becomes thirty seconds.

Every bit of that improvement rests on the word both, and both assumes the two failures are unrelated. That assumption is the entire product. Without it the calculation is not conservative or approximate; it is simply not about your network.

What a shared duct actually costs

Take two paths, each 99.9% available, and suppose they run through one duct that is itself 99.95% available. On a diagram they are two paths. In the ground they are two paths and a single point of failure.

The shared duct sits in series with the whole arrangement. When it is cut, both paths are down together, and the service is down for as long as the duct is. The redundancy either side of it changes nothing about that.

The naive parallel answer is 99.9999%, or about half a minute a year. The real answer is a little under 99.95%, or about four and a half hours a year. The design looks identical and is roughly five hundred times worse. That factor is not a rounding error or a safety margin; it is the difference between a service that meets its contract and one that does not.

Shared risk groups

The way to catch this is to stop asking whether the routes are different and start asking what they have in common. Tag every element with the physical things that could take it out, and look for tags that appear on both paths.

Ducts and trenches. The commonest genuine shared risk and the one most often missed, because two circuits ordered years apart from the same carrier frequently end up in the same subduct without anyone being told.

Bridges and structures. Routes that diverge by forty kilometres can still cross the same river on the same bridge. Geography concentrates risk in places a logical diagram has no way to show.

Building entries and chambers. Diversity across a country that arrives through one riser is diversity everywhere except where it matters.

Power. The one that catches people who did the fibre survey properly. Two genuinely diverse fibres landing on equipment fed from one rectifier are one path with extra steps.

Splice closures. A shared joint ties together the fate of everything inside it, including cables that are otherwise unrelated.

Geographic corridors. Separate ducts along the same road are separate until the road is dug up.

MTBF, MTTR, and the lever you actually hold

Steady-state availability is MTBF divided by MTBF plus MTTR. MTBF comes from the vendor and describes a population rather than your unit; you buy what is on the market and cannot move it much.

MTTR is set by spares holding, site access, contract terms and who is on call — all things you control. Halving repair time halves downtime exactly, and it does so whether or not there is a second path. For most services it is the cheapest availability you can buy.

What the arithmetic cannot tell you

Steady-state availability is a long-run average over many years. It says nothing about whether a particular year contains a bad month, and it does not model correlated weather, seasonal digging, wear-out, planned maintenance windows, or the time a protection switch itself takes to complete.

Treat the numbers as a way of comparing two designs rather than as a forecast of next year. The comparison is where the value is: it is what turns “we have diverse paths” into “we have diverse paths except for four hundred metres of duct, and separating that is worth four hours a year.”

Nothing leaves your browser

The calculator runs entirely on your device. Element names, availabilities and risk tags describe your network, and none of it is uploaded, stored or logged anywhere.

Frequently asked questions

Why do two 99.9% paths not give 99.9999%?

They do only if nothing can take both down at once. Share a duct that is itself 99.95% available and the service is capped a little under 99.95% — about four and a half hours a year instead of thirty seconds, roughly five hundred times worse for a design that looks identical on a diagram.

What counts as a shared risk group?

Anything that would take out more than one path at once: a duct, a trench, a bridge, a chamber, a riser, a splice closure, a power feed, a site, or a stretch of road. Diversity is the absence of shared groups, which is a statement about physical infrastructure rather than about how routes are drawn.

Is diverse fibre enough on its own?

No. Two separate fibre routes landing on equipment fed by a single rectifier are one path with extra steps. Power, building entry and the final chamber defeat diversity more often than the long-haul route does.

Should I spend on MTBF or MTTR?

MTTR, almost always. MTBF is set by the equipment you can buy. Repair time is set by spares, access, contracts and on-call cover, and halving it halves downtime exactly — usually cheaper than a second path, and it helps even when one exists.

How many minutes is each nine?

Two nines is 3.65 days a year, three nines 8.76 hours, four nines 52.6 minutes, five nines 5.26 minutes, six nines 31.5 seconds. Each nine is a factor of ten.

Why does adding a very reliable element make a chain worse?

Because everything in series must work. Availabilities multiply, and multiplying by anything below 1 lowers the result. A patch panel at 99.9999% still costs about thirty seconds a year.

Can this predict next year's downtime?

No. Steady-state availability is a long-run average across many years. It cannot say whether a given year holds a bad month, and it does not model correlated weather, wear-out, maintenance windows or protection-switch time. Use it to compare designs, not to forecast.

Open the Availability & Shared Risk Calculator →