The failure that lives between the reviews
One event strips protection from three services at once. Every service stays up, every dashboard is green, and the estate is one event from an outage.
Every service passes, and the estate is in trouble
A service review is scoped to one service, and correctly so. Are the working and protection paths diverse from each other? Does availability meet the target? Does the latency fit? Every one of those questions can be answered perfectly, for every service you run, while the network as a whole carries a risk nobody has looked at.
The reason is structural. Shared risk is a property of pairs — and a single-service review only ever examines one pair: that service's own two paths. It never looks at the pair formed by two different services, because there is no stage in the process where that is anybody's job.
The three failures that live between reviews
One cut, several services
Two services routed down the same duct on their working paths. Both switch to protection, or both go off. What every model treated as an isolated single-service outage is a multi-service incident, with a different escalation, a different set of customer calls, and a longer recovery because everything is competing for the same field engineer.
Nobody planned it and nobody was careless. Each service was routed sensibly, and the routes were sensible because the duct is where the ducts go.
One cut, no alarm, and the estate quietly changes state
This is the one worth building a tool for, because it produces no signal at all.
A single event strips the protection from several services simultaneously. Every service stays up. Every dashboard is green. No SLA is breached, no incident is raised, and nobody is paged.
But three services that were protected an hour ago are now running unprotected on single paths. A second event — anywhere, on any of their remaining paths — now has three targets instead of one, and the probability of a customer-visible outage has multiplied.
Nothing reported this, because protected and unprotected but working look identical from outside. The service is up either way. The only difference is what happens next, and monitoring does not measure what happens next.
The pair that was sold as diverse
A customer asks for two circuits, diverse from each other. Two orders are raised. Two designs are produced. Two sets of paperwork are correct. Both circuits run down the same duct.
Nothing in either design is wrong. The requirement was about the relationship between two orders, and relationships between orders are not what a design review examines. Unless somebody asks the question explicitly, as a separate question, nothing in the process will ever ask it.
Blast radius, not availability
Availability is computed per service and cannot express "this one duct takes three of them". The unit that can is blast radius: for every shared-risk group in the estate, which services lose their working path, which lose their protection, and which lose both.
Laid out as a matrix — one row per possible failure, one column per service — it makes two things visible at a glance that no per-service document contains: the rows with more than one service down, and the rows with several services exposed but nothing down. The second kind of row is the one that never appears anywhere else.
Concentration is the other question
Blast radius asks "what breaks this service". The inverse is also worth asking: what have we quietly put all our eggs in?
An element carrying most of the estate is a maintenance window nobody can schedule — there is no time when taking it down is acceptable — and a failure nobody can absorb. That is a planning problem rather than an incident, and it does not show up until the day somebody needs to do work in that duct.
The tags are a survey, not a fact
All of this rests on risk tags, and a tag is a claim somebody made about physical infrastructure. Analysis of declared tags is exactly as good as the survey behind them, and untagged shared infrastructure is invisible — which is the usual reason a protected service fails.
The things most often missed are ordinary rather than exotic: a shared chamber where two routes cross, a bridge carrying both over the same river, a single building entry, one power feed behind two genuinely diverse fibres, a stretch of the same road verge where separate ducts sit a metre apart.
Diversity that exists on a drawing and not in a trench is the most expensive kind of diversity, because it is paid for and does not work.
One failure at a time is an assumption
A blast-radius table assumes independent single failures. Real correlated causes do not respect that: a storm, a regional power event, or a contractor working the length of a street takes out several tagged groups at once.
When that happens the blast radius is the union of several rows, which is larger than anything in the table — and it is the case worth thinking about deliberately, because it is the one where the escalation plan matters more than the design.
Frequently asked questions
Why can't per-service reviews find these problems?
Because shared risk is a property of pairs, and a single-service review only examines one pair — that service's own two paths. It never looks at the pair formed by two different services, and there is no stage in the process where that is anybody's job.
What is blast radius?
For each shared-risk group in the estate, which services lose their working path, which lose their protection, and which lose both. Availability is computed per service and cannot express 'this one duct takes three of them'; blast radius can.
Why does a failure with nothing down still matter?
Because it can strip protection from several services at once. They all stay up, every dashboard is green, no SLA is breached and nobody is paged — but three services that were protected are now on single paths, and a second event anywhere has three targets instead of one. Monitoring cannot see it, because protected and unprotected-but-working look identical from outside.
How do two circuits sold as a diverse pair end up in the same duct?
Two orders, two designs, two correct sets of paperwork. The requirement was about the relationship between the orders, and relationships between orders are not what a design review examines. It has to be asked as a separate question or it is never asked.
Is a working-to-protection overlap as bad as two working paths sharing?
No, and the difference is worth keeping. Two working paths in one duct means a single cut takes both circuits off. A working-to-protection overlap means one cut degrades the pair — one service switches and survives, the other loses its protection.
What is concentration risk?
The inverse question: not what breaks a service, but what carries most of the estate. An element most services depend on is a maintenance window nobody can schedule and a failure nobody can absorb. It is a planning problem, and it stays invisible until somebody needs to work in that duct.
How reliable is this analysis?
Exactly as reliable as the risk tags. Untagged shared infrastructure is invisible and is the usual reason a protected service fails. The things most often missed are ordinary — a shared chamber, a bridge, a building entry, one power feed behind two diverse fibres, the same road verge.
Does this cover storms and regional events?
No. It assumes one failure at a time. Correlated causes take out several tagged groups at once, and the blast radius of that is the union of several rows — larger than anything in the table, and the case where the escalation plan matters more than the design.
Open the Network Designer Lab →