Transport Network Designer Lab

Several services over one physical network. Every one of them passes its own review — and the failures that matter most do not exist inside any single service.

The failure that lives between the reviews Service A’s working path and service B’s protection path both run through duct 7. Neither review mentions the other service; why would it. A digger finds duct 7, A switches to protection and survives, B quietly loses the protection it was sold, and nobody knows until the second event. The unit that catches this is not availability — it is blast radius, counted across the whole estate.

1The services

Risk tags are comma separated — a duct, a bridge, a building entry, a power feed, a chamber. Two paths are diverse only if they share none of them, and the tags are only as good as the survey behind them.

2Pairs sold as diverse

Two circuits described to a customer as diverse from each other. Each one’s own paperwork can be perfectly correct while the pair is wrong, because no single-service review looks at pairs.

3One failure at a time

service down up, but protection gone unaffected

4What only the estate can tell you

Why per-service review cannot see this

A service review is scoped to one service, and correctly so. It asks whether the working path and the protection path are diverse from each other, whether the availability meets the target, whether the latency fits. Every one of those questions can be answered perfectly while the estate is in trouble.

The reason is that shared risk is a property of pairs, and a single-service review only ever looks at one pair — a service's own two paths. It never looks at the pair formed by two different services, and that is where the interesting failures live.

The failure that takes several services

The obvious case: two services both routed down the same duct on their working paths. Both switch to protection, or worse, both go off. What was modelled as an isolated single-service outage is a multi-service incident with a different escalation path, a different customer conversation and a different recovery time.

Nobody planned it. Each service was routed sensibly on its own merits, and the routes were sensible because the duct is where the ducts go.

The failure nobody sees at all

The subtler case is worse, because it produces no alarm. One event strips the protection from several services simultaneously. Every service is still up. Every dashboard is green. No SLA is breached and no incident is raised.

But the estate has quietly changed state: a second event — anywhere, on any of those services' remaining paths — now has several targets instead of one. The risk has multiplied and nothing reported it, because "protected" and "unprotected but working" look identical from outside.

This is why the middle column of the matrix matters as much as the red one.

The pair that was sold as diverse

A customer asks for two circuits, diverse from each other. Two orders are raised, two designs are produced, two sets of paperwork are correct, and both circuits run down the same duct.

Nothing in either design is wrong. The requirement was about the relationship between them, and relationships between orders are not what a design review examines. It is worth checking explicitly, as a separate question, because nothing else in the process will.

Tags are a survey, not a fact

Everything here rests on the risk tags, and a tag is a claim someone made about physical infrastructure. Untagged shared infrastructure is invisible to this analysis and is the usual reason a protected service fails.

The things most often missed are not exotic: a shared chamber where two routes cross, a bridge carrying both, a building entry, a single power feed behind two diverse fibres, a stretch of the same road verge. Diversity that exists on a drawing and not in a trench is the most expensive kind.

What this does not do

It does not compute availability — that is the path designer's job, one service at a time, and the shared-risk tool's. It does not route anything, cost anything, or know where your ducts actually run.

It assumes one failure at a time. Correlated causes — a storm, a regional power event, a contractor working a whole street — take out several tagged groups at once, and the blast radius of that is larger than anything in the table.

Everything is calculated in your browser. Service names, route names and risk tags describe your network and are not uploaded.