Fault localisation · v3.8
"My internet is slow"

The useful answer is not
whether. It is where.

A support desk does not need to be told a subscriber has a problem — the subscriber just told them. It needs to know which segment to send the ticket to: the customer's own line, the operator's uplink, something upstream, or nothing in the network at all.

Every per-subscriber quality detector in this industry reads the subscriber's traffic. Encryption took most of that traffic away. So we stopped reading it — and started reading the heartbeat their own router already sends us.

0 → 1,132
subscribers measurable, same box, same moment
94%
of sessions carry the new signal
3
segments measured separately
7
verdicts, each naming who acts

Why the old approach ran out of road

Loss and jitter used to be measured from the subscriber's own packets: count the retransmissions, watch the sequence numbers. That works while the traffic is TCP and readable in the middle.

QUIC is now roughly two thirds of download traffic and its packet numbers are encrypted, so there is nothing to count. On the gateway used for the measurements in this document, the traffic-reading stack had usable data for zero subscribers. Not few. Zero.

Two things a gateway could read A traffic-reading detector depends on the subscriber's payload being inspectable, which encryption prevents. The access leg reads the PPP keepalive the customer's router sends regardless of what the traffic contains. Reading the subscriber's traffic needs sequence numbers in the clear encrypted transports show nothing quality depends on what they browse 0 subscribers measurable Reading the router's heartbeat the CPE already sends LCP echoes unencrypted, and always present independent of what they browse 1,132 subscribers measurable Same gateway. Same instant. The difference is not a better algorithm — it is looking at a signal encryption never touched.

The keepalive was always there. Every PPPoE router sends the gateway a small "still here?" message on its own schedule, and has done for decades. Nobody was measuring it, because while the traffic was readable there was no need to. It is the one signal a subscriber cannot encrypt, because it is addressed to us.

Three segments, measured separately

A ticket is only actionable once the fault has a location. The gateway measures three, and they answer different questions with different owners.

The three measured segments From the home to the gateway is the access leg, measured per subscriber. From the gateway to the first upstream hop is our uplink. Everything past that is beyond us. The last two are properties of the whole gateway, not of one customer. Home router / CPE The gateway where we measure Upstream transit / peering internet ACCESS LEG per subscriber OUR UPLINK whole gateway BEYOND The distinction that decides the ticket Our uplink is the piece we own and can fix. Beyond is everything we can only route around. Access is one customer's own line. Access is measured per subscriber. Both transit legs are properties of the gateway — identical for everyone on it, and never to be presented as something about one customer.

The verdict, and who it belongs to

Combining the legs produces one of seven answers. The value of the list is that every entry routes the ticket somewhere different.

VerdictWhat it meansWho acts
accessThe customer's line, their router, or something in the homeField technician
shared segmentOLT, backhaul or aggregation — several customers, one causeNOC, high priority
transitOur upstream is degraded; every subscriber on the box is affectedNOC — do not dispatch
cleanBoth legs measured and good — the network is not the problemClose the ticket
access only, okThe line is fine; transit could not be measuredPartial answer
access only, badThe line is degraded; transit could not be measuredField technician
unknownNeither leg could be measuredDo not guess

The last three matter more than they look. A monitoring system that only knows "fault" and "no fault" has nowhere to put a subscriber it could not measure, so it files them under healthy. Two of these seven verdicts exist purely to stop that happening, and every verdict carries flags saying which legs were actually measured. A partial answer is shown as partial rather than dressed up as a confident one.

What an operator actually looks at

The fault location panel: a green verdict reading NOT THE NETWORK, with three cards below for the access leg, our uplink and beyond us, each showing its measurement.
One panel, three legs, one verdict. Read the coloured line first — it is the answer. The cards below it are the evidence, and they are what you quote to the customer or to the upstream provider. Here both transit legs sit close to their own floor and the access leg is quiet, so the honest conclusion is that nothing in the network is responsible.

How to read it in ten seconds

1

The coloured line

Green means the network is measured and good. Red means a segment is at fault and the wording names which. Grey means something could not be measured — which is not the same as fine, and is never shown as a tick.

2

Compare against the floor

Transit legs are judged against their own best-observed value, not a universal threshold. A 22 ms path with a 21.6 ms floor is healthy; the same 22 ms on a network whose floor is 4 ms is not.

3

One customer or all of them

The access card is about this subscriber. The two transit cards are about the whole gateway. If several tickets in a row show the same transit degradation, it was never several tickets.

What it looks like on a real network

Measured on one gateway carrying roughly 1,600 subscribers:

MeasurementResultWhat it tells you
Sessions carrying the signal94%Nearly every router sends keepalives; the few that do not are shown as not measured
Access jitter, typical2.8 ms medianThe normal state of a healthy line, and the baseline everything else is read against
Access jitter, worst tenth40.9 msWhere the interesting lines live
Lines showing upload loss0.49% — 5 in 1,013A short, workable list rather than an alert storm
Uplink floor / beyond floor21.66 ms / 21.58 msThe gateway's own normal, learned rather than assumed
Router polling intervals seen60s, 10s, 5s, 30s, 50sEach CPE is measured on its own cadence, not an assumed one

0.49% is the number to notice. A detector that flags half a percent of subscribers has told you something. One that flags everybody has told you nothing — and we have shipped that mistake before, an indicator that returned a failure for every subscriber on every box and hid a real fault underneath it. If a panel lights up for your whole estate, suspect the panel.

What this does not do

Upload only

No round-trip time

We see the router's message arrive but never learn when it was sent, so this is upload-side variation and upload-side loss. It is not latency, and we do not present it as latency.

Needs a keepalive

A few clients send none

Around 94–99% of real CPE routers send LCP echoes. Some software diallers do not. Those subscribers show as not measured, which is the honest result rather than a gap disguised as health.

Never coming

Encrypted-traffic loss

QUIC packet numbers are protected by design, so per-subscriber loss inside that traffic is unobservable from the network and always will be. This measures a different thing rather than pretending to solve that one.

And one honest caveat about the thresholds. The levels at which jitter and upload loss are called degraded were derived from the shape of a healthy population, not yet from lines known to be physically broken. Treat a crossing as worth investigating, not as a diagnosis. That correlation is the next piece of work, and we would rather say so than let a number look more settled than it is.

The bottom line

Encryption quietly removed the measurement most quality monitoring was built on, and the industry's answer has largely been to keep showing the old screens with less behind them. The alternative is not a cleverer way to read traffic nobody can read. It is to measure something else — something the subscriber's own equipment has been sending all along.

The result is that a support desk stops asking whether there is a problem and starts being told which segment owns it: dispatch, escalate upstream, or close the ticket with evidence.

About the figures and the screen. All measurements come from a production gateway carrying live subscriber traffic, not a test bench. The panel shown is the real console; it contains no operator, site, gateway or subscriber identifier of any kind.