NOC2: the BNGSOFT network operations centre Path visibility · Plain-words verdicts · Incidents
NOC2 · Internet Observatory · New
Know where the problem is before your customer calls.
The Internet Observatory follows every gateway's path link by link, from the subscriber's line to the servers they actually use, and says in plain words which link is at fault. Your network, your upstream, or the internet itself: one page, one verdict, with the evidence underneath.
One gateway's path, live
!
Problem on access: one VLAN's line-fault share is 6× this gateway's other VLANs.Gateway, DNS and uplink are clean, so the core team is not paged.
!problem~watch✓clean?not measured, never green
7 links
subscriber line to destination, on one page
Plain words
a verdict per link, not a graph to interpret
“It's not you”
proof when the fault is beyond your network
1 incident
not 200 alerts for one cause
01The problem: everyone is looking, nobody can see
When subscribers say "the internet is slow", an operator usually has graphs for the parts it owns and nothing for the parts in between. So the call goes round the building.
Blind spots
?
Links nobody measures
Line quality, a single bad VLAN, the resolver subscribers really use, the answer time of the servers they visit. Each one can be the fault, and most dashboards show none of them.
Wrong team
→←
Blame lands in the wrong place
Access blames core, core blames upstream, upstream says it is clean. Hours pass before anyone proves where the loss is, and the customer hears "we are looking into it".
Alert storm
200
Messages for one fault
One upstream problem can raise an alert per gateway, per VLAN, per subscriber. The real cause is buried in the noise, and people learn to ignore the pager.
02Walk the chain: what each link measures, and what it says
Click any link (or use the arrow keys). Every verdict below is the kind of sentence NOC2 writes on the page.
What is measured
Per-subscriber line-fault detection read straight from the bngxdpd forwarding plane: loss and reordering, upload and download.
The fault is located: the subscriber's line, CPE or in-home side, versus our side.
No probe box, no agent on the customer's router.
What it says
!Line fault: 64.9% of packets lost on the way to this subscriber.
✓No fault found on our path.
Why it matters: a van goes out only when the fault really is on the line.
What is measured
A per-VLAN census of line faults, compared with the gateway's other VLANs.
A VLAN is flagged at ≥5 faults, ≥2% share and ≥3× the median of the other VLANs.
A flag must hold for 3 readings in a row (about 6 min) and clears with a dead band, so borderline VLANs do not flicker.
After a gateway restart nothing is judged for 30 min.
What it says
!VLAN 1204: line-fault share 6× the median of this gateway's other VLANs, held for 3 readings.
?Restarted 12 min ago: VLANs are not judged yet.
Measured on real restarts: persistence cut flag/clear flapping by about 91%, and the dead band by a further 36%.
What is measured
A health verdict per gateway.
Real RADIUS login time: successful logins answered in about 16 ms, measured apart from rejects, which RADIUS servers deliberately hold for about 1 s.
Address-pool fill and DDoS protection state.
What it says
✓Logins healthy: successful RADIUS answers in ~16 ms.
~Address pool filling: watch before new logins are refused.
Honest timing: deliberately delayed rejects never make your login time look slow.
What is measured
The gateway probes its own resolver, the one subscribers use, and each upstream, every 10 s.
A failing local resolver is a problem. Upstream trouble is a warning, unless all upstreams fail.
Fewer than 20 probes means "not judged yet", never healthy.
What it says
✓Local resolver answers in 0.07 ms.
~One upstream resolver is slow; the others are answering.
?Not judged yet: fewer than 20 probes.
What is measured
Each gateway probes its next hop and public targets such as 8.8.8.8 and 1.1.1.1.
Loss over a sliding window, and round-trip time against a rolling one-hour floor, so "slower than usual" means something.
A 24 h timeline and 30 days of history.
What it says
✓Uplink clean: next hop at its one-hour floor, no loss.
!Loss at the next hop: the problem starts on your own uplink.
What is measured
The same public targets, compared across gateways of different operators.
When a target is degraded on ≥3 gateways of ≥2 operators while their own uplinks are clean, the fault is outside every one of those networks.
Each operator sees only its own gateways; facts from other operators appear only as anonymous counts.
What it says
~Beyond our network: 8.8.8.8 is slow from gateways of more than one operator, and their uplinks are clean. It's the internet, not you.
Why it matters: you can tell customers the truth, and stop escalating to a team that has nothing to fix.
What is measured
Where subscribers' new connections go, per server address block, and who owns it: Google, Meta, Akamai, Amazon, Microsoft and more.
How fast those servers answer, measured at the gateway so the subscriber's own line is left out.
IPv4 and IPv6, TCP and QUIC. Percentiles from histograms, never averages.
Slow server answers are counted on their own and never called loss.
What it says
~Could be faster: Google answers in 20–40 ms over IPv6 but 1–2 ms over IPv4 on this gateway. IPv6 traffic likely bypasses the local cache.
03“It's not you”: proof that the fault is beyond your network
One gateway seeing 8.8.8.8 slow proves nothing. It could be your uplink. But when gateways of different operators, on different uplinks, all see the same target degrade at the same time while their own links are clean, the only shared part left is the internet.
≥3 / ≥2
The ruleSame public target degraded on at least 3 gateways of at least 2 operators, each with a clean uplink: NOC2 says beyond our network.
Try it: make your own uplink lossy and watch the verdict change. NOC2 never lets the internet take the blame for a fault at home.
Beyond our network.8.8.8.8 is slow from gateways of more than one operator, and every uplink is clean. Tell the customer it's the internet, and stop the escalation.
04Destinations: where your subscribers could be faster
The Observatory watches the servers your subscribers actually reach, and how fast they answer from your gateway. That turns "the network feels slow" into a finding with numbers attached.
Could be faster
Same gateway, same Google, two very different answers
Verdict: IPv6 traffic likely bypasses the local cache. Fix the IPv6 path, and every dual-stack subscriber gets the fast answer.
CDN & peering advisor
A cache or closer peering, with the size of the win
"23% of new connections go to Google at 20–40 ms. Gateways elsewhere reach Google in 1–2 ms. A local cache or closer peering could serve them."
23%of new connections affected
20–40 msyour answer time today
1–2 msachievable elsewhere
Benchmark
"Your Google answer time is slower than 72% of operators measured." Example reading. Shown only when at least 3 other operators measure the same destination, and never names them.
05Subscriber diagnosis: answer the support call in seconds
Type the caller's username and get one verdict with the evidence chain under it. Try it below; this demo runs entirely in your browser on anonymised examples.
Examples:
06The incident engine: one incident, not 200 alerts
Every finding above is evidence, not an alert. NOC2 groups related evidence into a single incident with a start and an end, a likely cause, an estimate of affected subscribers and the evidence that supports it.
!VLAN flagaccess
~Upstream findinguplink
~Slow destinationdestinations
!DDoSprotection
~Flooding devicesubscriber
!Gateway problemBNG
INCIDENT · OPEN
Access fault: line faults concentrated on VLAN 1204, gw-north-01
Started
14:02, still open
Likely cause
access segment behind one VLAN
Est. affected
~140 subscribers
Evidence
4 signals, 1 incident
VLAN flag held 3 readings · subscriber line faults on that VLAN only · gateway healthy · uplink clean. Unknown evidence never closes an incident.
200 alerts→1 incident, with a cause
Illustrative incident; names and numbers are examples.
07Built to be believed
A monitoring page is only useful if a green tile means something. Three rules run through every number in the Observatory.
Unknown is never green
A link without enough data says so: "not judged yet", "not measured". Silence is never painted as health, and unknown evidence never closes an incident.
Percentiles, not averages
Real example: an average of 12 ms hid that 99% of connections were under 5 ms. NOC2 reads histograms and reports the percentiles.
Coverage on every number
Every count says how much it is based on. You can see at a glance whether a clean result covers the whole gateway or just part of it.
Our promise: where the product is sure, it says so. Where it is likely, it says likely. Where it has not measured, it shows a question mark. Nothing unmeasured is ever shown as healthy.
08What operators get
The faulty link in seconds
Seven links, one verdict each, in plain words. No more reading twelve graphs to guess.
The right team, first time
Access, core or upstream: the evidence says which, so nobody spends the afternoon proving it was not them.
Proof it's not you
When the internet is at fault, you can tell customers so, backed by a rule, not a hunch.
Faster users, with numbers
See where a cache or closer peering would help, and by how much, before you sign anything.
Fewer alerts, better incidents
One incident with a cause, an estimate of who is affected and its evidence, instead of a storm.
A page you can trust
Unknown is shown as unknown. Green means measured and clean, every time.
See your own network's path.
We will walk through the Internet Observatory on a live NOC2: the chain, the "it's not you" verdict, the destinations advisor and a subscriber diagnosis, on your questions.
Examples on this page are illustrative and anonymised. Usernames, gateway names, VLAN numbers and addresses are invented. 8.8.8.8 and 1.1.1.1 are public resolvers used as probe targets. Google, Meta, Akamai, Amazon and Microsoft are trademarks of their respective owners, named only as destinations that subscribers reach; BNGSOFT is not affiliated with them.