NOC2: the BNGSOFT network operations centre
Path visibility · Plain-words verdicts · Incidents
NOC2 · Internet Observatory · New

Know where the problem is before your customer calls.

The Internet Observatory follows every gateway's path link by link, from the subscriber's line to the servers they actually use, and says in plain words which link is at fault. Your network, your upstream, or the internet itself: one page, one verdict, with the evidence underneath.
One gateway's path, live
!
Problem on access: one VLAN's line-fault share is 6× this gateway's other VLANs.Gateway, DNS and uplink are clean, so the core team is not paged.
!problem ~watch ✓clean ?not measured, never green
7 links
subscriber line to destination,
on one page
Plain words
a verdict per link,
not a graph to interpret
“It's not you”
proof when the fault is
beyond your network
1 incident
not 200 alerts
for one cause

01The problem: everyone is looking, nobody can see

When subscribers say "the internet is slow", an operator usually has graphs for the parts it owns and nothing for the parts in between. So the call goes round the building.

Blind spots
?

Links nobody measures

Line quality, a single bad VLAN, the resolver subscribers really use, the answer time of the servers they visit. Each one can be the fault, and most dashboards show none of them.

Wrong team
→←

Blame lands in the wrong place

Access blames core, core blames upstream, upstream says it is clean. Hours pass before anyone proves where the loss is, and the customer hears "we are looking into it".

Alert storm
200

Messages for one fault

One upstream problem can raise an alert per gateway, per VLAN, per subscriber. The real cause is buried in the noise, and people learn to ignore the pager.

02Walk the chain: what each link measures, and what it says

Click any link (or use the arrow keys). Every verdict below is the kind of sentence NOC2 writes on the page.

What is measured

  • Per-subscriber line-fault detection read straight from the bngxdpd forwarding plane: loss and reordering, upload and download.
  • The fault is located: the subscriber's line, CPE or in-home side, versus our side.
  • No probe box, no agent on the customer's router.

What it says

!Line fault: 64.9% of packets lost on the way to this subscriber.

✓No fault found on our path.

Why it matters: a van goes out only when the fault really is on the line.

What is measured

  • A per-VLAN census of line faults, compared with the gateway's other VLANs.
  • A VLAN is flagged at ≥5 faults, ≥2% share and ≥3× the median of the other VLANs.
  • A flag must hold for 3 readings in a row (about 6 min) and clears with a dead band, so borderline VLANs do not flicker.
  • After a gateway restart nothing is judged for 30 min.

What it says

!VLAN 1204: line-fault share 6× the median of this gateway's other VLANs, held for 3 readings.

?Restarted 12 min ago: VLANs are not judged yet.

Measured on real restarts: persistence cut flag/clear flapping by about 91%, and the dead band by a further 36%.

What is measured

  • A health verdict per gateway.
  • Real RADIUS login time: successful logins answered in about 16 ms, measured apart from rejects, which RADIUS servers deliberately hold for about 1 s.
  • Address-pool fill and DDoS protection state.

What it says

✓Logins healthy: successful RADIUS answers in ~16 ms.

~Address pool filling: watch before new logins are refused.

Honest timing: deliberately delayed rejects never make your login time look slow.

What is measured

  • The gateway probes its own resolver, the one subscribers use, and each upstream, every 10 s.
  • A failing local resolver is a problem. Upstream trouble is a warning, unless all upstreams fail.
  • Fewer than 20 probes means "not judged yet", never healthy.

What it says

✓Local resolver answers in 0.07 ms.

~One upstream resolver is slow; the others are answering.

?Not judged yet: fewer than 20 probes.

What is measured

  • Each gateway probes its next hop and public targets such as 8.8.8.8 and 1.1.1.1.
  • Loss over a sliding window, and round-trip time against a rolling one-hour floor, so "slower than usual" means something.
  • A 24 h timeline and 30 days of history.

What it says

✓Uplink clean: next hop at its one-hour floor, no loss.

!Loss at the next hop: the problem starts on your own uplink.

What is measured

  • The same public targets, compared across gateways of different operators.
  • When a target is degraded on ≥3 gateways of ≥2 operators while their own uplinks are clean, the fault is outside every one of those networks.
  • Each operator sees only its own gateways; facts from other operators appear only as anonymous counts.

What it says

~Beyond our network: 8.8.8.8 is slow from gateways of more than one operator, and their uplinks are clean. It's the internet, not you.

Why it matters: you can tell customers the truth, and stop escalating to a team that has nothing to fix.

What is measured

  • Where subscribers' new connections go, per server address block, and who owns it: Google, Meta, Akamai, Amazon, Microsoft and more.
  • How fast those servers answer, measured at the gateway so the subscriber's own line is left out.
  • IPv4 and IPv6, TCP and QUIC. Percentiles from histograms, never averages.
  • Slow server answers are counted on their own and never called loss.

What it says

~Could be faster: Google answers in 20–40 ms over IPv6 but 1–2 ms over IPv4 on this gateway. IPv6 traffic likely bypasses the local cache.

03“It's not you”: proof that the fault is beyond your network

One gateway seeing 8.8.8.8 slow proves nothing. It could be your uplink. But when gateways of different operators, on different uplinks, all see the same target degrade at the same time while their own links are clean, the only shared part left is the internet.

≥3 / ≥2
The ruleSame public target degraded on at least 3 gateways of at least 2 operators, each with a clean uplink: NOC2 says beyond our network.

Try it: make your own uplink lossy and watch the verdict change. NOC2 never lets the internet take the blame for a fault at home.

Beyond our network.8.8.8.8 is slow from gateways of more than one operator, and every uplink is clean. Tell the customer it's the internet, and stop the escalation.
Gateways of different operators see one public target slow Illustration: gateways belonging to different operators each have a clean uplink, yet all of them see 8.8.8.8 degraded. The shared part is the internet. THE INTERNET 8.8.8.8 degraded ✓ ✓ ✓ ✓ UPLINKS gatewayoperator A gatewayoperator B gatewayoperator C yourgateway Illustration. Other operators are never named.

04Destinations: where your subscribers could be faster

The Observatory watches the servers your subscribers actually reach, and how fast they answer from your gateway. That turns "the network feels slow" into a finding with numbers attached.

Could be faster

Same gateway, same Google, two very different answers

Google answer time from one gateway, IPv4 versus IPv6 Over IPv4 Google answers in 1 to 2 milliseconds; over IPv6 it answers in 20 to 40 milliseconds. 0 10 20 30 40 ms IPv4 IPv6 1–2 ms, served close by 20–40 ms Measured at the gateway, subscriber line excluded

Verdict: IPv6 traffic likely bypasses the local cache. Fix the IPv6 path, and every dual-stack subscriber gets the fast answer.

CDN & peering advisor

A cache or closer peering, with the size of the win

"23% of new connections go to Google at 20–40 ms. Gateways elsewhere reach Google in 1–2 ms. A local cache or closer peering could serve them."

23%of new connections affected
20–40 msyour answer time today
1–2 msachievable elsewhere
Benchmark

"Your Google answer time is slower than 72% of operators measured." Example reading. Shown only when at least 3 other operators measure the same destination, and never names them.

05Subscriber diagnosis: answer the support call in seconds

Type the caller's username and get one verdict with the evidence chain under it. Try it below; this demo runs entirely in your browser on anonymised examples.

Examples:

    06The incident engine: one incident, not 200 alerts

    Every finding above is evidence, not an alert. NOC2 groups related evidence into a single incident with a start and an end, a likely cause, an estimate of affected subscribers and the evidence that supports it.

    !VLAN flagaccess
    ~Upstream findinguplink
    ~Slow destinationdestinations
    !DDoSprotection
    ~Flooding devicesubscriber
    !Gateway problemBNG
    INCIDENT · OPEN

    Access fault: line faults concentrated on VLAN 1204, gw-north-01

    Started
    14:02, still open
    Likely cause
    access segment behind one VLAN
    Est. affected
    ~140 subscribers
    Evidence
    4 signals, 1 incident
    VLAN flag held 3 readings · subscriber line faults on that VLAN only · gateway healthy · uplink clean. Unknown evidence never closes an incident.
    200 alerts1 incident, with a cause

    Illustrative incident; names and numbers are examples.

    07Built to be believed

    A monitoring page is only useful if a green tile means something. Three rules run through every number in the Observatory.

    Unknown is never green

    A link without enough data says so: "not judged yet", "not measured". Silence is never painted as health, and unknown evidence never closes an incident.

    Percentiles, not averages A histogram where most answers are under 5 ms and a small tail is very slow; the average of 12 ms sits where almost no answers are 99% < 5 ms avg 12 ms

    Real example: an average of 12 ms hid that 99% of connections were under 5 ms. NOC2 reads histograms and reports the percentiles.

    Coverage on every number

    Every count says how much it is based on. You can see at a glance whether a clean result covers the whole gateway or just part of it.

    Our promise: where the product is sure, it says so. Where it is likely, it says likely. Where it has not measured, it shows a question mark. Nothing unmeasured is ever shown as healthy.

    08What operators get

    The faulty link in seconds

    Seven links, one verdict each, in plain words. No more reading twelve graphs to guess.

    The right team, first time

    Access, core or upstream: the evidence says which, so nobody spends the afternoon proving it was not them.

    Proof it's not you

    When the internet is at fault, you can tell customers so, backed by a rule, not a hunch.

    Faster users, with numbers

    See where a cache or closer peering would help, and by how much, before you sign anything.

    Fewer alerts, better incidents

    One incident with a cause, an estimate of who is affected and its evidence, instead of a storm.

    A page you can trust

    Unknown is shown as unknown. Green means measured and clean, every time.

    See your own network's path.

    We will walk through the Internet Observatory on a live NOC2: the chain, the "it's not you" verdict, the destinations advisor and a subscriber diagnosis, on your questions.

    Request a demo Browse all brochures

    Examples on this page are illustrative and anonymised. Usernames, gateway names, VLAN numbers and addresses are invented. 8.8.8.8 and 1.1.1.1 are public resolvers used as probe targets. Google, Meta, Akamai, Amazon and Microsoft are trademarks of their respective owners, named only as destinations that subscribers reach; BNGSOFT is not affiliated with them.

    Product page: Support & operations · Talk to us