The user reported “terrible internet speed” and suspected DNS. It is not DNS. 35 days of net-monitor data plus live capture show episodic multi-day WAN degradation on the fixed-wireless uplink (TP-Link NE200), upstream of the LAN and upstream of telep-router. DNS, IPv6 (the 2026-07-28 root cause), the LAN, and bufferbloat were each ruled out with evidence. A separate, unrelated hardware fault was found on the router’s 10g-sfp port. RF evidence captured 2026-08-31 confirms the root cause is the cellular radio link (see RF baseline captured 2026-08-31 — root cause confirmed).
RESOLVED 2026-09-04 — the WAN was replaced, not repaired
This investigation is closed out by the migration to Starlink in Bypass mode: see 2026-09-04-starlink-wan-migration-dish-telemetry. The NE200’s own RF logger wrote its final entry at
2026-09-04T02:15still degraded —RSRP -103 dBm, SINR 9, downlink QPSK— i.e. exactly the condition diagnosed below, right to the end. Post-migration the link measures 235-279 Mbps wired at 18.3-35.3 ms with 0% loss, and the NE200 at192.168.254.1is 100% unreachable and out of the path.The
ne200_signal.pylogger’s*/15cron line was removed, but the script and its 342 rows ofrflog.csvwere deliberately preserved — that data is the evidence base that justified the switch. Do not delete it.⚠️ Note that Starlink is still CGNAT (
100.64.0.0/10), so the “no inbound ports” consequence of the old double-NAT still applies.
The NE200 is a 5G NR cellular FWA CPE, NOT a WISP point-to-point link
Earlier vault text called this a “TP-Link NE200 outdoor fixed-wireless unit,” which read like a WISP/point-to-point radio. It is not. The NE200 is a 5G NR cellular fixed-wireless-access (FWA) CPE with a SIM card on Telekom Hungary (Telekom HU NET) — a mobile-network link, not a line-of-sight WISP bridge. Its uplink quality is governed by cellular RF metrics (RSRP/RSRQ/SINR) and by whatever carrier data-cap / fair-use policy the SIM plan carries. This reframes every “fixed-wireless” mention below.
For Agents
Verdict: episodic WAN degradation on the NE200 5G NR cellular FWA uplink (Telekom HU SIM). Throughput collapses 5-50x (normal 180-265 Mbit → 3-17 Mbit) in multi-day episodes separated by clean periods; latency stays healthy (~25-40 ms) throughout, so this is a capacity/loss fault, not a latency fault. DNS is fast (30 ms) — do not re-investigate DNS. The IPv6 fix is still holding and is NOT recurring. Evidence lives in
/home/levander/net-monitor/netlog.csvon telep-mainframe. RF baseline captured (2026-08-31): SS-SINR ~9.5 dB, SS-RSRP -102 dBm on N78 — weak cellular radio, root cause confirmed. An automated RF logger is now deployed (ne200_signal.pyon telep-mainframe, every 15 min →rflog.csv) — scripted NE200 login solved; see NE200 RF logger — deployed 2026-08-31. Real-time confirmation (~12:43-12:45): serving N78 cell on QPSK downlink with RSRP swinging -100→-108 dBm in 2 min — direct proof the radio is the cause; see Confirming evidence 2026-08-31 ~12:43-12:45 — real-time QPSK collapse. Open competing hypothesis: a carrier data-cap throttle explains the multi-day blocks better than pure RF jitter (777 GB used) — see RF baseline captured 2026-08-31 — root cause confirmed.
Verdict
NOT DNS. Root cause is episodic WAN degradation on the NE200 5G NR cellular FWA uplink (Telekom HU SIM, Double-NAT behind it), upstream of the LAN. Confirmed 2026-08-31 by RF stats read from the NE200 web UI: the modem is running on a weak cellular signal (SS-SINR ~9.5 dB, SS-RSRP -102 dBm on N78 3.5 GHz), which caps modulation and throttles throughput; see RF baseline captured 2026-08-31 — root cause confirmed. A carrier data-cap throttle remains a live competing explanation for the multi-day blocks.
What was ruled out (with evidence)
| Suspect | Evidence | Verdict |
|---|---|---|
| DNS | Resolution 30 ms consistently from all three resolvers: Tailscale 100.100.100.100, router 192.168.1.1, 1.1.1.1. curl -w showed time_namelookup ~0.03 s while time_connect was the slow part | ❌ ruled out |
| IPv6 / Happy Eyeballs | en8 has only a link-local fe80:: address — no global, no ULA. ndp -rn shows no RA on en8. curl -6 fails in 2-32 ms rather than stalling 2-7 s | ❌ not a recurrence; the 07-28 fix holds |
| LAN | Mac on router port lan3 at 1000baseT full-duplex, 0 interface errors either side, 0% loss to gateway (0.72 ms avg over 150 packets). Counter-delta test confirmed the router forwarded 111 MB to lan3 cleanly during a transfer | ❌ ruled out |
| Bufferbloat | Latency under load rose only 26.4 ms → 27.4 ms. SQM is working | ❌ excellent, ruled out |
| Mac-specific cause | Initially suspected (mainframe hit 177 Mbit while the Mac got 39-55 Mbit in one window) but did not hold up — later samples had the Mac at 80-99 Mbit and the mainframe showing 5-7 s TCP SYN-retransmit stalls | ❌ that gap was sampling noise |
Lesson: never conclude from a single paired comparison in an episodic-fault environment
The “the Mac is slower than the mainframe” hypothesis came from one simultaneous sample. Minutes later the ordering reversed. In an environment where the underlying fault is bursty, a single A-vs-B comparison measures when you sampled, not which machine is broken.
The actual pattern
From 35 days of net-monitor hourly data on telep-mainframe (/home/levander/net-monitor/netlog.csv), classifying samples < 100 Mbit as degraded (normal is 180-265 Mbit):
Multi-day episodes separated by clean periods:
| Date | % samples degraded |
|---|---|
| Aug 8 | 52% |
| Aug 19 | 55% |
| Aug 20 | 100% |
| Aug 21 | 100% |
| Aug 22 | 82% |
| Aug 23-25 | 0-9% (clean) |
| Aug 26 | 52% |
| Aug 27-30 | 0-9% (clean) |
| Aug 31 (today) | 33% (3 of 9 samples) |
Aug 19-22 is the worst episode: 4 consecutive days.
- Worst individual samples: 3.0, 3.1, 4.6, 9.1, 11.2, 14.1, 14.2, 17.2 Mbit.
- Packet-loss events logged Aug 26 include three 100%-loss samples.
Live capture during the investigation
- One window: 26.7% ICMP loss from the Mac, and TCP handshakes of 5.1 s / 5.1 s / 7.1 s from the mainframe (SYN retransmission) to the same host.
- Minutes later: 150-packet runs from both machines, and hop-by-hop from the router — NE200
192.168.254.1, ISP hops10.153.240.158and84.1.85.225, and1.1.1.1— all showed 0% loss.
That is the confirmation that the fault is bursty/episodic, not constant.
Latency baseline stays healthy throughout (~25-40 ms) even during throughput collapse — a capacity/loss fault, not a latency fault.
Treat single net-monitor samples cautiously
The probe’s 20 MB sample size is small enough that TCP slow-start inflates variance. Individual rows are noisy; the day-level aggregates above are the reliable signal.
Reusable diagnostic signature
Episodic WAN degradation vs. the 2026-07-28 IPv6 fault
Episodic WAN degradation looks like: DNS fast · latency normal · bufferbloat fine · 0% loss on most samples · but throughput collapsing 5-50x intermittently and TCP connects occasionally taking 1 s / 2 s / 4 s / 5 s / 7 s (SYN retransmit backoff).
Distinguish from the IPv6 fault with one command:
curl -6 <dual-stack-site>stalls (2-7 s, then times out) → IPv6 fault.curl -6 <dual-stack-site>fails instantly (2-32 ms) → not IPv6; look upstream.
Secondary fault found (separate, unresolved)
Router port
10g-sfpis badly faulted — almost certainly a bad or empty/oscillating SFP moduleFound while auditing telep-router port counters. Unrelated to the internet slowness, but it should be removed or replaced.
- 3,792
carrier_changesvs 26 on the next-worst port (lan2) over 5 days uptime. - 252
rx_crc_errors— the only port on the router with any CRC errors at all. - Currently
operstate=down,speed=65535(invalid). - Only 414 KB rx in 5 days — it carries essentially no traffic.
- Flapping up/down on a ~1-second cycle in bursts; it is a member of
br-lan. - STP is disabled on
br-lan(stp_state=0,topology_change=0), so it is NOT causing bridge-wide topology churn.
Assessment: spams the system log heavily; unlikely to be causing the internet slowness (no traffic, no STP). Remove or replace the module.
Also noted: router port lan5 is linked at only 100 Mbit (1 GB rx, so something real is behind it) — worth identifying.
Environment facts confirmed
- Double-NAT intact: router WAN
192.168.254.2behind the NE200 at192.168.254.1— see Double-NAT. - The NE200 is a 5G NR cellular FWA CPE, not a WISP radio: SIM on Telekom HU NET, connected over NR5G with LTE anchors (B3 1800 MHz ×2 CA + B8 900 MHz) plus 5G N78 (3.5 GHz). CGNAT WAN IP
10.182.14.228, gateway10.182.14.229, plus a routed public IPv6/64(2a00:1110:138:1496::/64). - NE200 exposes HTTP on port 80 and telnet on port 23 at
192.168.254.1(TP-Link web UI). RF signal/SNR stats now retrieved (2026-08-31) — see RF baseline captured 2026-08-31 — root cause confirmed. - Router uptime 5.04 days, WAN uptime 6.8 h (
wan carrier_changes=9, so the WAN ethernet link has flapped 9 times). The NE200’s own PS (packet) session duration was 11 h 19 min at capture — the cellular data session is short-lived and re-establishes. - net-monitor measures
dl_mbpsvia a 20 MB Cloudflare download, hourly (only when minute < 5).
RF baseline captured 2026-08-31 — root cause confirmed
Read from the NE200 status page at 192.168.254.1 (web UI, admin login) at ~mid-day on a day running 33% degraded. This is the previously-missing RF evidence.
| Field | Value | Reading |
|---|---|---|
| Internet Status | Connected; ISP Telekom HU NET | — |
| SIM Card Status | prepared | — |
| Network Type | NR5G (5G SA/NSA) | — |
| Bands | B3, B3, B8, N78 | LTE anchors B3 1800 MHz ×2 CA + B8 900 MHz, plus 5G N78 3.5 GHz |
| Signal Strength | 75% | vendor composite (optimistic) |
| SS-RSRP | -102 dBm | POOR for 5G NR (the -100…-110 weak band) |
| SS-RSRQ | -12 dB | FAIR |
| SS-SINR | 9.5 dB | FAIR, low end — the throughput-limiting metric |
| Data used | 777.405 GB total | relevant to the cap hypothesis below |
| PS Session Duration | 0 d 11 h 19 m | short-lived; matches router WAN uptime ~6.8 h + a re-session |
| WAN IP (CGNAT) | 10.182.14.228 (v4) + 2a00:1110:138:1496::/64 (v6) | carrier-grade NAT on v4; routed public v6 |
| Gateway | 10.182.14.229 | — |
| Carrier DNS | 84.2.46.1, 84.2.44.1 (v4); 2001:4c48:2::1, 2001:4c48:1::1 (v6) | — |
Root cause confirmed: the cellular radio link is weak
At SINR ~9.5 dB the modem cannot sustain high-order modulation (256-QAM) and is forced down to 64/16-QAM, cutting throughput several-fold. When SINR dips further — interference, weather, or cell congestion at peak — throughput collapses to the 3-15 Mbit seen in netlog. Good 5G wants SINR > 15-20 dB. RSRP -102 dBm on N78 (3.5 GHz, poor propagation) is genuinely weak. This is a weak-signal cellular CPE, and it explains the hour-to-hour throughput jitter.
Competing hypothesis for the MULTI-DAY blocks: a carrier data-cap / fair-use throttle
Pure RF jitter is noisy hour-to-hour. But the netlog shape of Aug 19-22 (~100% degraded for 4 straight days) then fully clean Aug 23-25 looks more like a billing-cycle / data-cap throttle than RF noise. 777 GB used. If the Telekom plan throttles after a monthly threshold, antenna aiming will NOT fix the multi-day blocks — only the hour-to-hour jitter. This must be checked (action 2 below) before investing in an antenna.
For Agents — 5G NR signal reference bands (for future triage)
SS-RSRP: ≥-80 excellent · -80…-90 good · -90…-100 fair · -100…-110 poor · ≤-110 none. SS-RSRQ: ≥-10 good · -10…-15 fair · -15…-20 poor. SS-SINR: ≥20 excellent · 13…20 good · 0…13 fair · <0 bad. Captured baseline sits in the poor RSRP / fair RSRQ / low-fair SINR zone.
NE200 RF logger — deployed 2026-08-31
The scripted-login blocker (error 71014, flagged unsolved below) is now fully solved, and an automated RF logger is built, deployed, and running on telep-mainframe. Action 3 above is done.
The NE200 web login ("GDPR encrypt" scheme) fully reverse-engineered
- Login endpoint:
POST http://192.168.254.1/cgi_gdpr?9— the?9suffix is REQUIRED (omitting it → error71014). Server speaks HTTP/0.9 (usecurl --http0.9); body istext/plainofsign=<hex>\r\ndata=<b64>\r\n.- Public key from
POST /cgi/getParm→ returnsnn,ee=010001,seq. RSA-512.- Sign = RSA raw “nopadding” (each 64-char chunk zero-padded to 64 bytes big-endian,
e^ mod n, 128 hex/chunk, concatenated) of:
- login:
key=<K16>&iv=<IV16>&h=<md5(name+pwd)>&s=<seq+len(datab64)>- subsequent:
h=<md5(name+pwd)>&s=<seq+len(datab64)>- AES-128-CBC / PKCS7;
key&iv= 16 ASCII digits each, freshly generated per session.- Username is
user(the page forcesadminType="user"), NOTadmin. Hash =md5("user"+password).- Login plaintext (AES-encrypted, NO trailing CRLF):
{"data":{"UserName":"<b64(user)>","Passwd":"<b64(pwd)>","Action":"1","stack":"0,0,0,0,0,0","pstack":"0,0,0,0,0,0"},"operation":"cgi","oid":"/cgi/login"}. Missing theoperation:"cgi"field → errorcode71011. Success response decrypts to$.ret=0;.- After login,
GET /(with the session cookie) and scrape a 30-hex-char token; send it as theTokenIDheader on all subsequent queries (without it → HTTP 406).- Data queries: plaintext
{"data":{"stack":"0,0,0,0,0,0","pstack":"0,0,0,0,0,0"},"operation":"gl","oid":"<OID>"}WITH a trailing\r\n(omitting the CRLF → HTTP 406). Signed with the non-login sign form.- Signal OID =
DEV2_LTE_SERVING_CELL_INFO(opgl/getList) → per-carrier array with RSRP/RSRQ/SINR/CQI/RSSI/downlinkModType/uplinkModType/MCS per cell.DEV2_ADT_WAN= WAN byte counters;DEV2_CELL_INTF= IMEI/access-tech.
One web session at a time
The NE200 permits only one web session (“only one device can log in at a time; force logout?”). Any automated login force-logs-out an active web-UI session and vice-versa — so the logger logs out after each run and runs on a gentle interval.
The logger (deployed)
/home/levander/net-monitor/ne200_signal.pyon telep-mainframe — pure stdlib +curl+opensslCLI, no pip deps (the mainframe has nopycryptodome). Reads the password from/home/levander/net-monitor/.ne200_cred(chmod 600). Appends one row per run to/home/levander/net-monitor/rflog.csv.- Columns:
timestamp_iso,nr_rsrp,nr_rsrq,nr_sinr,nr_cqi,nr_dl_mod,nr_ul_mod,lte_rsrp,lte_rsrq,lte_snr,lte_rssi,lte_cqi,lte_dl_mod—nr_*= the serving 5G NR/N78 cell,lte_*= the serving LTE B3 anchor. - Cron:
*/15 * * * *(every 15 min); stderr →/home/levander/net-monitor/rflog.err. Correlaterflog.csvagainstnetlog.csv(throughput) by timestamp. - A Mac-side working client is preserved in the session scratchpad; the mainframe copy is the deployed one.
Caveat — trust the modulation/RSRP, treat
nr_sinras relative only
nr_sinrraw-field units are unclear (raw values 20-70 while the web UI shows SS-SINR ~5-9.5 dB) — treatnr_sinras a relative trend, not absolute dB. The unambiguous degradation indicators arenr_dl_mod(QPSK vs 256-QAM),nr_rsrp/nr_rsrq(dBm/dB), andnr_cqi(0-15).
Confirming evidence 2026-08-31 ~12:43-12:45 — real-time QPSK collapse
Live per-cell readings from the new logger, several runs ~90 s apart, on a 33%-degraded day — the radio was fluctuating fast and running degraded:
- Serving 5G NR (N78 3.5 GHz): RSRP swung -100 → -106 → -108 dBm across 3 runs in ~2 min; RSRQ -12/-13; CQI 10-12; downlink modulation = QPSK on every sample (the lowest order — vs 256-QAM when healthy); uplink dropped 256-QAM → 64-QAM on one sample.
- Serving LTE anchor (B3 1800 MHz): RSRP -95 to -98, RSRQ -12/-13, RSSI -83 to -95, CQI 6, downlink QPSK.
- Web-UI summary this window: SS-RSRP -102 dBm, SS-RSRQ -12 dB, SS-SINR dropped 9.5 → 5 dB vs the earlier reading, Signal 75%, data now 784.5 GB.
Direct real-time confirmation of the root cause
QPSK downlink means the modem is forced to the lowest modulation by the poor channel — a ~4x throughput loss vs 256-QAM before any other factor. The fast RSRP swings (6-8 dB in ~2 min) explain the episodic collapses. The persistent QPSK + low CQI — not DNS, LAN, or router — is why the internet is slow.
Next actions
- [BIGGEST LEVER] Physically aim/adjust the NE200’s antenna to raise RSRP/SINR — the highest-impact fix for a weak cellular CPE. Only applies if the unit has directional/adjustable antennas.
- Check the Telekom plan for a monthly data cap / fair-use throttle (777 GB used) and correlate the cap-reset date against the multi-day degraded blocks in netlog. This decides whether an antenna is even worth it.
- ✅ DONE — RF logger deployed.
ne200_signal.pynow scrapes per-cell RSRP/RSRQ/SINR/CQI/modulation from the NE200 every 15 min intorflog.csv; see NE200 RF logger — deployed 2026-08-31. This is the measurement that disambiguates RF vs. carrier: next episode, does modulation/RSRP crater (RF) or stay healthy while throughput drops (cap/congestion/carrier)? - Remove/replace the faulty SFP module in the
10g-sfpcage (unrelated hardware fault). - The netlog CSV remains strong evidence for a carrier/ISP complaint if the plan has no cap — Aug 20-21 were 100% degraded all day.
SOLVED — scripted NE200 login working, RF logger deployed
The NE200 web login (“GDPR encrypt”: AES-128-CBC + RSA-512) is fully reverse-engineered and a pure-stdlib logger runs every 15 min on telep-mainframe. The earlier error
71014was the missing?9suffix on/cgi_gdpr(plus usinguser, notadmin). Full protocol and the deployed logger: see NE200 RF logger — deployed 2026-08-31.
Related
- 2026-09-04-starlink-wan-migration-dish-telemetry — ⭐ the resolution: WAN migrated to Starlink Bypass; also the successor telemetry logger, the
100.64.0.0/10Tailscale collision, and the NE200 logger’s retirement - 2026-07-28-net-monitor — the logger whose 35-day record is the primary evidence here; it was installed for exactly this suspicion
- 2026-07-28-ipv6-slow-internet — the previous “slow internet” root cause; ruled out as a recurrence, fix still holding
- telep-router — the router;
10g-sfpfault and LAN/port evidence live here - Double-NAT — the NE200 fixed-wireless double-NAT topology
- telep-mainframe — host running net-monitor
- telep-mainframe-handover
- homelab
- LOG
- TOPICS