Read-only investigation into the crypto_5m_algo (crypto_shortterm_latency_test) losses over 2026-07-08 → 07-10. The headline: there is no code regression and no signal failure — the loss is a 3-day sub-breakeven slide (not a one-day blowout), the money-at-risk edge is razor-thin (~1–3pt over breakeven), and the on-chain resolved_up label our monitoring is built on overstates the traded win rate by ~9–14pt on the filled subset, so every dashboard reads ~10pt healthier than the cash.

Provenance — un-peer-reviewed

This is a single-session, read-only analysis by the main agent (2026-07-10), all numbers from read-only duckdb→Postgres queries against the live shared project mkofmdtdldxgmmolxxhc (ATTACH … TYPE postgres, READ_ONLY). Not peer-reviewed, no teammate cross-check, no code changed. The one remaining hard confirmation (a CLOB tokens[].winner cross-check, see Open questions) was not run. Treat the direction as strongly evidenced and the exact magnitudes as this-session estimates.

For Agents — the seven findings

  1. No code regression. The 07-09 emit commits (E1 nats.rs Pub-Ts header, E4 emit.rs claim_ms/pub_lag_ms) are telemetry-only + dedup-safe (Nats-Msg-Id still = order.id, unit-tested); E2/E3 are analysis scripts, not in the trade path.
  2. Model/signal healthy. On-chain signal-win 85–88% every day incl. 07-09 (85.7%) and hour-by-hour through US-open; fire volume normal (~880–920/day); entry asks normal (~0.72). The loss is NOT a model or signalling failure.
  3. A 3-day slide, not a blowout. Live realized_pnl_usd (CLOB truth): 07-05 +15, 07-07 +120, 07-09 170 (partial). 07-08 was already negative — contradicts the repo HANDOVER.md “+448 on 07-09”.
  4. KEY: resolved_up overstates the traded win rate ~9–14pt on the FILLED subset. 07-08+07-09 fills: 111 on-chain-WIN-but-real-LOSS vs 13 the reverse, disagreements are near-ties (median 3.5 bps from strike) = asymmetric optimistic label noise worth −$1,011 over the 2 days. Real fill-win on loss days 64–72% vs on-chain 78–86%; breakeven ~73%. Monitoring (signal-win dashboards, CRM toxicity tab, break-even bookkeeping) reads ~10pt healthier than the money at risk.
  5. Loss fingerprint = side-specific INTRADAY adverse selection. On loss days one side’s real fill-win collapses below breakeven (Up 07-08 65% & 07-10 46%; Down 07-09 64%) while the other holds ~74–77%; profitable days both sides ~74–76%. Not a daily trend (market up-rate balanced ~48–50%) and not a chop regime (near-tie fraction did NOT track losses). ~50% of orders FAK-no-match daily = separate conversion loss.
  6. 07-09 amplifier. Wallets 3 & 4 (levi/gyula) were live only 07-08 20:02 → 07-09 12:58 UTC (152/154 fills each) — exactly spanning the blowout, then pulled. 07-09 ran on 4 wallets, doubling exposure into a correlated down-side loss.
  7. No live daily-loss circuit breaker. Caps are dead code (known/deferred) → the 3-day slide ran unchecked.

1. No code regression

The 07-09 emit changes are not in the money-critical decision path and do not change what gets fired or how it dedups:

  • E1 — nats.rs Pub-Ts header and E4 — emit.rs claim_ms / pub_lag_ms are telemetry-only. They add timing fields to the signal-emitted log; they do not touch the fire decision, the limit price, or the Directive contract.
  • Dedup is intact: Nats-Msg-Id still = order.id (= {source}|{window}), deterministic per window, and unit-tested. No double-fire path was introduced.
  • E2 / E3 are analysis scripts (the fanout query pack and the weekly depth guardrail), run offline against telemetry — not in the trade path at all.

The publish-first + DLQ latency work (crypto-shortterm-emit-publish-first-latency-2026-07-09) is flag-gated (EMIT_PUBLISH_FIRST, default OFF = byte-identical legacy path) and was rolled DARK — it was not live-active during the slide either. So the losses cannot be pinned on a 07-09 code change.

2. The model and signal are healthy

Whatever is bleeding money, it is not the predictor:

  • On-chain signal-win is stable 85–88% every day, including 07-09 at 85.7%, and holds hour-by-hour through the US-open window that the repo narrative blamed.
  • Fire volume is normal (~880–920 fires/day) — the selector is not misfiring or over-firing.
  • Entry asks are normal (~0.72) — we are not suddenly paying up.

This isolates the problem to execution/labeling on the filled subset, not to fair-value or fire selection.

3. The real picture — a 3-day sub-breakeven slide, not a one-day blowout

Using pmv2_autotrade_orders.realized_pnl_usd (mode='live', our lanes) — the CLOB-settlement truth established in crypto-shortterm-lane-state-adverse-fill-2026-07-05 and crypto-shortterm-pnl-attribution-corrections-2026-07-03:

DateLive realized P&L (our lanes)Shape
2026-07-05+$164profitable
2026-07-06+$15~breakeven
2026-07-07+$83profitable
2026-07-08−$120slide begins
2026-07-09−$3264-wallet US-open blowout
2026-07-10−$170 (partial)still bleeding on 2 wallets

Contradicts the repo HANDOVER.md

The 2026-07-10 HANDOVER.md “NUMBERS” section frames the week as “+448 US-open adverse-fill blowout.” On the CLOB-truth realized_pnl_usd series, 07-08 was already −$120 — the bleed started a day earlier and continues on 07-10 after wallets 3/4 were pulled. This is a 3-day slide, not a self-contained one-day event. (Exact daily figures differ from the HANDOVER by measurement basis — on-chain cash Σredeem+Σsell−Σbuy vs realized_pnl_usd, and lane subset — but the shape correction stands: the losing regime is not confined to 07-09, so a “one-off” mental model will mis-size the risk.)

4. KEY finding — resolved_up overstates the traded win rate ~9–14pt on the FILLED subset

This is the most consequential result and the reason the team’s read of the situation is too optimistic.

Our on-chain settlement proxy outcomes.resolved_up (settle ≥ K from Polygon Chainlink) is known to over-book near-tie wins vs the market’s actual resolution (crypto-shortterm-pnl-attribution-corrections-2026-07-03 found ~4–7pt per-bin inflation). On the filled, adversely-selected subset the gap is bigger — ~9–14pt — because the fills we win/lose are concentrated exactly in the near-ties where the two labels disagree:

  • For 07-08 + 07-09 filled orders: 111 rows were on-chain-WIN but real-P&L-LOSS, versus only 13 the reverse — a strongly asymmetric disagreement (optimistic label noise, not symmetric jitter).
  • The disagreements are near-ties: median 3.5 bps from strike — precisely the sub-5bps band where on-chain settle ≥ K and the market’s Chainlink Data Streams resolution diverge.
  • Dollar cost of the asymmetry: ≈ −$1,011 over those 2 days that on-chain accounting does not see as losses.
  • Win-rate gap on loss days: real fill-win 64–72% vs on-chain 78–86%; breakeven is ~73%. So on the honest label the loss days were below breakeven, while the on-chain label showed them comfortably above it.

Implication — the monitoring reads ~10pt healthier than the money at risk

Everything built on resolved_up — the signal-win dashboards (crypto-shortterm-perf-dashboard-2026-07-05), the CRM “toxicity” tab signal_win − fill_win (crypto-shortterm-empire-ui-crm-topology-2026-07-06), and break-even margin bookkeeping — is ~10pt optimistic on the filled subset. The real edge is razor-thin (~1–3pt over breakeven), consistent with the documented 73.7% WR vs ~73% breakeven (crypto-shortterm-pnl-onchain-cash-truth-2026-07-06). A thin edge plus a ~10pt optimistic gauge is exactly how a multi-day sub-breakeven slide can run while the dashboards look fine.

Tension with the FAK order-type verdict (2026-07-09)

crypto-shortterm-order-type-fak-verdict-2026-07-09 concluded “realized fills show no measured adverse selection — repriced fills win 87–88% ≈ population” and used that to keep FAK marketable limits unchanged. That 87–88% is on the on-chain (optimistic) label. On the traded realized_pnl_usd label, filled win on loss days is ~10pt lower and side-specifically below breakeven (finding 5). The order-type verdict is not necessarily wrong — but its “fills aren’t toxic” evidence base should be re-measured on realized_pnl_usd, split by side and day, before it is treated as settled. This note is the systematic, traded-label version of the open E2 question that verdict itself proposed.

5. Loss fingerprint — side-specific intraday adverse selection

The loss is not spread evenly; on each loss day one side’s real fill-win collapses below breakeven while the other holds:

DateUp real fill-winDown real fill-winCollapsed side
2026-07-0865%~74–77%Up
2026-07-09~74–77%64%Down
2026-07-1046%~74–77%Up
profitable days~74–76%~74–76%none (both hold)

Breakeven is ~73%, so a side at 46–65% is a structural money-loser that day. The adverse selection is visible even on the optimistic label: on the losing side, filled opt-win runs 11–22pt below all-signal opt-win (today’s Up: filled 56% vs all-Up-signal 78%) — i.e. we systematically fill the losers and miss the winners on that side, the btc-up adverse-fill signature generalised.

Why this is INTRADAY counter-trend, not a chop regime

Two negative controls rule out the obvious “bad day” explanations:

  • Daily market up-rate was balanced ~48–50% on the loss days — there was no daily directional trend for one side to fight; the collapse is intraday, not a whole-day drift.
  • Near-tie fraction did NOT track losses — 07-05 had the most near-ties (42.9%) yet was the most profitable day. So it is not a “choppy near-tie regime” story either.

The signature is therefore intraday counter-trend adverse fill: within the day, when one side is being pushed, our cheap-ask fills on that side land on the wrong outcome while the ask on the winning side runs away. The exact intraday trigger is the main open question below.

Separately, ~50% of orders FAK-no-match every day — a conversion loss (missed fills at validated prices), distinct from the realized loss on the fills we do get. This is the latency-conversion story from crypto-shortterm-fak-nomatch-root-cause-2026-07-02 and crypto-shortterm-order-type-fak-verdict-2026-07-09, and it is stable across profitable and losing days — so it is a capacity ceiling, not the cause of the slide.

6. The 07-09 amplifier — wallets 3 & 4 doubled exposure into the blowout

The 07-09 day was the worst because exposure was temporarily doubled right through it:

  • Wallets 3 & 4 (levi/gyula) were live only 2026-07-08 20:02 → 2026-07-09 12:58 UTC152 / 154 fills each in that window, and essentially none outside it.
  • That interval exactly spans the US-open blowout. On 07-09 the book ran on 4 wallets trading the same windows/sides, so a correlated down-side loss (finding 5) was booked at ~2× size.
  • The wallets were then pulled back to 2 (w1/w2) — but 07-10 still bleeds (−$170 partial), which is why the amplifier is a multiplier on the loss, not its cause. Removing it did not stop the slide.

This is the cross-asset / cross-wallet within-window correlation risk (the risk unit is the window, not the fill) meeting a live-loss regime with no size backstop. See crypto-shortterm-per-wallet-size-scale-2026-07-06 for the wallet/size_scale mechanics and the ended 1×-vs-2× A/B.

7. No live daily-loss circuit breaker

There is no live daily-loss kill switch in the path. The total_exposure_cap / daily_spend_cap / max_open / max_copies caps enforce nowhere — they are dead code (the reserve() cap block required Mode::Live but its only caller passed Mode::Dry), decided advisory / deferred by Andras. See pmv2-live-caps-are-advisory-not-blocking-2026-07-09. The real live bound is EMIT_MAX_FIRES × per-wallet size, floored by wallet pUSD — none of which reacts to a losing day. Consequently the 3-day slide ran unchecked; nothing in the system was positioned to halt it as it developed.

Open questions

Two items are not resolved by this session and gate turning any of this into an intervention:

  1. The exact intraday trigger of the side-specific fill collapse. Needs an hourly price-move vs fill-win-by-side cut: does the losing side’s real fill-win collapse specifically when the underlying is moving against that side intraday (counter-trend adverse fill), and by how much lead/lag? Until this is pinned, the mechanism is characterised but not causally nailed.
  2. A CLOB tokens[].winner cross-check on the disagreeing fills to 100% confirm the ~9–14pt gap is label noise (strongly indicated: the 111-vs-13 asymmetry at a 3.5 bps median near-tie is the exact resolved_up-over-books-near-ties signature) vs any residual PM realized-booking lag (the booking-lag class). The 24/24 sign-match spot-check in crypto-shortterm-lane-state-adverse-fill-2026-07-05 argues for label noise, but that was a small sample on a different day.

What this changes / doesn’t change

  • Do not “fix the model.” It is healthy; a recalibration would be chasing the wrong layer (and the frozen calibration.json is validated-frozen by design).
  • Re-base the health gauges on realized_pnl_usd, split by side and day. The signal-win / toxicity dashboards are ~10pt optimistic on the filled subset; a side-and-day-resolved traded-win view is the honest instrument.
  • Treat the edge as razor-thin. ~1–3pt over breakeven means side-specific intraday adverse fill can flip a day negative with no model change — sizing and any circuit-breaker decision should assume that.
  • The FAK order-type verdict’s “no adverse selection” needs re-measuring on the traded label before it is relied on (see the tension callout in finding 4).