The decision-grade close of the 2026-07-10 loss work: a purged/embargoed OOS test (4,773 live filled orders, real CLOB labels) finds the overnight-+EV / daytime-−EV split is DIRECTIONAL-REAL but NOT statistically structural on 8 days — underpowered, period-noise-consistent, leaning on one blowout day. The better-supported lever is A0/adverse-fill mitigation (PM #1191: the toxic fills are the A0 first-touch, −5.70¢/sh @ 69.4% win; A1+ repriced fills are +1.30¢ @ 75.3% — removing A0 flips the 2-day book +EV), which is SEPARATE from the clock, not a proxy for it. Decision: keep the overnight re-arm as a +EV/low-risk but PROVISIONAL tilt (re-test after ~2–3 wk forward data); the real work is the A0 fill-policy — NOT the clock, NOT depth-aware sizing (NO-GO). Live-state change: Andras re-armed the 7 lanes overnight-only (LIVE 20:00–07:00 UTC, DRY daytime; old 13–16 CEST slice removed; levi/gyula OFF; wallets 1×), superseding the earlier-07-10 all-lanes-dried state.

Provenance — single-session, read-only, un-peer-reviewed

This is a single-session, read-only analysis by the main agent (2026-07-10). The OOS numbers are read-only duckdb→Postgres pulls against the live shared project mkofmdtdldxgmmolxxhc, scored on ACTUAL CLOB resolution labels (the market’s realized winner), over 4,773 live filled orders where realized_pnl_usd sign agreed 100% with the CLOB winner (label-truth-grade at scale — the 24 sign check now confirmed over thousands of fills). The A0 first-touch decomposition is PM’s (babylon pmv2 #1191) — cross-check its exact ¢/share against PM’s source before treating as canon. Not peer-reviewed, no teammate cross-check. Treat the direction (overnight-better / daytime-worse; A0 toxic) as well-evidenced and the exact statistics as this-session estimates — the whole point of the verdict is that the overnight split is underpowered on 8 days. The live-state change (overnight re-arm) is a real config action by Andras via @deploy (#1194/#1195).

For Agents — the verdict, the lever, the state change

1. OOS VERDICT (PROVISIONAL): the overnight-+EV / daytime-−EV split is DIRECTIONAL-REAL but NOT statistically structural. On 4,773 filled orders / real labels: overnight (20:00–06:00 UTC) +560 / −2.15¢/sh; the gap is within-lane (+0.0352/fire within-lane vs +0.0372/fire aggregate — not lane-mix). BUT P(overnight>daytime)=0.949 with the difference CI including 0, and P(overnight>0)=0.875 — so “overnight is positive” is NOT significant; the robust leg is “daytime is worse.” Walk-forward met the bar in only 5/9 day-folds & 3/6 contiguous folds, WITH sign reversals on 07-03/04/08 (the time-of-day NO-GO flip signature). 67% of the daytime loss is the single 07-09 US-open blowout; dropping it collapses P 0.949→0.883. ⇒ UNDERPOWERED / period-noise-consistent. Overnight’s edge is largely that the overnight window runs the known-winner lanes (sol_up / eth_up / xrp_down / doge_down / eth_down, ~5 lanes).

  • 2. A0 IS THE REAL LEVER (mechanism CONFIRMED, PM #1191): the toxic fills are the A0 first-touch−5.70¢/sh @ 69.4% win (−EV); the A1+ repriced fills are +1.30¢ @ 75.3%; removing A0 flips the 2-day book +EV. Crucially, A0-toxicity is SEPARATE from time-of-day — daytime A0-share 82.8% vs overnight 80.1% (only +2.7pp), so the overnight gate is NOT a proxy for “avoid A0.” This sharpens the FILL adverse-selection verdict (where the toxicity lives in the fill sequence: the instant first-touch fill) and corrects/extends the FAK verdict (its “repriced fills win 87–88%, no adverse selection” is the A1+ subset — clean; the toxicity is the A0 blend it averaged over, on the optimistic label).
  • 3. DECISION: keep the overnight re-arm — +EV/low-risk but PROVISIONAL, NOT a validated structural rule — and re-test after ~2–3 wk forward data. The better-supported lever = A0 / adverse-fill mitigation (a fill-policy that avoids the toxic first-touch), NOT the clock, and NOT depth-aware sizing (NO-GO).
  • 4. LIVE-STATE CHANGE (2026-07-10): Andras re-armed the 7 lanes overnight-only (@deploy, babylon pmv2 1195): scheduler restored, LIVE 20:00–07:00 UTC (22:00–09:00 CEST), DRY daytime 07:00–20:00 UTC; the old 13–16 CEST daytime slice explicitly REMOVED (−EV); levi/gyula stay OFF; wallets 1×. This supersedes the earlier-07-10 all-lanes-dried state. PM’s 45fa0eb telemetry-fix (best_ask_at_snap was a req.limit proxy at attempt 0) rolls alongside so fills are measured real going forward.
  • Memory pointers: this note is the canonical overnight-edge-provisional-a0-is-the-lever anchor; it sharpens loss-is-adverse-selection-at-fill (A0 = where the toxic fill is) and sits beside depth-aware-sizing-nogo (the shelved sizing lever).

The two questions and the one answer — clock vs fill

flowchart TD
    Q["<b>Recent loss = FILL adverse selection</b><br/>(prior verdict, real labels)<br/><i>which lever fixes it?</i>"] --> C{"<b>Lever A: the CLOCK</b><br/>arm overnight, dry daytime?"}
    Q --> F{"<b>Lever B: the FILL</b><br/>avoid the toxic first-touch?"}
    C -->|"overnight +1.57¢ / daytime −2.15¢<br/>within-lane (not lane-mix)<br/>BUT P(on&gt;day)=0.949 CI∋0,<br/>P(on&gt;0)=0.875, folds 5/9<br/>reversals 07-03/04/08,<br/>67% of loss = 07-09 alone"| CA["<b>PROVISIONAL</b><br/>directional-real, NOT structural<br/>underpowered on 8 days<br/>robust leg = 'daytime worse'"]
    F -->|"A0 first-touch −5.70¢ @ 69.4%<br/>A1+ repriced +1.30¢ @ 75.3%<br/>remove A0 ⇒ 2d book +EV<br/>A0-share 80.1% vs 82.8%<br/>(clock is NOT a proxy)"| FA["<b>THE REAL LEVER</b><br/>fill-policy: avoid A0<br/>separate from time-of-day"]
    CA --> D["<b>DECISION</b><br/>re-arm overnight-only = low-risk tilt,<br/>PROVISIONAL (re-test ~2–3 wk)<br/><b>+ do the A0 fill-policy work</b><br/>(NOT clock, NOT depth-sizing)"]
    FA --> D
    style Q fill:#264653,stroke:#2a9d8f,color:#fff
    style FA fill:#264653,stroke:#2a9d8f,color:#fff
    style CA fill:#3d2020,stroke:#c1666b,color:#fff
    style D fill:#3d2020,stroke:#c1666b,color:#fff

The OOS verdict — overnight/daytime is directional-real, not structural

Purged/embargoed OOS over 4,773 live filled orders scored on ACTUAL CLOB resolution labels (not resolved_up), with realized_pnl_usd sign matching the CLOB winner 100% of the time:

Window (UTC)RealizedPer-shareRead
Overnight 20:00–06:00+$473+1.57¢/shpositive, but see the significance test
Daytime 07:00–19:00−$560−2.15¢/shthe loss window — the robust leg

The gap is NOT a lane-composition artifact: the overnight-minus-daytime edge is +0.0372/fire aggregate and +0.0352/fire within-lane (holding lane mix fixed) — essentially identical, so it is a genuine within-lane time effect, not “overnight just happens to run the good lanes.” (That said, the overnight window does run the known-winner lanes — sol_up / eth_up / xrp_down / doge_down / eth_down, ~5 lanes — which is why arming it is low-risk regardless of whether the clock is the cause.)

Why it is underpowered / provisional, not structural

The direction is real; the structure is not established on 8 days:

  • The positive leg is not significant. P(overnight > daytime) = 0.949 but the difference CI includes 0, and P(overnight > 0) = 0.875 — so “overnight is +EV on its own” does not clear a conventional bar. The statistically robust finding is the negative leg: daytime is worse.
  • Walk-forward is split, with the flip signature. Only 5/9 day-folds and 3/6 contiguous folds met the bar, and there are sign reversals on 07-03, 07-04, 07-08 — the exact time-of-day NO-GO “flip between periods” fingerprint that flagged hour-of-day P&L as noise last time.
  • One day carries the effect. 67% of the daytime loss is the single 07-09 US-open blowout (the 4-wallet-amplified −$326 day). Dropping 07-09 collapses P(overnight > daytime) 0.949 → 0.883 — the whole effect leans on one blowout.

Reconciling with the 2026-07-06 time-of-day NO-GO

This does NOT overturn the time-of-day NO-GO — it confirms its caution. The reversals (07-03/04/08) and the one-day dependence are the same “period-noise, efficiently-priced overlay” signature. The difference is operational, not statistical: the overnight re-arm is a risk-management tilt on the known-winner lanes during their observed-profitable window, explicitly labelled provisional, not a claim that the clock is a validated structural edge. The durable lesson stands — do not treat time-of-day as alpha on top of the frozen [60,150)s regime.

A0 is the real lever — the toxic fill is the first-touch

PM’s decomposition (babylon pmv2 #1191) splits the filled book by fill attempt index, and the toxicity is concentrated in A0 (the first-touch fill, at the emit-ask):

Fill classPer-shareWin rateRead
A0 (first-touch, attempt 0)−5.70¢/sh69.4%−EV — the toxic fills. Cheap-and-fills-instantly = the market knows it loses (classic adverse selection).
A1+ (repriced retry)+1.30¢/sh75.3%+EV — clean. The ask moved but we still caught it within the ceiling.
remove A0the 2-day book flips +EV.

This is the sharp, actionable form of the FILL adverse-selection verdict: the +EV lives in the unfilled/repriced tail, and the losers are the ones that fill instantly at the original ask (A0). A fill-policy that avoids or de-weights the A0 first-touch targets the loss directly.

For Agents — A0 is SEPARATE from the clock (not a proxy)

The overnight gate and the A0 lever are independent mechanisms, not two views of one thing: A0-share is nearly flat across the day — daytime 82.8% vs overnight 80.1% (only +2.7pp). So arming overnight does not meaningfully avoid A0, and avoiding A0 is not a time-of-day effect. The overnight re-arm and the A0 fill-policy are additive, orthogonal levers — and A0 is the better-supported of the two (a mechanism with a sign-flipping counterfactual, vs an underpowered clock split).

How A0 corrects/extends the FAK order-type verdict

The 2026-07-09 FAK verdict measured “repriced fills win 87–88% ≈ population, no adverse selection.” A0 explains the tension the loss investigation already flagged:

  • The FAK verdict’s clean “repriced fills” are the A1+ subset (+1.30¢ / 75.3% on real labels) — it was right about those.
  • What it averaged over (on the optimistic on-chain label) is the A0 first-touch blend, which on real labels is −5.70¢ / 69.4%.
  • The order-type conclusion still stands — do NOT switch to uncapped market/FOK or resting GTC (strictly worse). What is corrected is the characterisation of the blended filled population as “no adverse selection”: the adverse selection is real and concentrated in A0. The lever is a fill-policy (which attempts to accept), not an order-type swap.

Decision — provisional overnight tilt + A0 is where the work is

The two-part decision

  1. Keep the overnight re-arm as a +EV / low-risk but PROVISIONAL operational tilt — it arms the known-winner lanes during their observed-profitable window and dries the robustly-−EV daytime window. It is NOT a validated structural rule; re-test after ~2–3 weeks of forward data (the current 8-day sample is underpowered and one-day-dependent).
  2. Do the A0 / adverse-fill work — a fill-policy that avoids the toxic first-touch is the better-supported lever (mechanism-confirmed, sign-flipping counterfactual, orthogonal to the clock). This is NOT the clock and NOT depth-aware sizing (shelved — a size-up, not a de-risk).

Why the overnight tilt is reasonable despite being underpowered: the downside is bounded (the overnight lanes are the documented winners; daytime is where the money bleeds), it is fully reversible, and it buys ~2–3 weeks of live-forward data to actually power the test — while the real edge-recovery work (A0) proceeds in parallel and independently.

Live-state change — overnight-only re-arm (2026-07-10)

Andras re-armed the 7 previously-live lanes overnight-only (task to @deploy, babylon pmv2 #1194 / #1195), superseding the earlier-07-10 all-lanes-dried state:

AspectNew state (2026-07-10 re-arm)
SchedulerRESTORED (was removed when all lanes were dried)
LIVE window20:00–07:00 UTC (22:00–09:00 CEST) — the 7 lanes place real orders
DRY window07:00–20:00 UTC — daytime dry-shadow only
Old 13–16 CEST daytime sliceREMOVED (explicitly — it was −EV)
Lanesthe 7 previously-live (btc-down, eth-up/down, sol-up, xrp-down, doge-up/down); overnight +EV concentrates in the ~5 known winners (sol_up / eth_up / xrp_down / doge_down / eth_down)
levi / gyula (wallets 3 & 4)stay OFF (the 07-09 exposure amplifier — not re-added)
Wallet size (no size-up — consistent with the depth-sizing NO-GO)

Telemetry fix rolls alongside — fills measured real going forward

PM’s 45fa0eb corrects a telemetry defect where best_ask_at_snap was a req.limit proxy at attempt 0 (i.e. the recorded “ask at snapshot” was actually our own limit, not the book) — which would have muddied any A0/attempt-index and slippage analysis. Rolling it with the re-arm means the overnight-forward fills (the data that will power the ~2–3 wk re-test) are measured against the real book, not a self-referential proxy. Any A0 re-derivation on pre-45fa0eb data inherits this caveat.

Artifacts

For Agents — where the numbers live

  • OOS overnight/daytime: read-only duckdb→Postgres pulls over the 4,773 filled-order population on real CLOB labels (shared project mkofmdtdldxgmmolxxhc); bootstrap P-values (overnight>daytime 0.949, overnight>0 0.875), day-fold + contiguous-fold walk-forward, 07-09-drop sensitivity (0.949→0.883).
  • A0 decomposition: babylon pmv2 #1191 (PM-owned) — A0 −5.70¢/69.4%, A1+ +1.30¢/75.3%, remove-A0 flips +EV, A0-share 80.1% overnight vs 82.8% daytime.
  • Telemetry fix: PM commit 45fa0eb (best_ask_at_snap req.limit-proxy correction).
  • Deploy: babylon pmv2 #1194 / #1195 (overnight-only re-arm: scheduler + 20:00–07:00 UTC LIVE / daytime DRY, 13–16 CEST slice removed, levi/gyula OFF, 1×).