A small, a-priori regime sweep over the same K=6 purged/embargoed walk-forward as crypto-shortterm-phase0-walkforward. Selecting the high-conviction near-expiry regime raises t from 1.46→4.61 (5/5 OOS folds positive) — the sign/persistence of the edge is statistically robust OOS and clears the Bonferroni bar. But the magnitude is almost certainly a mirage: ~7%/trade implies the 1s top-of-book is not the fillable price. T19 is a CONDITIONAL pass — direction confirmed, fill realism still unverified.

For Agents

Status after this note: T19 = CONDITIONAL pass. Selectivity sweep (bin selectivity.rs, commit 2d52628) on the K=6 purged/embargoed walk-forward found a robust OOS signal in one regime — τ∈[60,150)s & conviction |p_cal−0.5|≥0.30: t=4.61, 5/5 folds positive, clears both the K=7 (t≥2.69) and broad N≈17 (t≥2.97) Bonferroni bars. Edge lives exclusively in the last 60–150s on clearly-ITM/OTM contracts; all τ∈[150,300)s cells fail (t<1.8). Caveat (key): magnitude (~5.5¢ net / ~80¢ contract ≈ ~7%/trade) is implausibly high for real binary fills ⇒ recorded 1s top-of-book ≠ fillable price (thin-book/snapshot optimism — the project’s recurring #1 risk). Cannot be resolved offline: observations don’t record book DEPTH. Do NOT promote to Phase 1 / build execution yet. Next: record book sizes → re-run depth-aware → only then a tiny real-fill test in the τ60–150s/c≥0.30 cell. Snapshot paper-lane still rejected (1s books can’t test true latency).

Setup

  • Binary: crates/analysis/src/bin/selectivity.rs
  • Commit: 2d52628
  • Report: docs/superpowers/phase0-selectivity-report.md
  • Walk-forward: identical K=6 sequential purged/embargoed folds as crypto-shortterm-phase0-walkforward (train on all prior, test on next block; purge/embargo at boundaries; τ≥60s base filter; isotonic calibration fit-on-train)
  • Anti-data-snooping design: a small, fixed, economically-motivated grid — not a search. 7 a-priori regime cells:
    • conviction |p_cal − 0.5| ∈ {0.10, 0.20, 0.30} (distance of the calibrated prob from a coin flip — how clearly ITM/OTM)
    • τ-band ∈ {[60,150)s, [150,300)s}
    • plus the all-τ≥60s baseline
  • Significance bar: Bonferroni-corrected. K=7 cells → t≥2.69; broad family N≈17 → t≥2.97. A cell must clear both to count.

Out-of-Sample Results (ranked)

RankRegimet-statOOS folds +TradesWindowsMean/windowSharpeVerdict
1τ∈[60,150)s, c≥0.304.615/5~34k8342.210.16PASS (both bars)
2τ∈[60,150)s, c≥0.204.10PASS
3τ∈[60,150)s, c≥0.104.04PASS
baseline all-τ≥60s2.51FAIL (< 2.69)
τ∈[150,300)s (all cells)< 1.8FAIL

Reading: the edge lives exclusively in the last 60–150s on clearly-ITM/OTM contracts. Every τ∈[150,300)s cell fails. The undifferentiated baseline also fails — selectivity is doing real work.

Why the t-stat jumped (1.46 → 4.61): pure variance reduction, not a bigger edge. Mean per-window PnL actually FELL (4.15 → 2.21/window) once selective — but σ dropped far more, so t rose. This is exactly the signature of a real-but-noisy effect being isolated: the sign and persistence survive OOS even as the headline magnitude shrinks.

The Caveat — magnitude is a mirage

Robust direction, but the dollar magnitude is almost certainly not fillable

The winning cell nets ~5.5¢ per trade on ~80¢ contracts ≈ ~7% per trade — implausibly high for real binary-contract fills. This strongly implies the recorded 1s top-of-book is NOT the price you’d actually fill at (thin-book / snapshot optimism — the project’s recurring #1 risk, see crypto-shortterm-phase0-codebase-complete’s 0.98-spread / empty-book WATCH and the ~$150–250 touch-depth finding). It cannot be resolved offline: the observations rows record top-of-book prices but not book DEPTH (sizes), so every PnL figure here implicitly assumes an infinite fill at top-of-book. The signal’s existence is robust; its realizable size is unknown.

Implication for T19 / phasing

T19 is a CONDITIONAL pass: a statistically robust signal direction was found in a specific regime, but live fill realism is unverified. This is stronger than the prior crypto-shortterm-phase0-walkforward “fragile/likely-overfit” verdict (the selectivity isolates a real OOS effect) but is NOT a GO.

Next steps:

  1. Add book-size recording to the recorder + the observations schema (small change → new deploy pin, coordinate with @deploy) → re-run selectivity depth-aware to watch the edge decay under realistic sizing.
  2. Only if it survives depth-aware sizing: a small real-fill test on a funded account in exactly the τ∈[60,150)s / c≥0.30 regime.
  3. The snapshot paper-lane idea remains rejected — 1s book snapshots can’t test true latency, which is the whole question for a latency-style edge.