UPDATE 2026-06-22 — pmv2 reorg + this lane is now weather_as in its own repo

This research belongs to the Asian-weather lane, now carved out as the weather_as producer: source weather_as, subject pmv2.order.weather_as.entry, in the PRIVATE repo pmv2-weather-as-algo at /coding/pmv2/as_weather_algo (committed e9415dc). The research record itself stays in the renamed US repo pmv2-weather-us-algo at /coding/pmv2/us_weather_algo (paths below resolve there, e.g. research/gate/.../coding/pmv2/us_weather_algo/research/gate/...). Full current state, the multi-producer / single-executor (position_manager) architecture, and the pmv2-contracts envelope: pmv2-reorg-and-weather-as-2026-06-22.

The discharge of caveat #1 of the shadow-rehearsal CONDITIONAL-GO (weather-asia-shadow-rehearsal-2026-06-19): “the gate selects on convergence, not win — re-score against a true win/+EV label before sizing capital.” Re-scoring on P(WIN) instead of P(convergence) collapses the headline EV from ~+0.14/**, and the win-label model loses to a per-city win-rate lookup. Verdict: CONDITIONAL-GO downgraded to LIVE-MEASUREMENT-ONLY — no capital, no Rust, until a live loop shows a repeatable edge over an ungated / per-city baseline. Full report: research/reports/weather-gate-winlabel-2026-06-19.md (now under /coding/pmv2/us_weather_algo/).

The convergence EV was an optimistic ceiling — confirmed and quantified

The shadow rehearsal’s +0.14/ vs win gate **+0.0604/`. The model-attributable edge is +0.06/$ and is not separable from trivial baselines. This does not kill the live dry-run (it measures real fillability, which no backtest can), but it shrinks the prize and makes the thin-cohort problem more binding.

For Agents

Five load-bearing facts:

  1. The honest win-label EV is +0.0604/$ at matched 8.5% admit (+0.0594 at its own EV-optimal 7.4% admit), zero-cost central. Positive, but barely.
  2. It loses to trivial baselines. At matched admit: per-city win-rate gate +0.0939 (win-gate loses), ex-ante-only model (no price/book/gap, AUC 0.54) +0.0772 (loses), ungated +0.0553 (ties), price-alone +0.0253 (win-gate beats only this). Signature of a disguised price-picker.
  3. The market price IS the win probability. entry_price alone has OOS AUC 0.7136higher than the win-model’s 0.6838. The model re-expresses price worse than price itself.
  4. Win base rate ≈ 0.83 (these are favorites, mean entry ≈0.79). The “skill” is separating the ~17% of favorites that lose — on a population the market already sorted by price.
  5. Verdict downgrade: CONDITIONAL-GO → LIVE-MEASUREMENT-ONLY. Keep the live dry-run as a cheap measurement of real fillability; treat the strategy as “favorite-buying with a cost cushion,” not a weather-timing model. No capital, no Rust productionization.

The answer

Selecting on P(WIN) rather than P(convergence), net-of-cost EV is +0.0604/ at its own EV-optimal threshold (7.4% admit), zero-cost central. Positive — but it does not survive the honest criterion in any way that matters: the win-label model adds approximately nothing over the raw market price, and is beaten by a per-city win-rate lookup.

Head-to-head, identical honest basis (resolved-only OOS, true Gamma payoffs, matched admit)

At matched admit ~8.5%, zero-costEV/$beats win-gate?
convergence gate (honest same-basis)+0.1826win-gate loses
per-city historical win-rate gate+0.0939win-gate loses
ex-ante-only model (no price/book/gap, AUC 0.54)+0.0772win-gate loses
win-label gate (this exercise)+0.0604
ungated (admit everything)+0.0553ties (+0.005)
price-alone gate+0.0253win-gate beats

Discrimination (OOS AUC on the WIN label)

Model / signalOOS AUC (win)Note
entry_price ALONE0.7136the market-implied win prob; beats everything below
win-label model — logistic (selected)0.6838loses to price-alone
win-label model — GBT0.6762
convergence gate re-scored on win0.6071trained on convergence, weak on win
ex-ante-only (city/hi_lo/unit/session/dow/local_hour)0.5393no price/book/gap
per-city historical win-rate0.4945~coin flip on win

Effect on the shadow-rehearsal CONDITIONAL-GO

This weakens the live-shadow GO and, on the strict reading of its own pre-conditions, kills the basis for the Rust-productionization track (pre-condition 2 — “re-score against a true win/+EV label” — is now met and the gate fails it):

  • The +0.14/$ headline was a convergence-selected, proxy-payoff ceiling — confirmed, inflation now quantified (the prior JSON’s −0.0798 ev_delta was itself a fabricated comparison against a hard-coded 0.14 literal on the full population with converged-as-win proxy payoffs; the real same-basis gap is −0.1222).
  • The gate does not clear the win/+EV bar — it loses to a per-city win-rate lookup and a price-free ex-ante model, and ties ungated. Not a learned edge.
  • The live dry-run GO is unaffected in mechanism but weakened in expected payoff — the loop is still the right next experiment (only it measures real fillability), but the prize shrank from ~+0.14/, and the thin-cohort problem (14–19 trades/latency) is now more binding because the per-trade edge halved.

Revised one-line verdict: CONDITIONAL-GO downgraded to LIVE-MEASUREMENT-ONLY. No capital, no Rust, until the live loop shows a repeatable edge over an ungated / per-city baseline — not over the convergence gate.

Honest caveats (load-bearing)

  1. True-resolution coverage is 74.7% on the OOS test, not 93.6%. Every EV here is conditional on resolution; the full-sample 0.9361 is not the comparison basis.
  2. Coverage bias correlates with the signal. Per-city OOS coverage ranges 47% (karachi) → 100% (seattle); unresolved rows have higher mean entry (0.8457 vs 0.7881) and far lower convergence (15.6% vs 42.7%). EV is conditional on a non-representative resolved subset.
  3. Single season. One historical regime; win base rate ≈0.83 is more stable than the (declining) convergence base rate but still one season.
  4. Thin live-cohort persists, now more binding — the per-trade edge we are trying to detect just halved (~+0.14 → ~+0.06), so the sample needed to detect it grew.
  5. Still hold-to-resolution. Payoff = Gamma terminal outcome; entry = price the copied wallet paid (shadow loop clamps to max(mid, trader_prior_paid)). The book reprices past us ~40–50% of the time (adverse selection).
  6. Early-fold instability reinforces the price-picker reading — first expanding fold AUC 0.555 (logistic) / 0.572 (GBT) vs ~0.68 later. The ~0.68 is the market price re-expressed, not a learned edge.

No leakage; the defect was the COMPARISON, not the model

All 9 features are entry-time-known (entry_gap == fair_gap, verified); the label is the true Gamma terminal outcome; the split is temporal and clean (last 25% OOS). The model was not refit to fix a leak (there was none) — the fixed defect was a hard-coded 0.14 literal computed on a different population with a proxy payoff model, replaced with a same-basis re-score, then benched against baselines that reveal the model adds no edge over the market price.

Repo pointers

Source files (now under /coding/pmv2/us_weather_algo/)

  • Full report: research/reports/weather-gate-winlabel-2026-06-19.md
  • Artifacts: research/gate/{build_win_gate.py, winlabel_ev.json, gate_win_metrics.json, gate_win_model.pkl}
  • Note: build_win_gate is research-only and is intentionally NOT carried into the as_weather_algo producer (baseline benchmarking runs in the research repo on an exported shadow_measure_log, not in the producer).