UPDATE 2026-06-22 — this EV belongs to the
weather_aslane (own repo, pmv2 reorg) and was later re-scoredThis net-of-cost pre-check is part of the Asian-weather chain now packaged as the
weather_asproducer: sourceweather_as, subjectpmv2.order.weather_as.entry, PRIVATE repopmv2-weather-as-algoat/coding/pmv2/as_weather_algo(committede9415dc). The research record stays in the renamed US repopmv2-weather-us-algoat/coding/pmv2/us_weather_algo—research/...paths below resolve there. Important: the +0.140/ (loses to a per-city win-rate lookup). Treat +0.140/$ as an upper bound. Full reorg + current state: pmv2-reorg-and-weather-as-2026-06-22.
The capstone of the fillability arc: the binding net-of-cost question that weather-lag-distribution-2026-06-19 left as open item #1. On the gate’s OOS test set, the gated hold-to-resolution strategy is positive-EV net of historically realistic costs — but the headline is measured on a single 11-day summer slice at the traders’ (favorable) fills, so it is an optimistic upper bound on what WE would capture. Verdict: GO-SMALL — a small, instrumented dry rehearsal is justified; scaled capital is not. Full report: research/reports/weather-netcost-precheck-2026-06-19.md.
GO-SMALL, not GO — the edge is real but the evidence is thin and the fills are not ours
The gated edge is positive-EV and survives the worst modelled cost corner (+0.045/ structural loser into +14c/$). But three caveats hold it to GO-SMALL, not a scaled GO:
- Single-season. The entire +0.140 headline (and the +0.045 floor) is measured on ONE 11-day summer window (2026-06-09 → 06-19, 100% June-2026). The data is exponentially front-loaded, so “last 25% by
entry_ts” collapses to this slice. The classifier AUC generalizes across time (expanding folds 0.66→0.78); the EV/payoff has never been evaluated on any other period. No evidence it holds in a winter cold-snap or shoulder-season.- Trader-fill optimism (unmodeled). The headline uses the traders’ fills — small size (median 33 shares), +2c over book, 24% at/below book. Our copy strategy detects the same lag and enters afterward, into an already-converging thinner book — so our realized entry is systematically worse and this EV is an upper bound.
netcost.pydoes not model this; the rehearsal is precisely the test of it.- Unverifiable depth. The dataset carries a single
book_at_entrylevel and no order-book-depth field; slippage was only directly tested to 2 ticks. Central breakeven is ~11 ticks, pessimistic ~5.7 — finite headroom, eroding fast.The rehearsal is not a test of whether the edge exists (this pre-check shows it does, on this slice) — it is a test of fillability and latency: can WE capture enough of the +14c to clear the +5.7c pessimistic breakeven after our real fill degradation.
For Agents
Five load-bearing facts:
- The gate is essential, not optional. Ungated OOS EV = -0.177/1, 93% win rate. Gate lift = +0.317/$1. Without the gate there is no strategy.
- Positive across the entire fee×slippage grid. Worst historical cell (fee 0.02, slip 0.02) = +0.091/$1. EV rises monotonically with gate score: P>=0.60 +0.113, P>=0.68 +0.140, P>=0.75 +0.185 → the classifier ranks by exactly the thing that pays (internal-consistency check passes).
- Pessimistic corner survives. fee 0.02 + slip 0.02 + every unresolved entry counted as a loss = +0.045/$1, win rate 0.894, n=611. This is the worst corner
netcost.pymodels (and it does NOT model fill degradation — see caveat 2).- NEW operational finding — the money is in ASIAN-AFTERNOON cities, not the US cities
wx-signaltrades. Even gated, EV concentrates in Seoul (+0.17), Wuhan (+0.23), Hong Kong (+0.27), Tokyo/Chengdu/Taipei. In NYC — which the live US engine trades — gated EV is NEGATIVE (-0.16) on this slice; Atlanta/Chicago thinly positive. Small-n/single-season so directional, but flags a city-universe mismatch.- Resolution is 75% real, not all proxy. 74.7% true-resolution coverage (Gamma
outcomePrices+ index-matched terminal token price); the rest fall back to a flagged “converged→win” proxy. The proxy decides only 3.8% of headline entries; flipping 30% of proxy wins to losses moves EV 0.140→0.126. Not driving the result.
The verdict in one line
If EV were negative even here — at the traders’ favorable fills — the idea would be dead and there would be no point rehearsing. It is not negative; it is +0.140 central / +0.045 pessimistic. Because it is positive and survives the pessimistic corner, the live dry rehearsal is justified — but a GO on the rehearsal is not a GO on capital.
This downgrades the JSON’s blanket GO to GO-SMALL. The JSON verdict (netcost.py L228-236) is generated by a rule that only checks central > 0 AND pessimistic > 0 — it has no guard for OOS-window degeneracy, fill optimism, or regime coverage. The three caveats above are those missing guards.
Headline economics (gate P>=0.68, the recommended threshold)
| Cohort | n | EV / $1 | win rate | mean entry |
|---|---|---|---|---|
| Gated headline (P>=0.68), zero-cost central | 611 | +0.140 | 0.931 | 0.817 |
| Ungated OOS (everything the traders did) | 7,073 | -0.177 | 0.661 | 0.803 |
| Gate EV lift vs ungated | +0.317 / $1 |
The gate flips a structurally money-losing population (-17.7c/, 93% win rate). That lift is the core result and it is large. It is also the reason the gate (weather-exante-gate-2026-06-19) is non-negotiable — the ungated population is uneconomic.
Cost grid — historical corner (gated headline EV/$1)
| slip 0.0 | slip 0.01 | slip 0.02 | |
|---|---|---|---|
| fee 0.0 | +0.140 | +0.128 | +0.116 |
| fee 0.01 | +0.128 | +0.116 | +0.103 |
| fee 0.02 | +0.116 | +0.103 | +0.091 |
Positive in every cell. Polymarket taker fee is historically ~0; fee is sensitivity-tested anyway. The premium-over-book is already inside entry_price, so the grid is pure incremental fee + slippage on top of the realistic entry.
Extended slippage axis + breakeven (recomputed, central, fee=0)
The historical grid stops at 2 ticks — too short for the thin late-day books these entries hit. Extended:
| slip | 0.02 | 0.03 | 0.05 | 0.08 | 0.10 | breakeven |
|---|---|---|---|---|---|---|
| central EV/$1 | +0.116 | +0.103 | +0.079 | +0.042 | +0.018 | +0.1144 (~11.4 ticks) |
| pessimistic EV/$1 (fee 0.02, unresolved=loss) | +0.045 | +0.033 | +0.008 | -0.029 | -0.053 | +0.0567 (~5.7 ticks) |
The central case tolerates ~11 ticks before breaking even; the pessimistic case ~5.7. Two ticks is not the stress ceiling — there is real headroom — but it is finite, and the pessimistic floor erodes fast.
The pessimistic corner
fee=0.02, slip=0.02, every unresolved entry counted as a loss: EV/$1 = +0.0449, win rate 0.894, n=611. At zero slippage but still counting unresolved-as-loss, pessimistic EV is +0.0694. The edge survives the worst corner the script models — but the script does NOT model the fill caveat below.
Threshold monotonicity (sanity check, central zero-cost)
| threshold | n | EV/$1 | win rate |
|---|---|---|---|
| P>=0.60 | 1,098 | +0.113 | 0.907 |
| P>=0.68 (recommended) | 611 | +0.140 | 0.931 |
| P>=0.75 | 265 | +0.185 | 0.962 |
EV rises monotonically with the gate score — the classifier is ranking by exactly the thing that pays. Good internal-consistency signal for weather-exante-gate-2026-06-19.
Whose EV is this? (honest framing — the fill-optimism caveat)
This pre-check measures the traders’ entries at the traders’ fills, held to resolution. That is the right object for a go/no-go on the idea, but it is an upper bound on our EV, for three reasons netcost.py does not capture:
- We fill later and worse. Median headline trade is 33 shares (p25 = 13), median
entry_gapis +2.0c over book, and 146/611 (23.9%) filled at or belowbook_at_entry. These are small, early, cheap fills into a still-lagging book. Our copy strategy detects the same lag and enters afterward, into an already-converging book → our realized entry is systematically higher. The headline bakes in their entry advantage. - Depth is unverified. The dataset carries a single
book_at_entryprice level and no order-book-depth field (confirmed:lag_distribution.jsonhas fair-gap and lag stats but no depth/notional). We cannot prove how many shares are available at book; the lag study established the gap, not the size behind it. Thin late-day books are exactly where fill degradation bites. - Live feature distribution shift.
book_at_entryandentry_gapare model features taken from the trader’s snapshot (predict.pyL42). In live use we substitute our own book read — a distribution shift on two of the gate’s inputs (entry_gapis the #2 permutation-importance feature afterlocal_hour). The gate’s live precision is therefore not guaranteed to equal its OOS precision.
Fill-degradation haircut (recomputed): if our realized entry is N ticks worse than the trader’s —
| haircut | central EV/$1 | pessimistic EV/$1 |
|---|---|---|
| 0 ticks | +0.140 | +0.094 |
| +2 ticks | +0.113 | +0.068 |
| +5 ticks | +0.074 | +0.031 |
The edge survives a few ticks of fill degradation but is not indestructible. This is the whole rehearsal question: the rehearsal passes only if post-haircut pessimistic EV stays > 0.
New operational finding — the edge is in ASIAN-AFTERNOON cities, not the US cities wx-signal trades
Even gated, EV is heavily concentrated in Asian-afternoon temperature books — which is exactly where the weather-exante-gate-2026-06-19 local_hour signal and the weather-lag-distribution-2026-06-19 city ranking already pointed.
| city | n (gated P>=0.68) | EV/$1 | win rate | note |
|---|---|---|---|---|
| wuhan | 67 | +0.231 | 0.955 | richest well-powered name |
| hong-kong | 22 | +0.272 | 1.00 | |
| seoul | 141 | +0.173 | 0.929 | largest cohort |
| tokyo | 10 | +0.309 | 1.00 | thin-n |
| chengdu | 18 | +0.231 | 1.00 | |
| taipei | 10 | +0.264 | 1.00 | thin-n |
| paris | 23 | +0.171 | 1.00 | |
| san-francisco | 29 | +0.146 | 1.00 | |
| atlanta | 13 | +0.146 | 0.923 | thinly positive |
| chicago | 12 | +0.026 | 0.917 | thinly positive |
| nyc | 21 | -0.164 | 0.714 | NEGATIVE — the live US engine trades this |
| london | 30 | -0.034 | 0.833 | thin negative |
| istanbul | 7 | -0.158 | 0.714 | thin negative |
| amsterdam | 5 | -0.099 | 0.800 | thin negative |
- Seoul + Wuhan alone = 34.0% of headline entries (208/611) and the richest names (combined EV +0.191 vs +0.115 for the other 403). The edge survives excluding them: the rest of the book is +0.115/ in the full pessimistic corner — i.e. the non-Seoul/Wuhan book is essentially break-even once you pile on max cost + unresolved-as-loss.
- Day-by-day EV is positive on all 11 OOS days (range +0.022 to +0.251) → not one outlier session carrying the slice.
- The low-n negatives (nyc, london, istanbul, amsterdam) should be excluded or capped — consistent with the tier-A/B recommendation in weather-lag-distribution-2026-06-19.
City-universe mismatch — the money is not where the (US) engine is wired → now addressed by the
weather_aslaneGated EV is negative in NYC (-0.16) and thin/negative in London on this slice, while strongly positive in the Asian-afternoon cluster (Seoul/Wuhan/HK/Tokyo/Chengdu/Taipei). The live
wx-signalengine (weather-bet-live-status) trades the US cities (NYC/Atlanta/Chicago/Miami). This is small-n and single-season, so directional, not a re-pointing mandate — but it flags that the measured edge sits in Asian-afternoon temperature books, not the US universe the engine is wired into. (2026-06-22) This mismatch is the reason the Asian cluster was carved into its ownweather_asproducer (pmv2-reorg-and-weather-as-2026-06-22) — a separate lane on its own Asian-afternoon cohort, rather than re-pointing the US engine.
Resolution quality — 75% true, not proxy-driven
- 74.7% true-resolution coverage (5,282/7,073 OOS rows): Gamma
/events?slugoutcomePrices+ terminal token price index-matched to the entry asset; >0.5 ⇒ won. The rest fall back to a flagged “converged→win” proxy. - The proxy is barely load-bearing: converged-and-truly-resolved resolved NO only 1.1% full-test / 1.7% headline; the proxy is the deciding factor for just 23/611 (3.8%) of headline entries.
- Stress: flipping 30% of proxy wins to losses moves headline EV 0.140 → 0.126. The pessimistic corner already counts all 29 unresolved entries as losses. The win-proxy is not driving the result.
Residual concerns (adversarial pass)
| # | Issue | Sev | Disposition |
|---|---|---|---|
| 1 | JSON verdict “GO” is unguarded | high | Downgraded to GO-SMALL + 3 caveats. Verdict generator (netcost.py L228-236) only checks central>0 & pess>0 — no OOS/fill/regime guard. |
| 2 | OOS “test split” collapses to ~11 summer days | high | Disclosed, central limitation. Full OOS test = 2026-06-09→06-19, 11 distinct days, 100% Jun-2026. Headline EV is single-season; AUC generalizes, EV does not. |
| 3 | Hold-to-resolution EV uses TRADERS’ fills | high | Caveated as upper bound + haircut applied. Median size 33, gap +2.0c, 23.9% at/below book. Rehearsal is the fillability test. |
| 4 | Slippage axis too short; depth unverifiable | med | Recomputed + breakeven published (central ~11.4 ticks, pessimistic ~5.7). Depth genuinely unverifiable — one price level in data. |
| 5 | Convergence-as-win proxy soundness | low | Validated. Proxy decides only 3.8% of headline; 30%-flip → 0.140→0.126. |
| 6 | EV is concentration-/circularity-driven | low | Transparency fix. Seoul+Wuhan 34%, EV survives excluding them (+0.115 / +0.004 pess); per-day positive on all 11 days; 23 entries use the gate’s own “converged” label (mild leakage). |
Scale-up gates (before any real capital)
This is a GO on the rehearsal, not on capital. Scaled deployment is gated on:
- Post-fill-degradation pessimistic EV stays positive. Instrument the rehearsal to measure our realized entry vs book at detection latency, recompute the corner with that empirical haircut. Pass only if post-haircut pessimistic EV > 0.
- Out-of-summer EV re-validation. Re-evaluate EV on an earlier held-out window (e.g. a winter slice) or use blocked/seasonal CV for the EV — not just the AUC. There is currently zero evidence the favorable EV holds outside June-2026.
- A real (wider) slippage curve. The 2-tick grid is too short for these thin late-day books; the rehearsal should produce an empirical slippage curve to replace the modelled one.
It becomes NO-GO for scaled capital if the rehearsal’s measured fill degradation pushes pessimistic EV ⇐ 0, or if the out-of-summer re-validation comes back negative.
Bottom line
The gate does real work (+0.317/), the edge clears cost in the pessimistic corner (+4.5c/$), and the win-proxy is not driving the result. But the headline is one 11-day summer slice at the traders’ favorable fills, with unverifiable depth, and the EV concentrates in Asian-afternoon cities the live US engine does not trade. That combination is exactly what a small, instrumented live dry rehearsal is for: GO on the rehearsal as a fillability/latency probe; gate any scaled capital on post-fill-haircut pessimistic EV staying positive and an out-of-summer EV re-validation.
This closes weather-lag-distribution-2026-06-19’s open item #1 (net-of-cost recovery) as ADDRESSED (conditional) — positive on this slice, but pending live fill-validation, so it is not “closed,” it is “addressed, conditional.”
Repo pointers
Source files
- Full report:
research/reports/weather-netcost-precheck-2026-06-19.md- Gate artifact:
research/gate/netcost.json(verdict generator + cohort/grid/corner figures)- Inputs:
research/gate/gate_metrics.json, raw lag shardsresearch/copytrade/top50/lag/measured_shard_*.json,research/gate/resolution_cache.json- Method:
research/gate/netcost.py(entry stake =entry_price; EV/$1 = mean PnL per share / mean stake; cost = fee + slippage subtracted from payoff). All figures in the report were reproduced from raw shards + gate model + resolution cache, not copied from the JSON.
Related
- pmv2-reorg-and-weather-as-2026-06-22 — the pmv2 reorg + the
weather_aslane this EV underwrites - weather-gate-winlabel-2026-06-19 — the WIN-label re-score showing this +0.140/)
- weather-lag-distribution-2026-06-19 — the unbiased reprice-lag study that named net-of-cost recovery the binding economic gate (its open item #1, now addressed-conditional here); established the ~2-tick premium entry and the ex-ante city universe
- weather-exante-gate-2026-06-19 — the lean
P(tradeable)classifier (gate P>=0.68) whose selection is what makes this EV positive (gate lift +0.317/$); this pre-check is the downstream economic test the gate note deferred to - weather-top50-decypher-2026-06-19 — the v2 copy-trade decypher that produced the entry population; its “≤0.95 fillable edge plausible but NOT proven” is now addressed-conditional (positive on this slice, pending live fills)
- weather-bet-live-status — the live
wx-signalengine + the dry rehearsal that must settle the fill-degradation haircut; note the city-universe mismatch (engine trades US; the measured edge is Asian-afternoon) - weather-strategy-validated-plan — the late-settlement / repricing-lag GO plan this whole arc sizes and stress-tests
- weather-bet — project overview