UPDATE 2026-06-22 — this EV belongs to the weather_as lane (own repo, pmv2 reorg) and was later re-scored

This net-of-cost pre-check is part of the Asian-weather chain now packaged as the weather_as producer: source weather_as, subject pmv2.order.weather_as.entry, PRIVATE repo pmv2-weather-as-algo at /coding/pmv2/as_weather_algo (committed e9415dc). The research record stays in the renamed US repo pmv2-weather-us-algo at /coding/pmv2/us_weather_algoresearch/... paths below resolve there. Important: the +0.140/ (loses to a per-city win-rate lookup). Treat +0.140/$ as an upper bound. Full reorg + current state: pmv2-reorg-and-weather-as-2026-06-22.

The capstone of the fillability arc: the binding net-of-cost question that weather-lag-distribution-2026-06-19 left as open item #1. On the gate’s OOS test set, the gated hold-to-resolution strategy is positive-EV net of historically realistic costs — but the headline is measured on a single 11-day summer slice at the traders’ (favorable) fills, so it is an optimistic upper bound on what WE would capture. Verdict: GO-SMALL — a small, instrumented dry rehearsal is justified; scaled capital is not. Full report: research/reports/weather-netcost-precheck-2026-06-19.md.

GO-SMALL, not GO — the edge is real but the evidence is thin and the fills are not ours

The gated edge is positive-EV and survives the worst modelled cost corner (+0.045/ structural loser into +14c/$). But three caveats hold it to GO-SMALL, not a scaled GO:

  1. Single-season. The entire +0.140 headline (and the +0.045 floor) is measured on ONE 11-day summer window (2026-06-09 → 06-19, 100% June-2026). The data is exponentially front-loaded, so “last 25% by entry_ts” collapses to this slice. The classifier AUC generalizes across time (expanding folds 0.66→0.78); the EV/payoff has never been evaluated on any other period. No evidence it holds in a winter cold-snap or shoulder-season.
  2. Trader-fill optimism (unmodeled). The headline uses the traders’ fills — small size (median 33 shares), +2c over book, 24% at/below book. Our copy strategy detects the same lag and enters afterward, into an already-converging thinner book — so our realized entry is systematically worse and this EV is an upper bound. netcost.py does not model this; the rehearsal is precisely the test of it.
  3. Unverifiable depth. The dataset carries a single book_at_entry level and no order-book-depth field; slippage was only directly tested to 2 ticks. Central breakeven is ~11 ticks, pessimistic ~5.7 — finite headroom, eroding fast.

The rehearsal is not a test of whether the edge exists (this pre-check shows it does, on this slice) — it is a test of fillability and latency: can WE capture enough of the +14c to clear the +5.7c pessimistic breakeven after our real fill degradation.

For Agents

Five load-bearing facts:

  1. The gate is essential, not optional. Ungated OOS EV = -0.177/1, 93% win rate. Gate lift = +0.317/$1. Without the gate there is no strategy.
  2. Positive across the entire fee×slippage grid. Worst historical cell (fee 0.02, slip 0.02) = +0.091/$1. EV rises monotonically with gate score: P>=0.60 +0.113, P>=0.68 +0.140, P>=0.75 +0.185 → the classifier ranks by exactly the thing that pays (internal-consistency check passes).
  3. Pessimistic corner survives. fee 0.02 + slip 0.02 + every unresolved entry counted as a loss = +0.045/$1, win rate 0.894, n=611. This is the worst corner netcost.py models (and it does NOT model fill degradation — see caveat 2).
  4. NEW operational finding — the money is in ASIAN-AFTERNOON cities, not the US cities wx-signal trades. Even gated, EV concentrates in Seoul (+0.17), Wuhan (+0.23), Hong Kong (+0.27), Tokyo/Chengdu/Taipei. In NYC — which the live US engine trades — gated EV is NEGATIVE (-0.16) on this slice; Atlanta/Chicago thinly positive. Small-n/single-season so directional, but flags a city-universe mismatch.
  5. Resolution is 75% real, not all proxy. 74.7% true-resolution coverage (Gamma outcomePrices + index-matched terminal token price); the rest fall back to a flagged “converged→win” proxy. The proxy decides only 3.8% of headline entries; flipping 30% of proxy wins to losses moves EV 0.140→0.126. Not driving the result.

The verdict in one line

If EV were negative even here — at the traders’ favorable fills — the idea would be dead and there would be no point rehearsing. It is not negative; it is +0.140 central / +0.045 pessimistic. Because it is positive and survives the pessimistic corner, the live dry rehearsal is justified — but a GO on the rehearsal is not a GO on capital.

This downgrades the JSON’s blanket GO to GO-SMALL. The JSON verdict (netcost.py L228-236) is generated by a rule that only checks central > 0 AND pessimistic > 0 — it has no guard for OOS-window degeneracy, fill optimism, or regime coverage. The three caveats above are those missing guards.

CohortnEV / $1win ratemean entry
Gated headline (P>=0.68), zero-cost central611+0.1400.9310.817
Ungated OOS (everything the traders did)7,073-0.1770.6610.803
Gate EV lift vs ungated+0.317 / $1

The gate flips a structurally money-losing population (-17.7c/, 93% win rate). That lift is the core result and it is large. It is also the reason the gate (weather-exante-gate-2026-06-19) is non-negotiable — the ungated population is uneconomic.

Cost grid — historical corner (gated headline EV/$1)

slip 0.0slip 0.01slip 0.02
fee 0.0+0.140+0.128+0.116
fee 0.01+0.128+0.116+0.103
fee 0.02+0.116+0.103+0.091

Positive in every cell. Polymarket taker fee is historically ~0; fee is sensitivity-tested anyway. The premium-over-book is already inside entry_price, so the grid is pure incremental fee + slippage on top of the realistic entry.

Extended slippage axis + breakeven (recomputed, central, fee=0)

The historical grid stops at 2 ticks — too short for the thin late-day books these entries hit. Extended:

slip0.020.030.050.080.10breakeven
central EV/$1+0.116+0.103+0.079+0.042+0.018+0.1144 (~11.4 ticks)
pessimistic EV/$1 (fee 0.02, unresolved=loss)+0.045+0.033+0.008-0.029-0.053+0.0567 (~5.7 ticks)

The central case tolerates ~11 ticks before breaking even; the pessimistic case ~5.7. Two ticks is not the stress ceiling — there is real headroom — but it is finite, and the pessimistic floor erodes fast.

The pessimistic corner

fee=0.02, slip=0.02, every unresolved entry counted as a loss: EV/$1 = +0.0449, win rate 0.894, n=611. At zero slippage but still counting unresolved-as-loss, pessimistic EV is +0.0694. The edge survives the worst corner the script models — but the script does NOT model the fill caveat below.

Threshold monotonicity (sanity check, central zero-cost)

thresholdnEV/$1win rate
P>=0.601,098+0.1130.907
P>=0.68 (recommended)611+0.1400.931
P>=0.75265+0.1850.962

EV rises monotonically with the gate score — the classifier is ranking by exactly the thing that pays. Good internal-consistency signal for weather-exante-gate-2026-06-19.

Whose EV is this? (honest framing — the fill-optimism caveat)

This pre-check measures the traders’ entries at the traders’ fills, held to resolution. That is the right object for a go/no-go on the idea, but it is an upper bound on our EV, for three reasons netcost.py does not capture:

  1. We fill later and worse. Median headline trade is 33 shares (p25 = 13), median entry_gap is +2.0c over book, and 146/611 (23.9%) filled at or below book_at_entry. These are small, early, cheap fills into a still-lagging book. Our copy strategy detects the same lag and enters afterward, into an already-converging book → our realized entry is systematically higher. The headline bakes in their entry advantage.
  2. Depth is unverified. The dataset carries a single book_at_entry price level and no order-book-depth field (confirmed: lag_distribution.json has fair-gap and lag stats but no depth/notional). We cannot prove how many shares are available at book; the lag study established the gap, not the size behind it. Thin late-day books are exactly where fill degradation bites.
  3. Live feature distribution shift. book_at_entry and entry_gap are model features taken from the trader’s snapshot (predict.py L42). In live use we substitute our own book read — a distribution shift on two of the gate’s inputs (entry_gap is the #2 permutation-importance feature after local_hour). The gate’s live precision is therefore not guaranteed to equal its OOS precision.

Fill-degradation haircut (recomputed): if our realized entry is N ticks worse than the trader’s —

haircutcentral EV/$1pessimistic EV/$1
0 ticks+0.140+0.094
+2 ticks+0.113+0.068
+5 ticks+0.074+0.031

The edge survives a few ticks of fill degradation but is not indestructible. This is the whole rehearsal question: the rehearsal passes only if post-haircut pessimistic EV stays > 0.

New operational finding — the edge is in ASIAN-AFTERNOON cities, not the US cities wx-signal trades

Even gated, EV is heavily concentrated in Asian-afternoon temperature books — which is exactly where the weather-exante-gate-2026-06-19 local_hour signal and the weather-lag-distribution-2026-06-19 city ranking already pointed.

cityn (gated P>=0.68)EV/$1win ratenote
wuhan67+0.2310.955richest well-powered name
hong-kong22+0.2721.00
seoul141+0.1730.929largest cohort
tokyo10+0.3091.00thin-n
chengdu18+0.2311.00
taipei10+0.2641.00thin-n
paris23+0.1711.00
san-francisco29+0.1461.00
atlanta13+0.1460.923thinly positive
chicago12+0.0260.917thinly positive
nyc21-0.1640.714NEGATIVE — the live US engine trades this
london30-0.0340.833thin negative
istanbul7-0.1580.714thin negative
amsterdam5-0.0990.800thin negative
  • Seoul + Wuhan alone = 34.0% of headline entries (208/611) and the richest names (combined EV +0.191 vs +0.115 for the other 403). The edge survives excluding them: the rest of the book is +0.115/ in the full pessimistic corner — i.e. the non-Seoul/Wuhan book is essentially break-even once you pile on max cost + unresolved-as-loss.
  • Day-by-day EV is positive on all 11 OOS days (range +0.022 to +0.251) → not one outlier session carrying the slice.
  • The low-n negatives (nyc, london, istanbul, amsterdam) should be excluded or capped — consistent with the tier-A/B recommendation in weather-lag-distribution-2026-06-19.

City-universe mismatch — the money is not where the (US) engine is wired → now addressed by the weather_as lane

Gated EV is negative in NYC (-0.16) and thin/negative in London on this slice, while strongly positive in the Asian-afternoon cluster (Seoul/Wuhan/HK/Tokyo/Chengdu/Taipei). The live wx-signal engine (weather-bet-live-status) trades the US cities (NYC/Atlanta/Chicago/Miami). This is small-n and single-season, so directional, not a re-pointing mandate — but it flags that the measured edge sits in Asian-afternoon temperature books, not the US universe the engine is wired into. (2026-06-22) This mismatch is the reason the Asian cluster was carved into its own weather_as producer (pmv2-reorg-and-weather-as-2026-06-22) — a separate lane on its own Asian-afternoon cohort, rather than re-pointing the US engine.

Resolution quality — 75% true, not proxy-driven

  • 74.7% true-resolution coverage (5,282/7,073 OOS rows): Gamma /events?slug outcomePrices + terminal token price index-matched to the entry asset; >0.5 ⇒ won. The rest fall back to a flagged “converged→win” proxy.
  • The proxy is barely load-bearing: converged-and-truly-resolved resolved NO only 1.1% full-test / 1.7% headline; the proxy is the deciding factor for just 23/611 (3.8%) of headline entries.
  • Stress: flipping 30% of proxy wins to losses moves headline EV 0.140 → 0.126. The pessimistic corner already counts all 29 unresolved entries as losses. The win-proxy is not driving the result.

Residual concerns (adversarial pass)

#IssueSevDisposition
1JSON verdict “GO” is unguardedhighDowngraded to GO-SMALL + 3 caveats. Verdict generator (netcost.py L228-236) only checks central>0 & pess>0 — no OOS/fill/regime guard.
2OOS “test split” collapses to ~11 summer dayshighDisclosed, central limitation. Full OOS test = 2026-06-09→06-19, 11 distinct days, 100% Jun-2026. Headline EV is single-season; AUC generalizes, EV does not.
3Hold-to-resolution EV uses TRADERS’ fillshighCaveated as upper bound + haircut applied. Median size 33, gap +2.0c, 23.9% at/below book. Rehearsal is the fillability test.
4Slippage axis too short; depth unverifiablemedRecomputed + breakeven published (central ~11.4 ticks, pessimistic ~5.7). Depth genuinely unverifiable — one price level in data.
5Convergence-as-win proxy soundnesslowValidated. Proxy decides only 3.8% of headline; 30%-flip → 0.140→0.126.
6EV is concentration-/circularity-drivenlowTransparency fix. Seoul+Wuhan 34%, EV survives excluding them (+0.115 / +0.004 pess); per-day positive on all 11 days; 23 entries use the gate’s own “converged” label (mild leakage).

Scale-up gates (before any real capital)

This is a GO on the rehearsal, not on capital. Scaled deployment is gated on:

  1. Post-fill-degradation pessimistic EV stays positive. Instrument the rehearsal to measure our realized entry vs book at detection latency, recompute the corner with that empirical haircut. Pass only if post-haircut pessimistic EV > 0.
  2. Out-of-summer EV re-validation. Re-evaluate EV on an earlier held-out window (e.g. a winter slice) or use blocked/seasonal CV for the EV — not just the AUC. There is currently zero evidence the favorable EV holds outside June-2026.
  3. A real (wider) slippage curve. The 2-tick grid is too short for these thin late-day books; the rehearsal should produce an empirical slippage curve to replace the modelled one.

It becomes NO-GO for scaled capital if the rehearsal’s measured fill degradation pushes pessimistic EV 0, or if the out-of-summer re-validation comes back negative.

Bottom line

The gate does real work (+0.317/), the edge clears cost in the pessimistic corner (+4.5c/$), and the win-proxy is not driving the result. But the headline is one 11-day summer slice at the traders’ favorable fills, with unverifiable depth, and the EV concentrates in Asian-afternoon cities the live US engine does not trade. That combination is exactly what a small, instrumented live dry rehearsal is for: GO on the rehearsal as a fillability/latency probe; gate any scaled capital on post-fill-haircut pessimistic EV staying positive and an out-of-summer EV re-validation.

This closes weather-lag-distribution-2026-06-19’s open item #1 (net-of-cost recovery) as ADDRESSED (conditional) — positive on this slice, but pending live fill-validation, so it is not “closed,” it is “addressed, conditional.”

Repo pointers

Source files

  • Full report: research/reports/weather-netcost-precheck-2026-06-19.md
  • Gate artifact: research/gate/netcost.json (verdict generator + cohort/grid/corner figures)
  • Inputs: research/gate/gate_metrics.json, raw lag shards research/copytrade/top50/lag/measured_shard_*.json, research/gate/resolution_cache.json
  • Method: research/gate/netcost.py (entry stake = entry_price; EV/$1 = mean PnL per share / mean stake; cost = fee + slippage subtracted from payoff). All figures in the report were reproduced from raw shards + gate model + resolution cache, not copied from the JSON.
  • pmv2-reorg-and-weather-as-2026-06-22 — the pmv2 reorg + the weather_as lane this EV underwrites
  • weather-gate-winlabel-2026-06-19 — the WIN-label re-score showing this +0.140/)
  • weather-lag-distribution-2026-06-19 — the unbiased reprice-lag study that named net-of-cost recovery the binding economic gate (its open item #1, now addressed-conditional here); established the ~2-tick premium entry and the ex-ante city universe
  • weather-exante-gate-2026-06-19 — the lean P(tradeable) classifier (gate P>=0.68) whose selection is what makes this EV positive (gate lift +0.317/$); this pre-check is the downstream economic test the gate note deferred to
  • weather-top50-decypher-2026-06-19 — the v2 copy-trade decypher that produced the entry population; its “≤0.95 fillable edge plausible but NOT proven” is now addressed-conditional (positive on this slice, pending live fills)
  • weather-bet-live-status — the live wx-signal engine + the dry rehearsal that must settle the fill-degradation haircut; note the city-universe mismatch (engine trades US; the measured edge is Asian-afternoon)
  • weather-strategy-validated-plan — the late-settlement / repricing-lag GO plan this whole arc sizes and stress-tests
  • weather-bet — project overview