UPDATE 2026-06-22 — this study underpins the
weather_aslane (own repo, pmv2 reorg)The fillability arc this study anchors is now packaged as the Asian-weather (
weather_as) producer: sourceweather_as, subjectpmv2.order.weather_as.entry, PRIVATE repopmv2-weather-as-algoat/coding/pmv2/as_weather_algo(committede9415dc). The research record stays in the renamed US repopmv2-weather-us-algoat/coding/pmv2/us_weather_algo— theresearch/...paths below resolve there, not the oldweather_bet. Full reorg + current state, the multi-producer / single-executor (position_manager) architecture, and the maturity caveats: pmv2-reorg-and-weather-as-2026-06-22.
The unbiased, full-population reprice-lag measurement that settles the fillability open items left by weather-top50-decypher-2026-06-19. Every cheap-Yes buy (0.55–0.95) across the top-50 was measured — no highest-paid-40 cap that biased the prior cohort — for 28,664 entries, then adversarially critic-reviewed. Verdict: weak-GO, materially weaker than the first pass. Full report: research/reports/weather-lag-distribution-2026-06-19.md.
Weak-GO — and the BINDING gate is net-of-cost recovery, not reach
The edge survives unbiased, horizon-coherent measurement, but on honest numbers the realized hit rate is ~34.4% (unconditional reach@15 = 0.344), not the 63–67% conditional figure — because ~49% of cheap entries never reprice to 0.97 at all. Size every position for a ~65% per-entry miss.
The strongest finding flips the prior thesis: winners pay a ~2-tick PREMIUM, not a discount (median
fair_gap+0.02, 75% above the same-token/same-instant book). This is a convergence-timing edge bought slightly rich, not a cheap fill. Therefore the binding economic gate is whether that +0.02 premium is recovered net of taker fee + slippage — a GO precondition, not an open item, and answerable only in the live dry rehearsal. If the premium is not recovered net, the edge is uneconomic regardless of reach.
For Agents
Final, complete-population revision of the fillability question. Five load-bearing facts to carry forward:
- Size on 0.344, not 0.675. Unconditional reach@15 (converge@180 AND lag>15) = 0.344; conditional
P(lag>15 | converged)= 0.675. The 1.96× gap is the 49.1% never-converge haircut. At entry you do not know whether a book converges → size on 0.344.- Thesis flip (most robust result): winners’ entry is a premium (
fair_gapp25 +0.00 / p50 +0.02 / p75 +0.06; 75.1% above book). NOT discount-to-fair. The decypher’s withdrawal of the discount claim was correct.- The gate must be EX-ANTE (city / session / METAR), never observed lag — lag is measured forward and gating on it is look-ahead (critic-killed). The lag distribution is descriptive only, for fill-window sizing.
- City universe is the implementable gate primitive: tier-A core = london, nyc, seoul, hong-kong, atlanta, paris, wuhan. Thin/provisional held out; fast-repricing excluded.
- Methodology is now sound: the 120/180 horizon bug (wx-lag-pipeline-bug) is fixed (aggregator horizon-authoritative); shard_2 (16.7%) was dropped then restored and the aggregate barely moved → robustness confirmed, prior numbers were not biased by the omission.
What it was
An unbiased full-population reprice-lag measurement, replacing the biased 10-wallet / 40-highest-paid-entry fillability cohort that Section 2 of weather-top50-decypher-2026-06-19 flagged as load-bearing-but-unproven.
- Population: every successfully-measured cheap-Yes buy in the 0.55–0.95 band across the top-50 wallets — 28,664 entries (
pull_failed=false) out of 28,839 built bybuild_entries.py(175 = 0.6% pull-failed). All 6 shards measured. - No sampling bias: the prior cohort capped at ≤40 highest-paid entries per wallet, which biased every sample-derived statistic (the paid>book finding, the lag distribution). This study removes that cap entirely.
- Measurement:
reprice_lag_min = first_at_or_after(hist, entry_ts, 0.97, 180)— the forward time until the book reprices to 0.97, within a unified 180-min horizon.fair_gap = entry_price − book_at_entrymeasured same-token / same-instant (the correction that settled the discount-vs-premium question). - Adversarially critic-reviewed: 8 residual concerns (R1–R8) raised and dispositioned; the two critical ones (conditional-vs-unconditional headline, look-ahead lag gate) rewrote the verdict.
Verdict — weak-GO, materially weaker than the first pass
The feared collapse (most lags below our reaction window) did not happen — among converging books ~68% still leave a >15-min window. But the prize is bounded by convergence, and that bound is the honest headline.
| metric | value | meaning |
|---|---|---|
| Unconditional reach@15 | 0.344 | of all 28,664 entries: converge@180 AND lag>15. Realized hit rate — size on this. |
Conditional reach@15 (P(lag>15 | converged)) | 0.675 | converging books only; window-shape, not ex-ante sizing |
| Conditional reach@10 | 0.730 | converging books only |
| Conditional reach@5 | 0.797 | converging books only |
| Convergence rate (180-min) | 0.509 | fraction that reprice to 0.97 at all |
| Never-converge fraction | 0.491 | conditioned away by the conditional metric |
The 1.96× gap between 0.344 and 0.675 is exactly the convergence haircut — it belongs in the headline, not a footnote. The engine must tolerate that ~65% of candidate entries will not pay off via reachable convergence and be sized for that miss rate. This is the central downgrade from the prior pass’s implied 63–67%.
Thesis flip — winners pay a ~2-tick PREMIUM (the most robust finding)
With the book measured same-token / same-instant, the winners’ entry is a premium (fair_gap > 0), not a discount-to-fair. This is the strongest, fully-robust result and the major correction to the prior thesis.
fair_gap = entry_price − book_at_entry; gap > 0 = premium / risk-taking, gap < 0 = discount-to-fair.- Distribution across converged entries: p25 = +0.00, p50 = +0.02, p75 = +0.06 — the premium sign holds across the bulk of the distribution, not a median artifact.
- frac_paid_above_book = 0.7514 — 75.1% paid strictly above the contemporaneous book.
- +0.02 is exactly 2 Polymarket ticks; only ~14% of entries sit within ±0.005 of book → a genuine, near-universal premium, not noise.
The 0.84–0.91 band the winners paid is not a discount the book had already validated. This is a lag/convergence-timing edge taken at a ~2c premium — the winners buy slightly rich into a book that is about to reprice, not cheap into a book that already agrees. The edge comes from the subsequent convergence to 0.97+, not from entering below fair.
Net-of-cost recovery is the BINDING economic gate
Because the entry is a premium, whether the +0.02 is recovered by post-entry convergence after taker fee + slippage is no longer an “open item” — it is a headline gating condition answerable only in the live dry rehearsal. If not recovered net, the edge is uneconomic regardless of reach. Do NOT instruct the engine to wait for a discount (
paid < book): that condition is present only ~25% of the time and filtering on it would discard most of the edge. Defensible rule: enter at/slightly-above book, capped at 0.95, only when the ex-ante predictor says convergence is physically plausible and expected post-entry move exceeds fee + slippage + the ~2c premium.
The gate must be EX-ANTE — observed lag is look-ahead (critic-killed)
A viable gate exists, but it cannot be “observed reprice lag” — that quantity is only known after the book has repriced. first_at_or_after(..., 180) is measured forward, so the prior draft’s rule “suppress entries where observed lag < p25” is not implementable and would make any backtest look-ahead-biased (critic R2).
The implementable gate is an ex-ante predictor:
- City / session tier — the per-city table is the gate primitive; city is known at entry. Restrict to the high-reach tier; exclude fast-repricing cities outright.
- METAR-implied convergence state — whether observed weather already makes the outcome near-certain (which drives fast repricing) is readable from our own METAR feed at entry; books already physically resolved reprice before we can react and should be suppressed.
- Reaction-window sizing, not lag-gating — treat the lag distribution as the payoff-window prior conditional on the chosen city/session, and size partial vs full fills against it.
Open requirement: calibrate the city/session/METAR predictor against realized lag and report its out-of-sample skill. Until then the universe restriction below is the gate, and the lag quantiles are descriptive only.
Converged lag distribution (unified 180-min horizon) — descriptive, for fill-window sizing:
| stat | minutes |
|---|---|
| p10 | 1.30 |
| p25 | 8.23 |
| p50 (median) | 37.13 |
| p75 | 81.63 |
| p90 | 127.60 |
| mean | 51.10 |
Heavily right-skewed; the fillable middle is not hollowed out (p25 8.2 → p50 37.1 is wide). The population edge survives because the right tail (p75 81.6, p90 127.6) supplies the reachable mass.
Per-wallet unbiased lag medians
Recomputed under the unified 180-min horizon (the prior draft’s per-wallet medians used the stale 120-min flag — the source of the now-retracted “Happening9014 casualty”):
| wallet | unbiased median, 180-min (min) | n_conv@180 | prior pop. median | prior sample median |
|---|---|---|---|---|
| HondaCivic | 8.63 | 189 | 24.9 | 4.7 |
| vip68 | 10.68 | 142 | 50.2 | 19.5 |
| Happening9014 | 20.60 | 459 | 39.4 | 38.9 |
- HondaCivic (8.6) and vip68 (10.7) sit inside the 15-min window even at 180-min → the edge cannot be gated on a per-wallet median for these two.
- Happening9014 (20.6, n=459) clears the 15-min window — the “typical lag clears the bar” claim from the prior pass is reinstated (the “casualty” verdict was a 120/180 horizon artifact, wx-lag-pipeline-bug).
- Reinforces the section above: the entry cannot be gated on future observed lag (per-wallet or per-entry) at all — only ex-ante on city/session/METAR.
City universe — which survive
Ranked by conditional reach@15 (P(lag>15|converged), n_converged ≥ 20 floor). Reliability tiers by converged sample size: A = n_conv ≥ 300, B = 100–299, C (thin) = 20–99. Complete 6-shard numbers (shard_2 restored).
Recommended universe (the ex-ante gate primitive)
Tier-A robust core —
london,nyc,seoul,hong-kong,atlanta,paris,wuhan. Add tier-Blagos,seattle,los-angeleswith size capped to their CI width. Select by the ex-ante city/session/METAR gate, not by US/intl label, and not by future observed lag.
| city | reach@15 | tier (n_conv) | lag median (min) | conv. rate | median fair_gap |
|---|---|---|---|---|---|
| london | 0.738 | A (1,855) | 45.8 | 0.547 | +0.022 |
| nyc | 0.741 | A (1,616) | 50.3 | 0.545 | +0.020 |
| seoul | 0.740 | A (1,179) | 41.7 | 0.629 | +0.013 |
| hong-kong | 0.764 | A (622) | 40.8 | 0.412 | +0.015 |
| atlanta | 0.719 | A (488) | 49.2 | 0.489 | +0.020 |
| paris | 0.760 | A (366) | 45.0 | 0.426 | +0.015 |
| wuhan | 0.747 | A (363) | 40.9 | 0.713 | +0.023 |
| seattle | 0.722 | B (273) | 35.4 | 0.587 | +0.015 |
| lagos | 0.785 | B (172) | 61.7 | 0.644 | +0.030 |
| los-angeles | 0.707 | B (140) | 49.3 | 0.471 | +0.022 |
Thin / provisional — held out of the firm core (high point estimate on thin support; wide Wilson CIs): istanbul (reach 0.792 but n=96, CI [0.700, 0.861] vs london [0.717, 0.757] at n=1,855), sao-paulo (n=97), lucknow (n=79), wellington (n=198). Do not read istanbul’s nominal #1 rank as equivalent to the tier-A cities — include only with size capped to CI uncertainty.
Fast-repricing — EXCLUDE (reach@15 below ~0.50, the book reprices before we can react): tel-aviv (0.468), taipei (0.467), cape-town (0.456), mexico-city (0.455), jakarta (0.441), buenos-aires (0.433), panama-city (0.425), kuala-lumpur (0.423), jeddah (0.407), denver (0.353), manila (0.270), karachi (0.254). Borderline (~0.50–0.61, keep out of the firm universe pending calibration): miami 0.579, moscow 0.604, austin 0.541, milan 0.506.
Prior-story reconciliation: London confirmed but not #1 (largest sample, robust member — prior over-weighted it); the Asian-hours cluster is confirmed and strengthened (hong-kong, wuhan, seoul lead; tokyo/qingdao/shenzhen/guangzhou survive — but “Asian-hours, specific cities,” not all of Asia: shanghai/chongqing/beijing/chengdu/taipei sit lower); NYC/Atlanta confirmed (refutes any wholesale US down-weight); Paris confirmed.
Methodology integrity
The study’s numbers are now coherent after two integrity fixes:
- 120-vs-180-min convergence-horizon bug — 5 shards baked
convergedat 120 min whilereprice_lag_minwas measured at 180 min, and the report’s published 180-min numbers were re-derived at analysis time but never persisted tolag_distribution.json. The aggregator now re-derives convergence at aggregation time (is_converged()againstCONVERGE_HORIZON_MIN=180), making it horizon-authoritative and immune to stale per-shard baked flags. Full backstory: wx-lag-pipeline-bug. Effect of unification: convergence 0.449 → 0.509, conditional reach@15 0.633 → 0.675, median lag → ~37 min (conservative for the GO call — the old number understated the edge — but now internally coherent). - shard_2 (16.7% of population) was dropped, then restored — re-measured at the unified 180-min horizon (4,807 rows, 50 pull-failed) and folded in (n_total 23,907 → 28,664). It carried the headline cities (london 510 / nyc 443 / seoul 292 / hong-kong 278 / atlanta 169 / paris 164). Adding it moved the aggregate negligibly (convergence 0.511 → 0.509, conditional reach@15 0.677 → 0.675) → robustness confirmed: the omission was a random 1/6 split and did not bias the prior numbers. Per-city estimates are no longer flagged provisional-pending-shard_2.
- fair_gap snap hygiene (R8) — the converged subset driving the headline is tightly snapped (median 15s, p90 27s, max 519s, 0% > 1800s); the 673,462s outlier lives in the non-converged tail and does not touch the headline. fair_gap is sound. Recommended non-blocking hygiene: drop entries with
snap_dt_s > 120sin a future re-run.
Open items — for the live dry rehearsal
- Net-of-cost recovery (the binding gate) — does the +0.02 entry premium survive taker fee + slippage on our actual reaction latency? These lags are book-reprice times, not our fill confirmations.
- Ex-ante predictor skill — out-of-sample skill of the city/session/METAR convergence predictor that replaces the non-implementable lag gate.
Item (3) from the draft — re-run shard_2 — is closed: the full 28,664-entry population is measured and aggregated; the per-city ranking is final on complete data.
Repo pointers
Source files
- Full report (final, 6-shard, N=28,664):
research/reports/weather-lag-distribution-2026-06-19.md- Aggregate artifact:
research/copytrade/top50/lag/lag_distribution.json- Pipeline (per-shard measurers + aggregator):
research/copytrade/top50/lag/(measure_shard_N.py,aggregate_lag.py) — see wx-lag-pipeline-bug for the inlined-fetch / no-lib.pygotcha
Related
- pmv2-reorg-and-weather-as-2026-06-22 — the pmv2 reorg + the
weather_aslane this study underpins - weather-gate-winlabel-2026-06-19 — the WIN-label re-score downstream of the gate this study motivated
- weather-top50-decypher-2026-06-19 — the v2 decypher whose Section 2 fillability open items this study settles; its “plausible but not proven” ≤0.95 edge is now resolved to weak-GO with a premium thesis
- wx-lag-pipeline-bug — companion data-integrity note: the 120/180 convergence-horizon baking bug and the aggregator-authoritative fix that made these numbers coherent
- weather-bet-live-status — live
wx-signalengine + the dry rehearsal that must settle net-of-cost recovery and the ex-ante predictor’s OOS skill - weather-strategy-validated-plan — the GO plan (late-settlement / repricing-lag); this study sizes the repricing-lag layer honestly (~34% realized hit, premium entry)
- weather-bet — project overview