The v2 copy-trade reconnaissance: the official Polymarket weather leaderboard (not the small weatherbot.fi board the v1 note used) → 87 candidates → 65-wallet activity gate → top 50 re-ranked by ROI, each decyphered by an LLM agent, with a 10-wallet CLOB fillability reconstruction and an adversarial critic pass (12 issues, all addressed). Headline: the money is in capacity, not return — whales win on turnover + execution at the close, not edge-per-dollar. The own-edge / fillability story is the load-bearing part and the honest answer leans NEGATIVE: the ≤0.95 fillable edge is plausible but not proven, and several specific claims are contradicted by the cohort’s own trade samples. Full report: research/reports/weather-top50-decypher.md.

UPDATE 2026-06-22 — head of the chain feeding the weather_as lane (own repo, pmv2 reorg)

This decypher seeds the fillability chain now packaged as the Asian-weather (weather_as) producer: source weather_as, subject pmv2.order.weather_as.entry, PRIVATE repo pmv2-weather-as-algo at /coding/pmv2/as_weather_algo (committed e9415dc). The research record stays in the renamed US repo pmv2-weather-us-algo at /coding/pmv2/us_weather_algoresearch/... paths here resolve there, not the old weather_bet. One framing update: the design-only copytrade source that “rides the existing WEATHER_SIGNALS stream” is superseded by the pmv2 reorg — copy-trade is now its own cohort_algo producer on pmv2.order.cohort.entry under a shared pmv2-contracts envelope. Full reorg + current state: pmv2-reorg-and-weather-as-2026-06-22.

For Agents

v2 of copytrade-top-weather-traders — supersedes/extends it with a far larger, gated, critic-corrected sample. Three things to carry forward: (1) the pool deploys 1.52M reconstructed PnL (8.7% cap-weighted, 11.5% median ROI), and ROI inversely correlates with absorbable capital — late-settlement whales hold 70% of capital at the lowest ROI (6.2%), longshots have the highest ROI (55.5%) on trivial size. (2) The fillability edge is NOT answered affirmatively — it is reframed as the experiment the dry rehearsal must settle (discount-to-fair vs residual risk-taking). (3) Copy-trade doctrine is “detect, don’t mirror” — we lose the latency race, so detect the same METAR lock from our own feed earlier and use cohort appearance as confirmation. A design-only copytrade NATS signal source spec exists (rides the existing WEATHER_SIGNALS stream). Every per-leg PnL is FIFO-inferred — there are no settlement outcomes in the raw data — so treat micro-stats as reconstructions.

Method

  • Source: official Polymarket weather leaderboard (category=weather), month + all-time union → 87 candidate wallets.
  • Activity gate → 65 passed (gate_thresholds in top50_profiles.json):
    • min_resolved_markets = 20 (raised from a phase-1 floor of 10 to kill thin samples).
    • min_capital_deployed = $10,000 (anchor ROI on a real book).
    • small_capital_ceiling = $50,000; below it a small-cap inflation cross-check rejects wallets whose reconstructed PnL exceeds the leaderboard PnL (max_pnl_to_leaderboard_ratio_when_smallcap = 4.0; killed to010 at 18.8×). Above $50k the ratio is deliberately not applied (large genuine winners show high ratios only because the public board exposes a partial/stale slice).
  • Re-rank the survivors by ROI → top 50. One LLM agent decyphered each of the 50 (interp/*.json: archetype, edge thesis, entry/exit rules, capacity, replicability_score, confidence, caveats).
  • Fillability: targeted CLOB price-history reconstruction on a 10-wallet cohort (fillability.json), each capped at ≤40 cheap entries, ≤15 displayed samples.
  • Adversarial critic: 12 issues raised, all addressed (table in The critic’s 12 issues).

Headline finding — capacity, not return

The 50-wallet pool deploys 1,522,590 reconstructed realized PnL: 8.73% capital-weighted, 11.51% median per-wallet ROI. The dollar-weighted return sits far below the median because deployment is dominated by a handful of eight-figure carry books. ROI and absorbable capacity trade off almost perfectly inversely.

ArchetypenTotal capitalTotal PnLMedian ROI
late-settlement-whale22$12.21M$752.9k6.2%
repricing-lag-nowcaster18$3.85M$409.7k15.4%
longshot-tail-lottery4$132.4k$65.6k55.5%
round-tripper-sell-the-pop3$867.3k$253.1k27.9%
early-forecast-value2$194.3k$28.2k16.7%
favorite-carry / round-trip (single, unverified)1$186.1k$13.1k7.1%

By median ROI: longshot (55.5%) > round-trip (27.9%) > early-forecast (16.7%) > repricing-lag (15.4%) > single (7.1%) > late-whale (6.2%, the floor). The capital ranking is the inverse.

  • late-settlement-whale (n=22) — buy the near-certain side (mostly the No leg of eliminated buckets, or the locked Yes landing bucket) at very high price within hours of (or after) the obs-window close, hold to settlement for the last fraction of a cent. Holds 70% of pool capital → 49% of pool PnL at the lowest ROI. Contains every eight-figure book (HondaCivic 1.59M, NoonienSoong 1.03M, vip68 0.82M). The edge is capital turnover across hundreds of markets, not per-trade margin. Exemplar NoonienSoong (0x38cc1d1f…): 1194 resolved markets, 99.06% win rate, median entry 0.993 at T−2.85h, ~266.8k reconstructed (the single largest in the set).
  • repricing-lag-nowcaster (n=18) — buy a cheap-to-mid bucket intraday and ride a sharp upward reprice as observed temps converge, exploiting the lag between METAR making the outcome obvious and the book catching up. This is the fillable engine for us. Exemplars Happening9014 (25.9% ROI, ~39-min median lag) and 0x686880… (46.3% ROI, RTR 0.75).
  • longshot-tail-lottery (n=4) — systematically buy many cheap (<0.15) OTM buckets, hold; winners pay 1/p. Highest ROI (55.5%) but it does NOT scale — cheap legs have no depth (books are $13–60k). Lowest replicability.
  • round-tripper-sell-the-pop (n=3) — buy the near-certain No leg, sell minutes later into the 0.998–0.999 microstructure pop. 27.9% median, but n=3 and two of the three are also classed as carry whales → not a robust archetype-level result (the v1 “sell-the-pop is the WORST / REFUTED” verdict is over-stated; the defensible claim is narrower — fast monetization of the obs-lock beat slow monetization on these specific large books).

This confirms (and sharpens) the v1 thesis

v1 said the whales win on nowcasting + execution at the close, not buy-cheap/sell-pop. v2 quantifies it on a bigger gated sample: the whales win on turnover + execution at the close, not edge-per-dollar — they hold the most capital at the lowest ROI. The high-ROI engines (longshot, round-trip) run trivial or crowded size and don’t scale.

Own-edge / fillability — the load-bearing part, and it leans NEGATIVE

The ≤0.95 fillable edge is PLAUSIBLE BUT NOT PROVEN — do NOT oversell it

Load-bearing caveat. The 99c carry that dominates leaderboard PnL is NOT fillable for us. The candidate edge is the repricing-lag layer (cheap-Yes 0.55–0.85 after the METAR lock, before the book reprices), but on this data it is thin and internally conflicted:

  • Only 3 of 10 cohort wallets clear the latency bar (HondaCivic, vip68, Happening9014) — a thin basis to “answer” the gate.
  • Reprice-lag is right-skewed. The population medians that justified a 20-min gate (24.9 / 50.2 / 39.4 min) are dragged up by 80–120 min outliers; the typical observed lag the engine sees live is <15 min (HondaCivic sample-median 4.7 min, 11/15 under 15). A 20-min gate would fire on a condition the engine rarely encounters. Only Happening9014’s typical lag genuinely clears the window.
  • The “cheap fill on a laggy book” claim is contradicted trade-by-trade. Winners paid above the contemporaneous book in 11–13 of 15 samples (HondaCivic median +0.046; two extreme fills paid ~0.93 while the yes-book sat at 0.74 / 0.10). This is either a measurement artifact (paid vs yes_book may reference different sides/instants — snap_dt_s is non-zero) or the trader buying into residual uncertainty (risk-taking), not harvesting a settled lag. The fill band is real as a price the winners paid; it is not established as a discount to fair.
  • Even a yes wallet is not “locked”: vip68’s convergence rate is 70% → ~1 in 3 of its sampled cheap buys did not converge.

The dry-rehearsal gate is NOT answered affirmatively. It is reframed as the experiment that decides whether the 0.84–0.91 band is a discount-to-fair or residual risk-taking.

UPDATE — settled by the unbiased full-population study (weather-lag-distribution-2026-06-19): measured on all 28,664 cheap entries (no highest-paid cap), the verdict is weak-GO. The discount-vs-risk-taking question is resolved: winners pay a ~2-tick PREMIUM (median fair_gap +0.02, 75% above the same-instant book), so this is a convergence-timing edge bought slightly rich, not a discount-to-fair. Realized hit rate is ~34% (not 63–67%) because ~49% never reprice to 0.97. The binding gate becomes net-of-cost recovery of that premium, answerable only in the live rehearsal.

Two axes the binary verdict conflates

The single our_latency_feasible yes/no collapses two things that should be reported separately:

  • Per-entry fillability — is this cheap entry reachable in our reaction window?
  • Book-fraction cheapness (pct_cheap) — what share of the wallet’s buys are cheap-≤0.95 at all?

They diverge: meropi has a no verdict but an 82% per-entry yes-fraction (higher than HondaCivic’s 75%) — its no is driven purely by pct_cheap = 2.0% (almost all its capital is 0.99+ carry, so there’s little cheap volume to copy even though the cheap entries it does make are individually feasible). Report the two axes separately; do not read the binary verdict as a clean go/no-go.

Best slice if the rehearsal is run

highest-temperature buckets, post-local-peak, ~4–9h pre-resolution, on specific liquid dailies — London (most-cited), Munich/Milan/Paris, the Asian-hours cluster (Shanghai/Chengdu/Chongqing/Shenzhen/Taipei/Seoul/Tokyo), and major US books (NYC/Atlanta/Dallas). Select cities by observed lag, not by US/intl label (the “disproportionately international” framing was withdrawn — NoonienSoong is US-majority us:1792/intl:925; HondaCivic is london/nyc/munich). Size small and wide: 1–2k on a high-conviction leg, capacity from breadth not depth (per-market capacity is low single-digit thousands). Gate on a robust lower quantile of observed lag (e.g. 25th pct), not the median, and not a hard 20-min line; suppress sub-7-min (carry-whale zone); cap entry at 0.95 and instrument whether the 0.84–0.91 band is actually a discount to fair.

Copy-trade — “detect, don’t mirror”

We lose the latency race, so the doctrine is: detect the same physical lock from our own per-minute METAR feed earlier, and use cohort appearance as confirmation — not mirror fills into a thin closing book. The shortlist is split into two tiers (a critic correction reconciling it with the own-edge section, which previously contradicted it by ranking carry whales as top copy targets).

Tier A — copyable now (survives the ≤0.95 fillability filter):

HandleAddressArchetypeNote
Happening90140x2d0fbf6995692cc56f21171a9993be04436a81aarepricing-lag (late tilt)Cleanest yes — only wallet whose typical lag (~39 min) clears our window.
vip680x8d0930676d559cc8fb7d8af0c555791c1820143frepricing-lagyes but 70% convergence (size for the 1-in-3 miss); Asian-hours.
EngineOfHondaCivic0x08cf0b0fec3d42d9920bb0dfbc49fde635088cbclate-settlement + lagmarginal (lag ~7 min) — partial-fill only; demoted from v1’s #3 copy slot.
WeatherTraderBot0xacc8e9dcabf9d65a5c78e3bec6941ed53a2b7d08repricing-lag + near-certaintyClean NYC/London sleeves; risk = we compete directly with this bot.
0x686880…51980x686880eea810fa9141be46a1acec9eee41755198repricing-lag-nowcaster46.3% ROI, RTR 0.75, fast-converging, confidence: high.
0bot0xf787b8a3a4c6990f48614a0b9af8422e3e7a1b63repricing-lag (obs convergence)US-focused, 5–6 liquid cities, clean late-day convergence.
cmcbrown0x182addd8096bbd0fc0f609a6e3887f67fe82f847repricing-lag-nowcasterCleanest single-archetype signal; good Asian-hours confirmer.
Pityok10x82f3458a6a4bec624e692c92d9ed14597f509f82repricing-lag + late lockPure physical-obs edge; our per-minute latency likely beats a manual book-hitter.

Tier B — confirmation-only carry whales (NOT a ≤0.95 fill target):

HandleAddressWhy Tier B
NoonienSoong0x38cc1d1f95d12039324809d8bb6ca6da6cbef88eBest edge×capacity in the pool ($267k realized) but fillability no (1.4% cheap, ~1-min typical lag) and a naive carry is negative-EV. Cohort-confirmation only.
1-800-LIQUIDITY0x584eee598b341109592b985c1a253ab044fa090fHighest raw replicability (0.82) but pure final-minutes 0.99 carry (~5.7 min pre-resolution). We lose that latency race.

Two wallets corrected from v1

  • BeefSlayer (0x331bf91c…) — v1 tagged it a late-settlement whale on a ≤0.99 near-certainty trigger; its profile contradicts that (median entry 0.5461, win rate 0.5713, US-majority) → it is mid-range directional / repricing-lag risk, removed from the near-certainty detector. Eligible only for the repricing-lag detector with a mid-price model, sized small.
  • ColdMath (0x594edb91…) — not a clean negative control. Reconstructed realized 136,018 leaderboard (a ~170× divergence) on HTR 1.0 / entry 0.0685 → the signature of a held-to-resolution longshot book whose settled wins the FIFO recon fails to capture. Treat as a reconstruction-failure case, not “zero realized edge.”

copytrade NATS signal source (design only)

A new source (not a new stream) riding the existing JetStream WEATHER_SIGNALS / subject weather.signals, per docs/signal-schema.md (strategy-agnostic envelope, consumers route on (source, action)). It is an independent confirmer, not a mirror: it watches the Tier A wallets’ on-chain positions and emits a cohort-first-mover signal when K-of-N targets open a fresh position on a market our wx-signal thesis also covers, within rolling window W (e.g. K=2 in W=20 min; K=1 allowed only for the highest-confidence Tier A wallets as a lower-weight lead).

  • Source-routed: source = "copytrade", no consumer change beyond adding the route.
  • Poll each target via the Polymarket Data API every 30–60s (edge persists minutes-to-hours; sub-second is unnecessary and rate-limit-hostile) with an after/cursor watermark; local seen-set of (wallet, fill id) so re-polls are idempotent.
  • max_price ≤ 0.95 — the cohort’s worst observed entry clamped to 0.95, so we never chase past where the cohort filled.
  • Dedup: id = UUIDv5 of "{source}|{market_slug}|{action}"; Nats-Msg-Id = that id so server + app dedup align. At most one copytrade BUY per (market, BUY); a future SELL/exit cohort uses the same envelope with action="SELL" and a distinct id. Additive — the existing publisher hardening (bounded retry on ack-confirmed publish, --smoke) applies unchanged.

Reconcile with the live-status note

weather-bet-live-status records that a 2nd source (cohort first-mover) already rides the same stream (it names subject cohort.signals). This report’s copytrade spec is the design-of-record for that source and routes it onto weather.signals with source="copytrade". The subject naming is the one detail to confirm against the live deploy; the dedup discipline (uuid5("{source}|{market_slug}|{action}") = Nats-Msg-Id) matches the live engine exactly.

Critical caveats

All per-leg PnL is FIFO-inferred, not measured

The raw files carry no settlement outcomes and no per-trade realized PnL. Reconstructed PnL therefore diverges from the leaderboard both ways: ColdMath ~170× under; the longshot books ~1/3 of leaderboard (Joe 51.8k; Henry 47.3k). Micro-stats — “91% of trades ≥0.95”, “~0.7c/turn”, convergence rates, reprice-lag figures — are reconstructions, not measurements. A naive NoonienSoong carry is even slightly negative-EV at its stated 0.9906 win rate (≈ −0.0024/share), so its positive reconstructed PnL must come from legs the FIFO attributes, not the headline behavior.

Truncation + biased fillability samples

  • The 3,500-trade/wallet API cap truncates high-volume whales (HondaCivic, opopv, vip68 all show >3,000 buys) → per-wallet totals, RTR, bucket splits are lower bounds, reconstructed PnL a partial slice.
  • The fillability cohort caps n_entries_analyzed at 40 (bhuumi only 11) and shows ≤15 samples — the highest-paid subset, which biases every sample-derived statistic (the paid>book and lag-distribution findings above).

The critic’s 12 issues

All 12 were raised and addressed. The high-severity ones that shaped the conclusions:

#ConcernDisposition
1”Cheap fill on laggy book” contradicted (paid > book 11–13/15)Discount-to-fair assertion withdrawn pending paid-vs-yes_book reconciliation.
220-min lag gate calibrated on a right-skewed medianRe-spec’d to a robust lower quantile; hard 20-min line dropped.
3n=3 fillable wallets on a 40-entry cap is too thinStated explicitly; gate treated as an experiment, not a settled affirmative; two axes separated.
5BeefSlayer mis-labeled (0.55 entry, 0.57 wr)Re-classified mid-range directional; removed from near-certainty trigger.
6Copy shortlist contradicted own-edge (ranked carry whales as top copy targets)Tier A / Tier B split; NoonienSoong + 1-800-LIQUIDITY → confirmation-only.
8ColdMath called a negative control despite 170× recon gapFlagged as reconstruction-failure, not a negative control.
9”sell-the-pop REFUTED as worst” over-stated (n=3, two are carry whales)“REFUTED” withdrawn; claim narrowed.

Open items the data cannot close (carry into the dry rehearsal): whether 0.84–0.91 is discount-to-fair or risk-taking; the full reprice-lag distribution (to set the quantile gate); settlement-truthed per-leg PnL to replace FIFO; whether n=3 yes wallets generalizes.

Repo pointers

Source files

  • Full report: research/reports/weather-top50-decypher.md
  • 50 wallet profiles + gate config: research/copytrade/top50/top50_profiles.json
  • 10-wallet fillability cohort: research/copytrade/top50/fillability.json
  • Per-wallet interps: research/copytrade/top50/interp/*.json
  • Signal schema (the envelope copytrade rides): docs/signal-schema.md