For Agents
The central strategic correction for this project and the Hormuz-specific data hazard. The market resolves on PortWatch number, not on physical reality — so the primary data source must be PortWatch itself, NOT a third-party AIS feed. Using a different AIS provider introduces basis risk (right about reality, wrong about settlement), which is severe in the Strait of Hormuz because it is the global epicenter of AIS evasion (dark vessels, GPS jamming, AIS spoofing). This note records the no-yes-man correction to the original premise and the verified evasion data.
The single most important strategic decision in the project, plus the reason third-party ship data is a trap here.
The Correction to the Original Premise
Model PortWatch, not reality
The original idea was “use some ship data API (AIS) to count ships.” That is subtly wrong for this market. The market resolves on PortWatch published number, NOT on physical reality. Counting ships yourself with a different third-party AIS provider (Datalastic, MarineTraffic, Kpler, …) means you are measuring a different quantity than the one that settles the contract. You can be right about reality and wrong about settlement. That gap is basis risk, and here it is large.
The correct design follows directly:
- Primary data source = PortWatch itself. Model the PortWatch transit-call series and its weekly publication mechanics (see hormuz-bet-portwatch-data). The full 2019-to-present history is free, which is what makes the backtest tractable (see hormuz-bet-backtesting).
- A third-party real-time AIS feed is only useful for the harder Phase-2 play: predicting PortWatch number before its weekly Tuesday publication. There it is a leading indicator of the settlement number, never the settlement number itself. See hormuz-bet-ais-providers.
Dark-Vessel / AIS Basis Risk (Hormuz-specific, severe)
The Strait of Hormuz is the global epicenter of AIS evasion
Mass “going dark”, GPS jamming, and AIS spoofing by Iranian / sanctioned tankers are routine here. A third-party AIS feed is doubly unreliable in this strait, which is the second reason (after basis risk) not to base settlement-tracking on one.
Verified data points:
- Windward SAR study (2026-05-05): of 167 vessels detected by satellite radar, 146 were dark (~87%) — i.e. not broadcasting AIS.
- GPS jamming (2026-03-01): >1,100 vessels reported GPS-jammed in the Gulf within a 24-hour window.
- PortWatch own caveat: counts may be distorted by jamming, spoofing, and dark vessels (see hormuz-bet-portwatch-data).
The Subtlety: This Hazard Helps Phase 1
The evasion problem cuts an interesting way. Because PortWatch is what settles the market, PortWatch own measurement quirks (how it does or does not catch dark/jammed/spoofed vessels) are part of the signal, not noise to be corrected. A third-party feed that “more accurately” counts reality would actually increase your settlement error. So:
- For Phase 1, you do not want a better ground-truth count; you want to model PortWatch number including its biases. The settlement series is the target, warts and all.
- For Phase 2, a third-party feed only helps insofar as it predicts how PortWatch will count, which means you must model the relationship between the feed and PortWatch number (a learned mapping), not assume they are equal.
Non-Stationary Wartime Regime
Compounding everything: the series is in a structural break right now (n_total ~3-10/day in May 2026 vs ~55-75/day in 2019). Evasion intensity, jamming, and the war regime all move together, so the PortWatch-vs-reality gap is itself time-varying. Backtests must respect the regime break and not assume a stationary mapping. See Current Regime (as of 2026-06) and hormuz-bet-backtesting.
Strategy C Kill-Tests PASSED → GO (conditional) (2026-06-08)
Strategy C is worth building — conditionally
Strategy C is the Phase-2 play stated sharply: predict PortWatch weekly count ahead of its Tuesday publish via live AIS, and trade the weekly markets before they reprice. Before building any AIS collector, two zero-cost paper kill-tests were run on real free data (the PortWatch daily series + Polymarket CLOB price history) to decide if C is even worth pursuing. Both passed. Verdict: GO, conditional on one remaining gate (can live AIS actually track PortWatch’s daily count — see below).
This is distinct from the Phase-1 backtest verdict. The hormuz-bet-build-spec1 live run came back INCONCLUSIVE for trading the count markets on the public PortWatch number the market already sees. Strategy C is the opposite informational bet: get the number before the market does, via AIS. The kill-tests below test exactly that thesis.
Test 1 — outcomes ARE sensitive to pending-day info (the bet is live). The weekly-total Hormuz markets resolved within 2–4 ships of the bucket boundary: week of May 11 had a 7-day sum of 44 vs the 40-boundary → settled in the 40-59 bucket; week of May 18 summed 42 vs 40. The running cumulative total crossed the boundary around Day 6 of the 7-day window — so the outcome is effectively locked in before PortWatch ever publishes. Knowing the pending days’ transits changes the answer, which is the precondition for any nowcast edge.
CRITICAL EXCLUSION — the max-daily THRESHOLD markets have ZERO AIS edge
The single-day threshold legs (40+/60+/80+ on any day) are pure regime-change bets in the current crisis regime: the daily count never exceeds 29 (April 18 was the outlier at 29). Crossing 40/60/80 requires a regime break, not a marginal day. Live AIS gives you no edge on these — exclude the threshold markets from Strategy C entirely. Strategy C is the weekly-TOTAL (and 7-day-MA) bucket markets only.
Test 2 — the info is NOT priced (the key finding; refutes the “it’s already priced” worry). On both markets with usable CLOB price data, the eventual winning bucket stayed cheap through the window and past the weekend, then jumped only the day AFTER the Tuesday PortWatch drop:
- Week of May 11: the total crossed into 40-59 on Sat May 16; the market still priced that bucket at ~19.5% and did not reprice to 90%+ until May 20 — the day after the Tue May 19 publish. A ~4-day lag; ~5x gross return if entered May 16 on AIS knowledge.
- June 1: opened flat at ~52.5% the day PortWatch published the prior week; an AIS reader already knowing the answer would have captured ~+90% in 5 days.
Prices REACT to PortWatch — they do not ANTICIPATE it
This is the empirical heart of Strategy C. The market does not front-run the weekly number; it re-rates only after the Tuesday release. An AIS reader who knows the count ahead of the publish is trading against a book that is still pricing the prior week. That is the exploitable lag.
The reframe (important). AIS gives the actual daily transit count in real time, not a noisy estimate. So the success bar is “count within the ~2–4 ship margin,” NOT “within a bucket width.” That is a sharp, testable target — and it converts a vague “predict the bucket” goal into a concrete tracking-accuracy spec.
Honest caveats — what is still UNPROVEN
- Tiny sample: only n = 2 markets had usable price data. This is directional, not statistically robust — need 10+ markets for confidence, and ~25+ drops / several months for a robust trading verdict.
- Small capacity: thin count books, realistic ~1k per clip (see hormuz-bet-liquidity).
- Regime-dependent: the tight margins are a feature of the current near-boundary low-count wartime regime; they may not survive a regime shift (see the non-stationary-regime section above).
- The one gate the kill-test could NOT test with free data is now THE question: can our live AIS count actually track PortWatch’s daily
n_totalwithin ~2–4 ships? Terrestrial AIS may undercount mid-strait (PortWatch is satellite-fed), so a satellite coverage-probe is now first-class, not optional (see hormuz-bet-ais-providers).
Red-team reconciliation. An independent pre-mortem had argued Strategy C was likely dead (info priced, unmonetizable, underpowered). Test 2 empirically REFUTED the strongest objection (priced). The pre-mortem’s valid surviving points are folded in as the conditions on the GO: the AIS-accuracy gate, small-sample / low power, operational data-integrity for the pilot, capacity, and regime-dependence. None of those are fatal; all are now explicit gates rather than unknowns.
Next — Milestone 1 = an AIS pilot. Single success criterion: “does our live AIS count track PortWatch’s daily n_total within ~2–4 ships?” Build: terrestrial aisstream (free) + a sparse satellite coverage-probe (~EUR 40–75/mo), forward collection, deployed on the polymarket-infra fleet via @deploy. Design spec: docs/superpowers/specs/2026-06-08-hormuz-strategy-c-ais-calibration-design.md (updated with this kill-test result).
For Agents — Strategy C state machine
Phase-1 count-market trading = INCONCLUSIVE (shelved as a standalone edge; engine kept). Strategy C (Phase-2 AIS nowcast of the weekly TOTAL/MA buckets) = GO (conditional). The single open gate is AIS-vs-PortWatch tracking accuracy (≤ ~2–4 ships). Exclude the daily-threshold markets from C (zero AIS edge). Proceed to Milestone 1 AIS pilot; do not deploy capital until the tracking gate passes and the sample grows to ~25+ drops.
Related
- hormuz-bet — project index (the one-thing-to-internalise callout summarizes this note).
- hormuz-bet-portwatch-data — the settlement series this argues you should model directly.
- hormuz-bet-resolution-mechanics — confirms settlement is the PortWatch number via UMA.
- hormuz-bet-ais-providers — the Phase-2 leading-indicator feeds (terrestrial-only, doubly exposed mid-strait); the satellite coverage-probe lives here.
- hormuz-bet-backtesting — why the regime break and the PortWatch-as-target framing shape the validation; the Strategy C kill-tests are recorded there too.
- hormuz-bet-build-spec1 — the Phase-1 INCONCLUSIVE backtest that Strategy C is the informational inverse of.