For Agents

Reverse-chronological session log. Newest entries at top, grouped by date (## YYYY-MM-DD). Each bullet: one piece of work, short summary, wikilinks to docs touched. Updated by obsidian-documenter on every project doc write. Read by historian at bootstrap.

2026-07-22

  • MAKER-SIM GATE — BUILD PHASE COMPLETE, READY-FOR-GATE-RUN (subagent-driven; gate runs blocked on infra-VM data fetches). The offline maker-sim gate harness is built + per-task-reviewed + final-reviewed, deterministic end-to-end — asks whether a maker/resting-order posture recovers the edge taker fills lose to last-look. No gate has run yet; lane stays SHELVED regardless. Built: crates/analysis/src/bin/makersim.rs (quote engine on lagged snapshots/pulls; bracketed fill engine [pessimistic: adverse-certain/at-quote φ=0.25]; inventory cap ±3 hold-to-settlement; day-clustered bootstrap z; fixed 48-config grid; 15/15 tests, clippy/fmt clean, deterministic) + three fetch scripts (makersim_resolve_markets/makersim_clob_winners/makersim_tape_backfill + makersim_common DRY extraction), all fixture-proven offline. Review caught 2 bugs in the plan’s OWN reference code: (a) unsorted-prints self-inconsistency → fixed with internal stable sort; (b) unresolved-CLOB-market silent Down-label → fixed with exactly-one-winner guard. Pre-registration (before any data existed): 0-print (untraded) windows COUNT in the bootstrap denominator (survivorship rationale); all 7 accumulated minors triaged DEFER. Fee verdict (Task 4): live 5m market advertises maker_base_fee=1000 == taker_base_feemaker_fee_zero: falsethe 0.07-charged variant is the PRIMARY gating number (a maker posture does NOT escape the fee; data/makersim/fee_verification.md). Environment gotchas (durable): (1) *.polymarket.com Cloudflare TLS-resets this workstation since ~11:00Z 07-22 (after ~3k morning API calls) AND blocks agent sandboxes → all live fetches rerouted to deploy/infra VM (babylon 1243 staged handoff; scripts resumable); (2) sqlx + transaction-mode pooler (:6543) breaks prepared statements at process rerun → gate runs must use session-mode pooler (:5432); same latent issue exists in walkforward.rs. State: waiting on deploy’s VM runs (markets→winners→tape, ~10-25k polite calls, tape ~0.5-1.5GB) → Task 8 sanity STOP-gate → Task 9 calib gate (KILL if no lag==1 config reaches z≥3.07 pessimistic 0.07-fee) → Task 10 single confirmatory run. A full pass clears only the gate — a separate Phase-2 decision is required before any live posture; lane stays shelved with observe recording. Artifacts: plan docs/superpowers/plans/2026-07-22-maker-sim-gate.md; spec v2+addendum; ledger .superpowers/sdd/progress.md. — crypto-shortterm-maker-sim-gate-build-complete-2026-07-22, crypto-shortterm-maker-sim-gate-spec-grounding-2026-07-22, crypto-shortterm-w0-fee-rebate-audit-2026-07-22, crypto-shortterm-15m-resite-premise-falsified-2026-07-21, crypto-shortterm-fill-space-closed-15m-pivot-proposal-2026-07-18, crypto-shortterm-backtesting, direct-db-ipv6-use-pooler
  • MAKER-SIM GATE SPEC GROUNDING (v2, pending Andras review) — 4 verified 5m data facts + 5 spec corrections. Grounded the proposed offline maker-sim gate study (docs/superpowers/specs/2026-07-22-maker-sim-gate-design.md v2; asks whether a maker/resting-order posture recovers the edge taker fills lose to last-look) — nothing built, lane still shelved. Data facts: (1) 5m near-tie rate 53% — 53.25% of btc 5m windows settle within ±5bps of strike, 78.39% within ±10bps (n=6,043, 21d) vs 22–31% at 15m → on-chain resolved_up unusable as a 5m scoring label, CLOB tokens[].winner backfill mandatory for ALL scored 5m windows. (2) gamma slug/condition_ids query BROKEN for 5m updown (returns []; works for 15m) → resolve via emitted_signals.condition_id → CLOB /markets/{condition_id} (confirms slug btc-updown-5m-{window_id}). (3) data-api /trades?market={conditionId} verified as 5m tape — full tape, ~2k/page paging, unix-second ts, taker-side only (no maker flag); sampled window 4,786 prints; backfill budget 15–35M rows for ~7.5k windows. (4) pm_book_snaps 5m cadence ≈1.3s (~228 snaps/window, sizes 100% populated; btc 5m spans 06-20→present, 8,259 windows, complete). Spec corrections in v2: 7-tick/+0.07 reprice figure lives in the A0 scope doc (07-13 follow-up), NOT latency_budget.rs; maker-rewards citation was wrong (competitor’s were TAKER rebates, maker-rewards applicability to 5m unverified); sim latency tiers → cadence multiples {1.3,2.6,3.9}s (no sub-second tiers from 1.3s data). Babylon: PM answered #1189 (sub-500ms feasible after fee-param caching, sub-300ms needs WS book), built fee fix fix/venue-cash-fee-coefficient @eee211e (books 0.07, removes 2 fee RTTs, closes zero-fee fail-soft; off PROD ca07ba5); crypto asked PM to verify 0.07 on own cash before rolling (#1239); 5,322-row historical fee backfill gated on Andras. — crypto-shortterm-maker-sim-gate-spec-grounding-2026-07-22, crypto-shortterm-w0-fee-rebate-audit-2026-07-22, crypto-shortterm-15m-resite-premise-falsified-2026-07-21
  • W0 FEE/REBATE AUDIT CLOSED — two empirical findings, verdict UNCHANGED (lane stays shelved). Ran the one free W-item from the 07-21 shelve verdict (also answers the standing “do WE accrue taker rebates?” follow-up from the competitor refutation). (1) pmv2 FEE-BOOKING BUG (fleet-wide): pmv2_autotrade_orders.fee_usd = exactly 0.10·p·(1−p)·size on all 5,322 fee>0 rows, but the venue’s ACTUAL charge — fitted on 3,399 of our own live BUY fills via the data-api usdcSize − price·size identity — is exactly 0.07 (median=p10=p90=0.0700, ZERO dispersion; matches the competitor’s 190k-fill fit). Booked fees 43% too highrealized_pnl_usd systematically pessimistic for EVERY pmv2-booked strategy; our lane 368 true (−0.92¢→−0.47¢/share). Our own observe::fee (0.07) was always correct — the validated net-edge gate’s fee term needs NO change. Reported to positionmanager as babylon #1233. (2) TAKER REBATES real but TINY for us: 16 TAKER_REBATE rows = 22.55, balu 1–2 credits + two ~1,163 modeled fee ≈ +0.05¢/share — 5× below the earlier +0.24¢ bound; maker-rebate $0 (pure takers); credits not represented anywhere in pmv2’s orders table. Net effect: true economics less bad than booked (−0.47¢ + 0.05¢ rebate) but still net-negative on real labels → the structural last-look ceiling is unaffected by a fee re-base → lane stays shelved, observe keeps recording. Provenance: data-api activity paging (scratchpad rebate_tally.py/fee_fit.py), wallets from public.pmv2_wallets (all dry_run); artifact data/w0_rebate_fee_reconciliation_2026-07-21.csv; HANDOVER.md 07-21/22 updated. — crypto-shortterm-w0-fee-rebate-audit-2026-07-22, crypto-shortterm-15m-resite-premise-falsified-2026-07-21, crypto-shortterm-competitor-wowitsamazing-refuted-2026-07-03, crypto-shortterm-polymarket-data-api-gotchas, crypto-shortterm-strategy-design, crypto-shortterm-lane-state-adverse-fill-2026-07-05

2026-07-21

  • GROUNDING-PASS VERDICT — the 15m re-site PREMISE is FALSIFIED; the pre-committed EXIT RULE FIRED; recommendation = SHELVE the trading lane (pending Andras). SUPERSEDES the 07-18 W0–W3 proposal’s conclusion. Andras asked for a grounding pass on the approved slow-clock re-site spec (“make sure it will actually produce reliable data”); three read-only passes (SQL via pgq.sh + Polymarket gamma/CLOB) falsified the core premise before ANY build.No live state changed; observe (5m+15m) keeps recording. PASS A — conviction gate SATURATES 15m into a dead ask zone: |p_fair−0.5|≥0.30 ∧ τ≤300s qualifies ~98% of 1,733 windows/asset, first-qualifying asks 0.96–0.97 — worse than the 5m τ∈[150,300) trap (0.90–0.93); by the time the signal is confident the market has fully priced it. PASS B — the competitor’s cushion band ([5,10)bps, τ≤300s) IS priced at his-era levels on OUR books (btc15 median ask 0.875 n=878; eth15 0.800 n=788; his bucket ~0.820) — the only non-dead pocket found in the whole W0–W3 program → motivated Pass C. PASS C (DECISIVE) — true CLOB label backfill of all 3,131 first-trigger windows (slug {asset}-updown-15m-{window_id}conditionIdtokens[].winner, 0.03% miss): EVERY cushion bucket net-negative on BOTH assets; pooled day-clustered own-ask EV −1.14¢ 95% CI [−3.0,+0.5]; calibration segment +1.6¢ in-sample flips to −2.2¢ OOS (n=1,209) on both assets independently. THE NEAR-BREAKEVEN ILLUSION WAS LABEL ERROR: on-chain resolved_up disagrees with the true CLOB winner on 18–23% of near-ties; true near-tie win 0.63–0.69 (~coinflip), NOT 0.77 — the same on-chain-optimistic mask manufactured a fake tradeable 15m pocket; the competitor’s +3.8¢ tape prior was a selection ceiling, our books price his pockets fee-negative. DURABLE METHOD FINDINGS: (a) 15m label trap — near-ties are 22–31% of late-window populations, outcomes stores no condition_id → a CLOB backfill is MANDATORY for any 15m scoring (off near-ties CLOB vs sign(settle−k) agree 99.1%); (b) 07-09 fill-regime BREAK — fill rate 4–7% → 13.6% → 27–35% in stages (not a smooth trend); pm_fak_attempts spans 07-02→07-16 (best_ask_at_snap 100% from 07-03) and lives in pmv2’s schema → the [[crypto-shortterm-backtesting|07-03 harness spec’s crypto_shortterm qualifier is WRONG]] (fill truth needs a cross-schema read). CORRECTIONS RETRACTED VISIBLY (in the rewritten spec):p_cal replicates his side skill” (the 94% is agreement with the MECHANICAL side; his +3pp discretionary skill is NOT reproducible — that was W2’s whole substitution bet, it fails); “eth15 as co-primary” (his eth-15m lost 14.4k, TAKER_REBATE rows exist), dormant pre-registered A0/drift 5m re-tests. REOPENERS: materially faster PM execution path, venue structure change, a new auto-settled market family with non-repriced books, or a non-latency-locked signal. Artifacts: data/fifteen_first_trigger.csv, data/fifteen_clob_labels.csv, data/fifteen_first_trigger_labeled.csv (3,131 rows each); verdict doc docs/superpowers/specs/2026-07-18-slow-clock-resite-program-design.md (rewritten with the verdict). — crypto-shortterm-15m-resite-premise-falsified-2026-07-21, crypto-shortterm-fill-space-closed-15m-pivot-proposal-2026-07-18, crypto-shortterm-competitor-wowitsamazing-refuted-2026-07-03, crypto-shortterm-15m-observe-nogo-2026-07-06, crypto-shortterm-loss-investigation-real-vs-onchain-edge-2026-07-10, crypto-shortterm-honest-recalibration-edge-gate-verdict-2026-07-11, crypto-shortterm-backtesting, crypto-shortterm-market-expansion-scan-2026-07-09, onchain-label-overstates-traded-winrate, signal-healthy-fill-is-the-problem

2026-07-18

  • SOLUTION-DESIGN SESSION — the 5m [60,150)s fill-structuring space is FORMALLY CLOSED (settled), and a W0–W3 pivot to a 15m offline gate is PROPOSED (pending Andras’s decision — NO decision taken, NO code/config/live change).STATUS: all live lanes remain dry_run (canary [[crypto-shortterm-canary-confirmed-fill-ev-live-2026-07-16|#1210]] aftermath confirmed closed, kill confirmed, sources back to dry); the program below is a recommendation awaiting go/no-go. (1) SETTLED CONCLUSION — space closed: every taker fill-lever is now negative or shave-to-breakeven — A0-avoidance (regime-confounded), exits ×2 (NO-GO), chase/limit-raise (OOS-rejected), market/FOK/GTC-maker (NO-GO), co-lo/latency (shelved OOS), depth-aware sizing UP (NO-GO), size-down/depth-selectivity (it’s liquidity not size), recalibration (NO-GO for EV), EMIT_MAX_ASK tighten, time-of-day (OOS-reversed). Mechanism NAMED = structural last-look: our PUBLIC-info signal (Binance+Chainlink) is repriced by sub-second makers inside the venue’s mandated ~250ms taker protective delay; clean-window A1+ reprice fills ≈ −0.5¢ breakeven (A0 scope) ⇒ even INFINITE producer-side speed caps at breakeven, so the ceiling is a property of trading public info as a protected-book taker, not of any one lever → no further taker bid-structuring studies on this band. Confirmed live by the canary 13.6pt gap + the front-loaded ask-run (speed non-binding). (2) NEW VERIFIED FACTS (read-only pass 2026-07-18): (a) btc15/eth15 observe runs since 2026-07-02 19:45Z (NOT 07-09) = 1,490 windows/asset, 96/day, outcomes 100% settled (1490/1490 both), both books ≥94% populated → an offline 15m gate is runnable NOW (thin: 2 assets, ~2.7-day purged folds); (b) 5m slow-clock τ∈[150,300)s population sufficient (~3,311 windows/14d pooled) BUT median first-qualifying ask 0.90–0.93 (the −EV pocket; only thin [210,300) reaches 0.86–0.90) — conviction and ask co-move → downgraded to cheap falsification; (c) the 2026-07-03 backtest-fill-harness spec was NEVER built — current walkforward assumes CERTAIN fill at ask+1¢ flat, resolved_up labels (~10pt optimistic), zero latency/book-evolution = the root cause of “offline +EV, live −EV” (deps refitlib/exitlib on branch feat/calibration-refit; fill truth pm_fak_attempts/pmv2_autotrade_orders needs cross-schema read); (d) emit’s 15m block is a one-line compiled guard (emit.rs:86-93 emit_supported), gated on a validated 15m calibration+regime, never before. (3) PROPOSED PROGRAM (W0–W3 + exit rule): W0 taker-rebate check (free, named twice, never run); W1 build the honest-fill harness per the 07-03 spec (fill model calibrated on our ~1,400 recorded fills, self-validated on held-out data, won_up labels only); W2 PRIMARY = 15m offline gate: fresh fit-then-frozen 15m calibration (days 1–5), OWN regime band + ask cap (do NOT inherit 5m values), ONE confirmatory test on remaining days (minimal-trials), honest fills at measured latency, per-asset, deflated bar — existence proof to beat: competitor’s btc-15m late band +3.8¢/share CI[+2.2,+5.4] at HIS fills (a ~1s taker like us); refutation to overcome: mechanical trigger −1.1¢ (his edge was side/timing SKILL; our candidate replacement = p_cal, 94% side-agreement on 5m); W3 5m τ∈[150,300) falsification through the same harness (1-line filter, expectation LOW). Pre-committed EXIT RULE: if W2+W3 both fail the bar under honest fills → shelve the trading lane (observe keeps running as a data asset; the two pre-registered 5m re-tests stay dormant unless PM ships a materially faster path or venue structure changes). (4) ALTERNATIVES REJECTED: canary #2 now (breakeven-bounded; can ride a future 15m canary), maker/MM flip (no speed/info edge — we’d BE the stale quote; rejected 07-02 & 07-09), further latency spend (ask-run front-loaded + last-look ⇒ speed non-binding), shave-stack re-arm (breakeven ceiling, stacked post-hoc filters = overfit), longshot-fade (bias is favorite-UNDERpricing near settlement), HYPE/7th asset now (multiplies broken fill economics). (5) BABYLON: open task #1189 (positionmanager exec-latency decomposition, READY @ 45fa0eb) = the latency-parameter source for W1; reply deferred until direction set. — crypto-shortterm-fill-space-closed-15m-pivot-proposal-2026-07-18, crypto-shortterm-canary-confirmed-fill-ev-live-2026-07-16, crypto-shortterm-size-depth-study-liquidity-not-size-nogo-2026-07-17, crypto-shortterm-loss-decomposition-fill-adverse-selection-2026-07-10, crypto-shortterm-a0-fill-policy-offline-scope-nogo-2026-07-12, crypto-shortterm-discriminator-catchability-ask-drift-2026-07-13, crypto-shortterm-15m-observe-nogo-2026-07-06, crypto-shortterm-competitor-wowitsamazing-refuted-2026-07-03, crypto-shortterm-market-expansion-scan-2026-07-09, crypto-shortterm-backtesting, crypto-shortterm-order-type-fak-verdict-2026-07-09, crypto-shortterm-phase1-execution-decision, signal-healthy-fill-is-the-problem

2026-07-17

2026-07-16

2026-07-13

  • DISCRIMINATOR RE-VERIFICATION + ASK-DRIFT CATCHABILITY — the “unfilled-would-win” discriminator is VALIDATED (opposite of A0), BUT the missed winners are largely NOT physically catchable. Read-only follow-up to the A0 offline-scope NO-GO; same per-asset/per-regime scrutiny re-run on the underlying discriminator. (1) Discriminator HOLDS (unlike A0): on data/signal_fill_decomposition_2026-07-10.csv (5,555 fires, 07-04→10, real_win=CLOB truth, validated 100% vs winner CSVs) unfilled fires win 80.6% vs filled 73.4% = ~7.2pt gap (pooled ≈6 SE, 3,484 vs 2,066). Survives every cut A0 failed: positive in ALL 6 assets (btc 9.6, doge 9.3, eth 8.9, bnb 8.9, xrp 8.2, sol 4.0); positive on ALL 7 days green+red (5.7–11.3pt, never near zero — where A0 flipped +0.4→−7.2, this is structural not a drawdown artifact); survives conditioning on conviction (7.5pt within the dominant p_cal stratum, n=5,418 → adverse selection on the unobservable, not “unfilled were higher-conviction”). (2) Missed winners are +EV-capturable AT THE QUOTED ASK — edge over breakeven: ask 0.5→+17.6pt, 0.6→+14.2, 0.7→+10.4, 0.8→+5.0, 0.9→−4.6 (the >0.90 pocket EMIT_MAX_ASK already blocks); prize biggest at cheap asks <0.75. (3) BUT ask-drift kills the speed lever: on real best_ask_at_snap no_match rows (n≈1,810 reprice attempts, 07-09→11) 99.9% had the real ask ABOVE our limit — not phantom liquidity, not a lost race, the price genuinely repriced up; median +0.07 above the emit ask (p90 +0.10–0.11), would_win 83%. The run is FRONT-LOADED — drift-vs-emit identical at attempt 1 and 2 (p50 0.07 both) → the ~7-tick lift completes during the emit→first-touch latency window and plateaus; only ~20% still within 2 ticks of emit → a tighter FAK ladder can’t help (run already done by attempt 1). CONSEQUENCE: catchable fills = the stale-ask losers (A0); +EV fills require filling at the emit ask before a near-instant ~7-tick lift = a sub-latency race (residual RTT 117–260ms; co-lo shelved as not clearing the OOS bar). Chasing the +7 ticks ≈ breakeven at 0.72, negative above 0.78 (why the break-even limit correctly refuses). NO cheap fill-capture lever adds EV. Honest fill levers = quality-not-quantity: EMIT_MAX_ASK 0.85→0.78–0.80, refuse the thin-book toxic pocket (doge/xrp). Independently re-confirms co-lo on EV grounds from the producer side. This VALIDATES the standing FILL adverse-selection thesis and the cheap-asks-are-the-durable-edge read; chasing the reprice up was already a NO-GO. Caveats: clean attempt-0 first-touch drift needs post-2f0ec6ad real best_ask_at_snap (accrue forward — current inference rests on the reprice-plateau); real-ask data red-days-only (07-09+) but drift is microstructural; decomp CSV ends 07-10 so 07-11 attempt-0 rows don’t join to labels. Source: docs/superpowers/a0-fill-policy-offline-scope-2026-07-12.md (Follow-up 2026-07-13). — crypto-shortterm-discriminator-catchability-ask-drift-2026-07-13, crypto-shortterm-a0-fill-policy-offline-scope-nogo-2026-07-12, crypto-shortterm-loss-decomposition-fill-adverse-selection-2026-07-10, crypto-shortterm-honest-recalibration-edge-gate-verdict-2026-07-11, crypto-shortterm-disagreement-ceiling-nogo-2026-07-03, crypto-shortterm-order-type-fak-verdict-2026-07-09, crypto-shortterm-phase1-execution-decision

2026-07-12

  • A0-AVOIDANCE FILL POLICY — offline scope NO-GO; “A0 first-touch is structurally toxic” is a TELEMETRY/REGIME ARTIFACT. Single-session, read-only. VERDICT: do NOT build the A0-avoidance (skip/delay first-touch) policy — it would be a 6th in-sample-overfit NO-GO (overfits 2 drawdown days). The confound: pm_fak_attempts.attempt_no=0 telemetry was only added with PM’s 45fa0eb, so A0 rows exist only from 07-09 onward = exactly and only the 3-day directional drawdown (07-09/10/11, the week’s worst P&L days); ZERO A0 data on the 5 green days → “A0 toxic” cannot be separated from “the 07-10/11 regime was toxic to everything.” Pre-07-09 pm_fak_attempts has count(attempt_no=0)=0 (proof). The like-for-like correction: the prior split compared red-day A0 fills vs green-day A1+ reprice fills — the flattering “+1.30¢ A1+” was green-day reprices; inside the clean 07-09→11 window A1+ = −0.5¢ (breakeven), NOT the +1.30¢ that made the book look like it “flips +EV when A0 removed.” A0 not stable across days: +0.36¢ POSITIVE on 07-09, −7.21¢ (07-10), −6.67¢ (07-11) — a structural pathology would show day 1. A0 asset-concentrated (thin-book signature, not universal first-touch selection): btc −1.10¢, eth −1.97¢ (~breakeven) vs doge −10.45¢, xrp −9.13¢ (echoes the depth NO-GO). Counterfactual unobservable (FAK stops on fill): optimistic bound (all A0 re-fills at A1+ economics) only reaches ~−0.5¢ (breakeven-negative); pessimistic keeps just ~15% of fill volume — skip-A0 produces NO +EV book in the only regime with A0 data. EV reconstruction validated vs booked realized_pnl_usd (A0 booked −6.21¢/share vs reconstructed −5.35¢). This REFINES the standing A0 finding: the general loss-at-fill thesis STANDS; the A0-SPECIFIC structural claim is downgraded to regime-confounded, and the “A0 is the real lever” framing is parked. Pre-registered re-test: 45fa0eb now records A0 daily — once ~5 green-regime days of A0 telemetry accrue, re-run this exact within-regime per-asset A0-vs-A1+ comparison. Independent open levers unaffected: EMIT_MAX_ASK 0.85→0.78–0.80, the doge/xrp thin-alt fill pocket. Source: docs/superpowers/a0-fill-policy-offline-scope-2026-07-12.md. — crypto-shortterm-a0-fill-policy-offline-scope-nogo-2026-07-12, crypto-shortterm-loss-decomposition-fill-adverse-selection-2026-07-10, crypto-shortterm-honest-recalibration-edge-gate-verdict-2026-07-11, crypto-shortterm-depth-aware-sizing-nogo-2026-07-10, crypto-shortterm-overnight-gate-a0-lever-rearm-2026-07-10

2026-07-11

  • HONEST-RECALIBRATION EDGE-GATE VERDICT — signal HEALTHY, recalibration NO-GO for EV + a VISIBLE RETRACTION of the same-session “edge eroding” read. Single-session, read-only, un-peer-reviewed; no live change, nothing merged, calibration.json untouched. ⚠ RETRACTION (prominent): an earlier same-session read that “the signal edge is being priced out (real win 80→73%, at-ask ceiling +7¢→~0)” was WRONG — it was a coarse CROSS-ASSET by-day aggregate skewed by SOL (07-10 sol −11.2¢ n=62, a PARTIAL day) while BTC was actually +9.2¢ and eth +3.8¢ the same day. Lesson (binding): measure erosion PER-ASSET, never cross-asset aggregate — a single weak/partial-day lane in the pool manufactures a fake decay trend (same asset-blind-aggregate family as the window_id cross-asset join trap). VERDICT (rigorous refit study, 5 purged/embargoed folds, 5,671 OOS windows, real BTC CLOB labels): MARGINAL/UNDERPOWERED — the edge PERSISTS, NOT priced out. (1) Signal HEALTHY: the late window (≥07-05) is the STRONGEST, honest win-rate stable ~0.77–0.79, +3.46¢/fire; all 6 assets +EV late week. (2) Honest recalibration NO-GO for EV: frozen 2,633 fires / 76.3% / +37.13 (own_z 2.30) → Δz −0.30 (tie); best of 16 variants +0.16 « deflated bar 3.99. The frozen map IS overconfident ~5–11pt/bin (claims 0.912 where real 0.808) BUT the gate absorbs it at the ~0.72 ask → recalibration recovers ZERO EV, it only refuses ~40% near-breakeven expensive-ask fires. (3) The loss is FILL-side (A0 first-touch), NOT the signal — reconfirms the standing FILL adverse-selection / A0-lever finding; lever = fill-capture. DO NOT shelve. ACTIONABLE: (a) A0 fill fix = #1 lever; (b) durable edge is in cheap asks (<0.75), the 0.78–0.85 tail is thin near-breakevenlower EMIT_MAX_ASK 0.85→~0.78–0.80 + re-gate when re-arming; (c) recalibrate ONLY for sizing-honesty (frozen over-sizes ~10pt), never for EV. Genuine-erosion re-check trigger: frozen-gated FILL win-rate <73% sustained over a full ≥5-day window (per-asset, not a partial day, not a pool). CAVEAT: the rigorous folds are BTC-only (binance_ticks BTC-only); the 5 alts were cross-checked (+EV late week) but NOT independently refit — resolve before any live GO. LIVE-STATE: all lanes DRY (no live exposure); the overnight re-arm was PULLED after night-1 lost −$189 (directional) — within the underpowered-noise band that note flagged. Artifacts: branch feat/calibration-refit-degraded-rerun (commit ae31a0c, not pushed/merged); report docs/superpowers/refit-honest-degraded-verdict-2026-07-11.md; scripts/gate.py, scripts/gate2.py; calibration.json untouched. — crypto-shortterm-honest-recalibration-edge-gate-verdict-2026-07-11, crypto-shortterm-loss-decomposition-fill-adverse-selection-2026-07-10, crypto-shortterm-overnight-gate-a0-lever-rearm-2026-07-10, crypto-shortterm-disagreement-ceiling-nogo-2026-07-03, crypto-shortterm-depth-aware-sizing-nogo-2026-07-10, crypto-shortterm-loss-investigation-real-vs-onchain-edge-2026-07-10, crypto-shortterm-risk-management, signal-healthy-fill-is-the-problem, loss-is-adverse-selection-at-fill, onchain-label-overstates-traded-winrate

2026-07-10

  • OVERNIGHT-GATE OOS VERDICT (PROVISIONAL) + A0 IS THE REAL LEVER + OVERNIGHT RE-ARM. Single-session, read-only, un-peer-reviewed; purged/embargoed OOS on 4,773 live filled orders, REAL CLOB labels (realized_pnl_usd sign 100% agreement with the CLOB winner = label-truth-grade at scale). VERDICT: the overnight-(20–06 UTC)-+EV / daytime-(07–19)-−EV split is DIRECTIONAL-REAL but NOT statistically structural on 8 days. Aggregate overnight +560 / −2.15¢/sh; effect holds within-lane (+0.0352/fire within-lane vs +0.0372 aggregate — not lane-mix). BUT P(overnight>daytime)=0.949 (diff CI includes 0) and P(overnight>0)=0.875 (overnight-positive NOT significant — the robust leg is “daytime worse”). Walk-forward met the bar in only 5/9 day-folds & 3/6 contiguous, WITH sign reversals 07-03/04/08 (the time-of-day NO-GO flip signature); 67% of the daytime loss is the single 07-09 US-open blowout — dropping it collapses P 0.949→0.883UNDERPOWERED / period-noise-consistent. Overnight’s edge is largely that it runs the known-winner lanes (sol_up/eth_up/xrp_down/doge_down/eth_down, ~5). A0 MECHANISM CONFIRMED (PM #1191): the toxic fills are the A0 first-touch−5.70¢/sh @ 69.4% win (−EV); A1+ repriced +1.30¢ @ 75.3%; removing A0 flips the 2-day book +EV. A0-toxicity is SEPARATE from time-of-day (daytime A0-share 82.8% vs overnight 80.1%, only +2.7pp) → the overnight gate is NOT a proxy for avoid-A0 (orthogonal, additive levers). This sharpens the FILL adverse-selection verdict (pinpoints where in the fill sequence the toxicity is) and corrects/extends the FAK verdict (its clean “repriced fills win 87–88%” = the A1+ subset; the A0 blend it averaged over is −EV on the real label; order-type conclusion still stands). DECISION: keep the overnight re-arm — +EV/low-risk but PROVISIONAL, NOT a structural rulere-test after ~2–3 wk forward data; the better-supported lever = A0/adverse-fill mitigation (fill-policy: avoid the toxic first-touch), NOT the clock, NOT depth-aware sizing (NO-GO). LIVE-STATE CHANGE: Andras re-armed the 7 lanes overnight-only (@deploy, babylon pmv2 1195) — scheduler RESTORED, LIVE 20:00–07:00 UTC (22:00–09:00 CEST), DRY daytime 07:00–20:00 UTC; the old 13–16 CEST daytime slice explicitly REMOVED (−EV); levi/gyula stay OFF; wallets 1×. This supersedes the earlier-07-10 all-lanes-dried state. PM’s 45fa0eb telemetry-fix (best_ask_at_snap was a req.limit proxy at attempt 0) rolls alongside so overnight-forward fills are measured against the real book. Reconciles with (does NOT overturn) the 07-06 time-of-day NO-GO: it confirms the caution (period-noise, reversals) — the re-arm is a provisional risk tilt on known-winner lanes, not a validated clock edge. — crypto-shortterm-overnight-gate-a0-lever-rearm-2026-07-10, crypto-shortterm-loss-decomposition-fill-adverse-selection-2026-07-10, crypto-shortterm-order-type-fak-verdict-2026-07-09, crypto-shortterm-time-of-day-filter-nogo-2026-07-06, crypto-shortterm-depth-aware-sizing-nogo-2026-07-10, crypto-shortterm-loss-investigation-real-vs-onchain-edge-2026-07-10, crypto-shortterm-risk-management, overnight-edge-provisional-a0-is-the-lever, loss-is-adverse-selection-at-fill
  • LOSS DECOMPOSITION → decisive verdict: the 07-08→07-10 loss is a FILL problem (adverse selection), NOT a bad signal + Andras DRIED ALL LIVE LANES. Single-session, read-only, un-peer-reviewed; 100% real CLOB-label coverage (4857/4862 fired windows). SIGNAL CEILING (all fires @ emit-ask, real labels) +3.68¢/share / 77.1% real-signal-win, stable ~77%/day, no collapse (full 07-04→now +4.46¢/78%) → signal is +EV at a clean fill. REALIZED (live fills, net fee) −3.87¢/share / 70.8% fill-win / −219, xrp-down −78 ≈ 641). The +EV lives in the UNFILLED tail (winners that reprice above our limit before we fill). KEY CORRECTION to prior canon: the ~50%/day FAK no-matches are NOT “the EV guard correctly refusing −EV trades” (as the 2026-07-09 FAK verdict framed them) — they are the +EV winners (84.9%) we MISS because the winning-side ask runs away (latency); same mechanism, opposite sign (cost of missing winners, not benefit of dodging losers). Added a non-destructive forward-pointer correction callout to the FAK note (body/verdict left intact — order-type conclusion still stands). TIES THE DAY: explains why depth-aware sizing was NO-GO (wrong lever — winner-capture not depth-reach), why the 2 extra wallets were exonerated (regime/exposure amplifier, not a defect), and confirms the signal is healthy (don’t “fix the model”). Only lever = fill-capture (cut time-to-fill to catch the winning tail) — NOT signal/depth/size, and NOT pricing (slippage negligible ~0.08¢). Caveat: on the filled subset the at-ask ceiling is already ~−1.9¢**, so the +EV is almost entirely in the unfilled tail (thin/latency-locked; chasing it up was a prior NO-GO). LIVE-STATE CHANGE: Andras dried ALL live crypto_5m_algo lanes (task to @deploy, babylon pmv2 #1187) — the 7 live lanes (btc-down, eth-up/down, sol-up, xrp-down, doge-up/down) → dry_run; other 5 already off. Reason: all 8 fill-bearing lanes −4.6 to −18.1¢/share on real labels (migration 0011). Reversible — re-arm lane-by-lane when a lane’s real edge crosses +EV; monitored via depthwatch + dry fires scored on real labels (the manual daily-loss backstop the system structurally lacked). Migration 0011 adds real-edge columns to crypto_lane_health (edge_per_share_cents/filled_shares_24h/avg_fill_px) + marks signal_win/toxicity_gap on-chain-optimistic — committed on branch worktree-agent-af81d3911fd13808c, NOT yet applied live. Artifacts: data/signal_fill_decomposition_2026-07-10.csv, data/{asset}5m_actual_winners_2026-07-10.csv. — crypto-shortterm-loss-decomposition-fill-adverse-selection-2026-07-10, crypto-shortterm-loss-investigation-real-vs-onchain-edge-2026-07-10, crypto-shortterm-order-type-fak-verdict-2026-07-09, crypto-shortterm-depth-aware-sizing-nogo-2026-07-10, crypto-shortterm-lane-state-adverse-fill-2026-07-05, crypto-shortterm-risk-management, loss-is-adverse-selection-at-fill, onchain-label-overstates-traded-winrate
  • DEPTH-AWARE DYNAMIC SIZING → offline edge-gate NO-GO; shelve it (supersedes repo HANDOVER.md item 5 “build depth-aware sizing post-canary”). Single-session, read-only, un-peer-reviewed; nothing merged (branch worktree-agent-a4946f9b9c75fafd9, commit 663b8af, NOT pushed). Tested per-window clamp(DEPTH_FRACTION × side_best_ask_sz, 5, EMIT_MAX_SHARES) vs fixed 1× (S=5) and retired 2× (S=10) — new bin crates/analysis/src/bin/depthsizing.rs, paired per-window, ACTUAL CLOB labels (NOT resolved_up), K=6 purged/embargoed walk-forward, clustered UTC-day bootstrap (2000 resamples, seed 42, deterministic), Down = symmetric down-book BUY; DRY module crates/analysis/src/depthlib.rs lifted byte-identical from depthaware.rs. Best vs 1×: FRAC=1.0 / MAX=10 → ΔPnL/window +0.141, delta t=2.60 (d_z=2.24, 3/5 folds) — BELOW the Bonferroni t≥3.08 bar (24 candidate×baseline trials); vs 2×: 1.0/15 → +0.139, t=2.54 — also below. DECISIVE: every candidate that out-earns 1× raises per-window exposure 1.9×–4.7× (best-t ~1.97×, ~4.1/window) with ¢/share 2.32–2.83 vs 2.74 baseline = NO per-share edge → it’s an offensive size-up, not a de-risk. Live shadow confirms: observe-only crates/observe/src/bin/depthwatch.rs~90% of live fires would size UP, avg 15.3 sh vs fixed 5 (~3×) at frac=1.0/MAX=20. Wouldn’t have helped the bleed anyway: depth-reach is a non-problem at 1× (~5–10 sh vs median book ~25+); the real drivers are adverse fill + the ~10pt label mask + a razor-thin edge. Caveats: the correct down-buy pass has 934 windows but only 5 UTC days (down_best_ask_sz populated only from ~06-28) — but a full-range up-book pass (2325 windows / 13 days, Down as sell-up) ALSO fails (best t=1.55), so not a 5-day artifact (reopener: backfill pre-06-28 down-book depth, re-run); depth_capped_fill models the 3¢ depth PRICE-penalty this policy targets but NOT the reprice/miss hazard (fill-RATE optimistic → biases against the NO-GO). Gates green (127 tests incl. 11 new depthlib); spec docs/superpowers/specs/2026-07-10-depth-aware-sizing-study-design.md, report .../depthsizing-report-btc.md. — crypto-shortterm-depth-aware-sizing-nogo-2026-07-10, crypto-shortterm-loss-investigation-real-vs-onchain-edge-2026-07-10, crypto-shortterm-phase0-depthaware, crypto-shortterm-per-wallet-size-scale-2026-07-06, depth-aware-sizing-nogo
  • LOSS INVESTIGATION — NO code regression, NO signal failure; a 3-day sub-breakeven SLIDE (not a one-day blowout), a razor-thin edge, and monitoring that reads ~10pt too healthy. Read-only duckdb→Postgres, un-peer-reviewed (single-session, main agent). (1) No regression: the 07-09 emit commits (E1 nats.rs Pub-Ts header, E4 emit.rs claim_ms/pub_lag_ms) are telemetry-only + dedup-safe (Nats-Msg-Id still =order.id, unit-tested); E2/E3 are offline analysis scripts; publish-first is flag-OFF/dark. (2) Model healthy: on-chain signal-win 85–88% every day incl. 07-09 (85.7%) + hour-by-hour through US-open; fires ~880–920/day; asks ~0.72. (3) The real picture: live realized_pnl_usd (CLOB truth) 07-05 +15 / 07 +120 / 09 −170 (partial)07-08 was ALREADY negative, contradicting the repo HANDOVER.md “+448 on 07-09” — a slide, not a blowout. (4) KEY — resolved_up overstates the TRADED win ~9–14pt on the FILLED subset (sharpens the 4–7pt per-bin since fills concentrate in near-ties): 07-08+07-09 fills 111 on-chain-WIN-but-real-LOSS vs 13 the reverse (median 3.5 bps from strike) = −$1,011 unseen; real fill-win on loss days 64–72% vs on-chain 78–86% (BE ~73%) → every gauge on resolved_up (signal-win dashboards, the CRM signal_win−fill_win toxicity tab, break-even bookkeeping) reads ~10pt healthier than the cash; the edge is razor-thin ~1–3pt over BE (the documented 73.7% vs ~73% BE). ⚠ Tension: the FAK verdict’s “fills win 87–88%, no adverse selection” is on the OPTIMISTIC label — re-measure on realized_pnl_usd split by side+day. (5) Fingerprint = side-specific INTRADAY adverse fill: on each loss day one side’s real fill-win collapses below BE (Up 07-08 65% & 07-10 46%; Down 07-09 64%) while the other holds ~74–77% (profitable days both ~74–76%); visible even on the optimistic label (losing-side filled opt-win 11–22pt below all-signal; today’s Up filled 56% vs all-Up-signal 78%) — generalises the btc-up signature. NOT a daily trend (up-rate ~48–50%) or chop regime (near-ties did NOT track losses — 07-05 had the most, 42.9%, yet was most profitable); ~50%/day FAK no-match = separate conversion loss. (6) 07-09 amplifier: wallets 3 & 4 (levi/gyula) live only 07-08 20:02 → 07-09 12:58 UTC (152/154 fills each), exactly spanning the blowout → 07-09 ran on 4 wallets = 2× exposure into a correlated down loss; pulled to 2, but 07-10 still bled (amplifier, not cause). (7) No live daily-loss circuit breaker — caps are deferred) → the slide ran unchecked. OPEN: the exact intraday trigger of the side-collapse (needs hourly price-move vs fill-win-by-side); a CLOB tokens[].winner cross-check to 100% confirm the ~10pt gap is label noise (strongly indicated by the 111-vs-13 / 3.5 bps asymmetry) vs residual PM booking lag. — crypto-shortterm-loss-investigation-real-vs-onchain-edge-2026-07-10, crypto-shortterm-lane-state-adverse-fill-2026-07-05, crypto-shortterm-pnl-attribution-corrections-2026-07-03, crypto-shortterm-order-type-fak-verdict-2026-07-09, crypto-shortterm-perf-dashboard-2026-07-05, onchain-label-overstates-traded-winrate

2026-07-09

  • ORDER-TYPE QUESTION SETTLED — keep FAK marketable limits at the break-even ceiling; do NOT switch to market/FOK or resting GTC. PM places FAK (immediate-or-cancel) MARKETABLE LIMIT orders, price-pinned (market_order().price(limit).order_type(OrderType::FAK), position_manager/crates/polymarket-clob/src/lib.rs:502 buys / :513 sells; hardcoded, no per-source order-type config; vendored rs-clob-client-v2 supports GTC/GTD/FOK if ever needed); limit = floor_tick(min(directive max_price, fresh best_ask)), FAK-v2 retries reprice upward capped at max_price (pricing.rs:58-64) — invariant held 0/9,047 attempts ever above the ceiling. “Adverse fills” decomposes: (a) ~50%/day FAK no-match = missed fills, 90.5% price-refusals where the BE ceiling correctly refused a repriced ask (avg +5.4¢ above) = the EV guard working, a latency-conversion loss (fix in flight: R-items, config cache 733ms→10ms, babylon 1157); (b) realized fills show NO measured adverse selection — repriced fills win 87–88% (#942, ~2,200 fires) ≈ population, and BE-priced limits guarantee ≥3¢ model edge/fill; (c) partial/under-fill = the open E3 watch (p10 touch thinner than clip). Alternatives lose: uncapped market/FOK buys the above-BE −EV pocket (the pre-07-02 38%-of-fires-−EV lesson); resting GTC maker quotes in τ∈[60,150)s = free option to sub-second Binance-watching bots vs our 2.78s p50 requote path, no cancel/replace infra, and maker fills void the validated taker economics (full gate re-run required). Open cheap E2 question: win rate of price-refused vs filled directives (adverse selection on CONVERSION) — quantifies what recovered conversion is worth. — crypto-shortterm-order-type-fak-verdict-2026-07-09, crypto-shortterm-fak-nomatch-root-cause-2026-07-02, crypto-shortterm-lane-state-adverse-fill-2026-07-05, pmv2-open-multi-wallet-liquidity-ceiling-brainstorm
  • SUPABASE INDEX AUDIT (shared project mkofmdtdldxgmmolxxhc) — DB 8.4GB (crypto_shortterm 4.3 / public 4.0); top-10 statements ≈ 80% of server exec time, PM’s schema owns ~65%. Biggest sinners: TEXT end_date regex sweeper 6.5s×1,371 (16.9%); pm_trades fat select 10,465 rows/call ×20k (19.8%, index fine — app query shape); pm_trader_metrics GROUP BY MAX read 2.6B tuples via 19,928 seq scans; pmv2_market_resolutions DISTINCT ON 549ms×3,927 with a useless 155MB index (4 lifetime scans) on an ~800k-rows/day insert table. Per-owner DDL posted to babylon #1163 (question #1162): PM 2 partial indexes + 3–4 drops + 2 query rewrites; deploy realtime WAL polling ~2.8% + dead event_2026_05/06 partitions 131MB; weather snapshots seq-scan pattern. Crypto side: migrations/0010_index_hygiene.sql written, NOT applied (obs_asof_idx 126MB/946-scans btree → BRIN); writers healthy (0.1–0.2ms); the real crypto lever = retention/partitioning on observations+pm_book_snaps (the 912MB book_snaps pkey IS the ON CONFLICT dedup key over 77-char TEXT token_ids — only retention shrinks it). Decision left to Andras. GOTCHA (unresolved correctness question): binance_ticks pkey = (trade_ts, price) with NO asset column — if multi-asset ticks land there, cross-asset same-price ticks silently dedup (same asset-blind-key family as the window_id trap). Repro: zsh scripts/db_index_audit.sh (duckdb postgres ATTACH, read-only, masked conn); full report docs/superpowers/db-index-audit-2026-07-09.md. — crypto-shortterm-db-index-audit-2026-07-09, crypto-shortterm-observations-window-id-not-asset-unique, crypto-shortterm-data-sources, crypto-shortterm-empire-ui-crm-topology-2026-07-06
  • PRODUCER-SIDE LATENCY LEAK MEASURED + FIXED (flag-gated, not rolled) — the emit DB claim INSERT gated the NATS publish. crates/observe/src/bin/emit.rs does a synchronous INSERT INTO crypto_shortterm.emitted_signals … ON CONFLICT DO NOTHING (the claim, logged as claim_ms) and only calls publish_entry (the NATS Directive) after it commits — so the claim latency sits on every fire’s critical path, ahead of the publish. Deploy grepped live logs: 234–301ms median, p90 ~1–2.5s, max ~3s. Cause = cross-region: emit runs on pmv2-zurich (GCP europe-west6) but Supabase mkofmdtdldxgmmolxxhc is eu-west-1 (Dublin) → cross-region round-trip + TLS/pooler overhead paid synchronously before the Directive leaves the box. Blows the <50ms budget on its own; plausible real contributor to the adverse-fill loss (5m adverse selection is in the final seconds). This is the producer-side complement to PM’s consumer-side R0–R6 work; measured (log grep), not assumed. FIX — feat/emit-publish-first-dlq @ 6cd9487 (pushed, NOT rolled): behind env flag EMIT_PUBLISH_FIRST (default OFF = byte-identical legacy C1 path); ON = publish the Directive FIRST, record emitted_signals async (tokio::spawn) off the critical path; on insert failure → append JSONL DLQ (EMIT_DLQ_PATH, default /tmp/emit_dlq_<label>.jsonl) + rate-limited (1/5min) Telegram alert via telegram.outbound; a drainer retries every 60s (ON CONFLICT DO NOTHING) and atomically rewrites survivors. New module crates/observe/src/dlq.rs (append/read_all/atomic-rewrite, corrupt-line skip, 4 TDD tests); shared helpers insert_signal+emit_directive (DRY); shared async mutex dlq_lock serializes DLQ file access. SAFETY LINCHPIN (OPEN, babylon pmv2 #1160): publish-first surrenders the C1 persist-before-publish DB-claim cross-process dedup — now relies on (a) per-instance in-memory fired_windows + (b) Nats-Msg-Id (= source|window) → JetStream stream dedup. Safe to flip ONLY if JetStream message-dedup with a duplicate window ≥ ~2min is confirmed on the crypto entry stream (PMV2_ORDERS); if off/tiny, publish-first can double-fire during a deploy roll. Third chapter of the ordering saga: bc38734 publish-then-mark → C1 claim-then-publish (never-duplicate, bought the claim_ms cost) → publish-first (buys latency back, moves dedup DB→JetStream). ROLLOUT (gated on #1160 + Andras): roll dark → 1-lane canary (watch pub_lag_ms drop, no double-fire, DLQ empty) → fan out. Op note: default /tmp DLQ is tmpfs → lost on restart; set EMIT_DLQ_PATH to a persistent volume (loss = one analytics row, never a trade). — crypto-shortterm-emit-publish-first-latency-2026-07-09, crypto-shortterm-down-side-money-path-fixes-2026-06-28, pmv2-measure-before-mitigating-priority-tiers-2026-07-09
  • MARKET-EXPANSION SCAN (10-agent workflow) — NEGATIVE RESULT, MOAT IDENTIFIED: 0 of 25 candidates across 8 Polymarket families survived; the 5m/15m Chainlink-settled crypto family is the ONLY automatic, dispute-free, minutes-to-redemption settlement on the venue. Everything else is UMA optimistic-oracle (proposal 12–16 min post-close + ~2h liveness) → kills capital velocity for our BUY-only hold-to-resolution pipeline; the lane’s real moat = settlement velocity (~2 min to redeemable), not just the reprice-lag model. Hourly up/down (btc-up-or-down-hourly id 10114, 20–150/day gross), resident fast bots confirmed (0xee55214e… 53/500 trades), BTC hourly tick 0.001 not 0.01 (fresh calibration + walk-forward gate required if ever revisited) → strictly dominated by 5m; observe-only sketch + kill criterion recorded. Monthly touch barriers REFUTED by natural experiment (reach-62.5k, 07-03): no post-touch lag window (1.8k all @0.999, every economic entry ≥0.90 > our const-asserted EMIT_MAX_LIMIT 0.88 (un-fireable by design), a single MM parks ~4.5–22k books), no model edge (winner is a lookup), ~2h15m release-to-redemption. Watchlist (passive, zero build): HYPE (Hyperliquid) newly listed as the 7th up/down underlying at 5m/15m/4h/daily with ~0.5–2.2k/day; WTI monthly touch settles on Pyth ($392k vol24h). Weather map delivered for the weather peer: 51 cities (only NYC/London are series objects, rest only via tag_slug=weather), NO intraday products, next-day resolution, per-city heterogeneous settlement (Wunderground airport stations; Hong Kong = HKO), Gamma endDate = display artifact. Source: docs/superpowers/market-expansion-scan-2026-07-09.md (workflow wf_3a86c588-6be, all numbers live pulls 16:45–16:58 UTC). — crypto-shortterm-market-expansion-scan-2026-07-09, crypto-shortterm-15m-observe-nogo-2026-07-06, crypto-shortterm-bnb-expansion-eval-2026-07-05
  • CONFIRMED SAME DAY — the multi-wallet ceiling is PRICE, not book depth (the decisive query ran; depth story REFUTED at current clip sizes). Read-only duckdb join pm_fak_attempts × crypto_shortterm.emitted_signals × pmv2_autotrade_orders over the telemetry week: 4,098/4,528 (90.5%) final retry attempts refused with best_ask_at_snap > repriced_limit — ask avg +5.4¢ ABOVE our ceiling, having moved avg +9.7¢ since emit; only 21 (0.46%) show ask ≤ limit without a fill, and most of those are same-second wallet1+wallet2 pairs where NEITHER filled (maker pull in flight — sibling consumption would leave the first filled, not both empty). The ~1-2-wallet depth cap is disproved directly: in the 4-wallet era (since 07-08 20:00Z) fills-per-directive is BIMODAL100 directives filled 0 of 4, 88 filled ALL 4 (71% all-or-nothing, only 29% mixed); the 88 all-four fills absorbed ~60-80 shares against the 658-share median touch. When price is there everyone fills, when it isn’t nobody does. The 400ms tier-2 stagger is COUNTERPRODUCTIVE as shipped: within mixed directives P(fill) by fire order = 72.4% / 51.3% / 38.2% / 28.9% — a time-spread-vs-moving-ask gradient, so staggering buys more of the exact thing killing the fill. Latency, not fan-out, is the failure mode: PM emit→first-retry-attempt p50 ~4-4.5s, p90 ~8s, stable every day since 07-02; the 4th wallet moved p50 only 4.38→4.54s; ground truth pmv2_autotrade_orders shows ~50% of live buys fail error_class='other' EVERY day since 07-02 (predates wallets 3/4). Every refused directive was fillable at emit by construction (max_price ≥ ask_at_emit + 0.02) → the ~50% daily refusals are latency-conversion losses at already-validated prices, not depth losses. BINDING — strategic: capacity is LATENCY-bound conversion, not depth-bound; the fee/MLM “hard capacity cap” premise is UNSUPPORTED at current clips (40-80 shares vs 658-share median touch) → agenda directions (a)/(d)/(e) rested on a premise now measured false; the lever is the limit/fill path + end-to-end latency, not wallet coordination. Two PM-side flags raised: pm_fak_attempts stopped writing ~06:57 while orders kept recording ~2h longer (silent telemetry gap → under-counts), and classify_error should give the FAK no-match string its own class instead of falling through to "other" (that fall-through is what enabled the original misdiagnosis). — pmv2-open-multi-wallet-liquidity-ceiling-brainstorm, crypto-shortterm-fak-nomatch-root-cause-2026-07-02, crypto-shortterm-phase0-depthprofile
  • Picked up PM’s multi-wallet “book size bug” brainstorm — appended COUNTER-EVIDENCE that the ceiling is PRICE, not book depth (OPEN, unverified). PM’s note (pmv2-open-multi-wallet-liquidity-ceiling-brainstorm) attributes heavy RecordedFailure("other") on wallets 3-4 to fan-out cannibalization (tier-1 FAKs exhaust the thin book, tier-2 hits an empty one). Three independent lines say otherwise: (1) FAK semantics — a thin-but-affordable book yields a PARTIAL fill, not an error; a hard no-match means nothing rested at-or-below our price = a price test, not a size test (PM directive_limit_price = min(max_price, fresh best_ask), crates/position_manager/src/pricing.rs:58book depth is never an input). (2) Prior measurement (crypto-shortterm-fak-nomatch-root-cause-2026-07-02, babylon #942, ~2,200 fires): ask side empty on 1 of ~2,200 no-matches; the dominant mechanism is the ask repricing ABOVE our limit within ~0.7s (19-25% UP, 29-38% DOWN) — and this predates wallets 3-4 entirely. (3) Arithmetic: 4 wallets × 10 shares = 40 shares ≈ 6% of the 658-share MEDIAN touch depth in the exact fire regime (τ∈[60,150)s, conviction ≥0.30; p90 5,109; 3¢ cum. depth 6,171 median — crypto-shortterm-phase0-depthprofile) → BTC moving in the final 60-150s explains the reprice; our own fills cannot. Consequences if confirmed: (a) the v1 400ms stagger may make tier-2 strictly WORSE — it walks tier-2 deeper into the reprice window against the same frozen shared max_price (#942: ~2/3 of repriced asks still above the old limit at +1-3s); (b) the strategic ceiling reframes to capacity ≈ touch depth ÷ clip ≈ ~65 wallets at 10-share clips if wallets fire within the fresh-price window → the binding constraint is fan-out latency spread, not depth = far less restrictive for the fee/MLM plan. Caveat: depthprofile publishes median/p90 but NOT p10 — a thin-tail minority of genuine exhaustion can’t be excluded; claim is about the dominant mechanism. BINDING — decisive query pending before ANY fix is built: join RecordedFailure("other") fires × pm_fak_attempts (ts_ms, best_ask_at_snap, fill_status, fill_price, filled_size) / book snaps — ask_at_fire > max_priceprice ceiling (fix = fresh per-wallet pricing, stagger rethink); touch_size < clip or partials present ⇒ true depth exhaustion (fix = size-aware coordination / partitioning / capacity caps). Inference from code + prior measurements, NOT a fresh measurement.pmv2-open-multi-wallet-liquidity-ceiling-brainstorm, crypto-shortterm-fak-nomatch-root-cause-2026-07-02, crypto-shortterm-phase0-depthprofile
  • SCHEMA GOTCHA documented — crypto_shortterm.observations.window_id is NOT asset-unique; any join to observations/outcomes needs an asset predicate. window_id = bare window-open Unix epoch (format!("{}", open_ts().timestamp()), observe/src/lib.rs:165), identical across every asset lane (btc/eth/sol/xrp/doge/bnb + btc15/eth15) at the same instant; observations PK is (asset, window_id, as_of) (migrations/0007_multi_asset.sql:3; outcomes = (asset, window_id)). A join on window_id alone + ORDER BY abs(Δt) LIMIT 1 returns a near-arbitrary lane’s row. The correct reference is depthprofile.rs:72 (WHERE o.asset = $1). emitted_signals has NO asset column (PK (source, window_id); was (window_id) alone in 0003); asset lives only in source via an IRREGULAR map (ingest/src/asset.rs) — BTC→crypto_5m_algo_up (prefix crypto, NOT btc), btc15→btc_15m_algo_up label btc15, eth15→eth_15m_algo_up — so split_part(source,'_',1) mislabels btc/btc15/eth15 (corrects the crypto-shortterm-pnl-attribution-corrections-2026-07-03 heuristic); use an explicit CASE. Real impact (2026-07-09): scripts/depth_guardrail.sh (commit 326bc2f) omitted the predicate → asset-blind subquery reported median_touch=33 (vs BTC’s in-regime ~658) → tripped alert=TRUE and drove a spec rewrite to rev 2.1. Diagnostic signature: per-asset medians suspiciously uniform across assets known to differ by an order of magnitude ⇒ suspect a missing asset predicate (grouping the outer query by the fire’s asset does NOT fix an asset-blind subquery). Adjacent facts verified: emitted_signals.size was DROPPED in migrations/0005_emitted_signals_v3.sql (per-directive size unrecoverable); emit_ts + as_of are both timestamptz (instant-based, immune to the session-TZ=+02 hazard). — crypto-shortterm-observations-window-id-not-asset-unique, crypto-shortterm-pnl-attribution-corrections-2026-07-03, crypto-shortterm-phase0-depthprofile

2026-07-06

  • CORRECTION — pmv2_wallets.size_scale is BOOT-LOADED, NOT hot-read (fixes an earlier claim in crypto-shortterm-per-wallet-size-scale-2026-07-06). Verified from code + live DB: main() in position_manager/crates/position_manager/src/bin/position-manager.rs builds a scale_by_wallet map (~L204) and passes each wallet’s scale as a static Decimal into consumer::run_wallet(…, scale) (~L322) which runs for the process lifetime — so a DB write to pmv2_wallets.size_scale does NOTHING until position_manager is RESTARTED. EVIDENCE (2026-07-06): @deploy set size_scale 1.0→2.0 for both wallets (DB confirmed 2.0) but live fills stayed filled_size=10 / size_usd≈$8 in pmv2_autotrade_orders (doubled would be ~20 / ~$16) → the running process kept the boot value. CONTRAST: source mode on pmv2_autotrade_sources IS hot-read per directive (why mode toggles go live without a restart) — size_scale is the exception. BINDING for babylon pmv2 #1065: the size_scale command must ALSO trigger a wallet-registry reload (or the consumer must re-read per order) or it writes the DB and the running executor silently ignores it. Full sizing: size_usd = limit_price × size_shares × size_scale, size_shares = ceil(min_size) × EMIT_SIZE_MULTIPLIER (=2 → 10 shares) → TWO compounding size levers: EMIT_SIZE_MULTIPLIER (crypto emit env, needs an emit redeploy) + size_scale (DB, needs a position_manager restart). — crypto-shortterm-per-wallet-size-scale-2026-07-06
  • EMPIRE UI = the Levandor CRM — crypto lane-ops is a NEW CRM TAB, not new React (topology reference). The central “empire” UI is the Levandor CRM (React + Vite + Supabase + Clerk, built by @crm); it already renders pmv2_*/pm_* tables directly (via as never casts — absent from database.types.ts) in a Polymarket surface (tabs Scout / Copy Lane / Longshot / Bets). One Supabase project holds everything: mgmt ref mkofmdtdldxgmmolxxhc hosts the CRM, the pmv2/pm_* tables AND our isolated crypto_shortterm schema (crypto_algo .env SUPABASE host = same ref); the CRM reads the DB directly. Implication: a crypto lane-ops UI = one more tab; crypto’s deliverable = public SQL views/RPCs (bridge crypto_shortterm→PostgREST — likely via public, the schema is isolated) + a cash-truth ingester + a CRM_DASHBOARD_PROMPT.md-style build-prompt for @crm → likely no React written by crypto. Code pattern web/src/lib/* + web/src/lib/hooks/* + web/src/pages/polymarket/* w/ unit tests. Pending: exact CRM repo path/remote; is @crm on babylon?; is crypto_shortterm PostgREST-exposed or must bridge via public. Source: polymarket_fetch/CRM_DASHBOARD_PROMPT.md. — crypto-shortterm-empire-ui-crm-topology-2026-07-06
  • PER-WALLET BET SIZING ALREADY EXISTS in pmv2 — pmv2_wallets.size_scale; no crypto code needed (reference). pmv2_wallets carries per-wallet size_scale NUMERIC NOT NULL DEFAULT 1.0 (per-wallet GLOBAL multiplier — BOOT-LOADED, NOT hot-read: a DB change needs a position_manager RESTART; see the 2026-07-06 correction entry above) + per-wallet mode (off/dry_run/live — already a ledgered command) + caps daily_spend_cap_usd/total_exposure_cap_usd; pmv2_source_wallets only maps source→wallet (NO per-source scale). So scaling crypto’s bet = one DB column — no crypto code, no migration — only a command surface. Requested from @positionmanager on babylon pmv2 #1065: a size_scale command as the sibling of the mode command (writes pmv2_wallets.size_scale, ledgered like mode_control_ledger, clamped [0.5, 3]; the command MUST force a wallet-registry reload — a bare DB write is ignored by the running executor until restart). GOTCHA: a size_scale bump must be matched by that wallet’s daily_spend_cap_usd or the cap (a reject-GATE, not a resizer) silently clips it. Use case: de-risk the ~2026-07-09 week-read 2× decision as a control experiment — wallet1.size_scale=2.0 (= 2× across ALL crypto; both wallets crypto-dedicated) vs wallet2=1.0 baseline. Composes atop the producer-side per-trade size_shares (crypto-shortterm-phase1-bet-sizing); source migration 20260701000000_pmv2_wallets.sql. — crypto-shortterm-per-wallet-size-scale-2026-07-06
  • P&L MEASUREMENT — on-chain realized CASH is the truth; three metrics disagree (reference). Trustworthy realized P&L = on-chain cash = Σ(REDEEM)+Σ(SELL)−Σ(BUY) from the Polymarket data-api /activity per proxy wallet (net of fees, NO mark-to-market) — persisted scripts scripts/wallet_lane_cashflow.py (per-lane; maps redeems→lane via the buy’s conditionId) + scripts/wallet_pnl_reconcile.py. The PM lb-api/user-pnl-api/UI “profit” is INFLATED 2–6× (unrealized mark-to-market + gross of fees — wallet1 UI ~90); do NOT use the UI number as realized P&L. realized_pnl_usd is correct but LAGS at resolution then reconciles — the wallet2 “2× under-booking” (babylon #1054) was a BOOKING LAG, not a code bug (@positionmanager: the winner-flag sweep is wallet-agnostic, can’t be wallet2-only; overnight booked 288). NOT the #1031 mechanism. Confirming query: 836/837 on-chain conditions matched PM-placed → also refutes the “Balu manual activity” hypothesis (majors algo-dedicated, size_scale=1.0, only 1 manual Paraguay bet). Current state: combined realized cash **0 idle USDC. — crypto-shortterm-pnl-onchain-cash-truth-2026-07-06
  • TIME-OF-DAY FILTER — NO-GO (in-sample great, OOS reversed). In-sample (July live ~5d) the algo looked like it “printed overnight” (01–05 CEST +245, hour 15:00 5), overnight (01–05) **−0.018/fire (−10)** — hour-of-day P&L flips sign between periods = **noise, not structure**. BNB May showed no US-open dip; the July −165 was period-specific (Jul 4 −$124). Do not arm a time filter — 4th in-sample-great/OOS-fail this session (after raise-EMIT_MIN_EDGE, disagreement-ceiling, don’t-chase). Durable lesson: the edge is the [60,150)s pocket; the market efficiently prices every overlay (selection, execution, time-of-day). — crypto-shortterm-time-of-day-filter-nogo-2026-07-06
  • 15m MARKET — efficiently priced, NO tradeable regime (early NO-GO; keep observing). First read of our btc15/eth15 observe data (332 windows each, 5d): the fair-value model HAS strong signal on 15m (p_fair monotonic; conviction p_fair≥0.80 windows resolve up 98% actual; under-confident tails like 5m). BUT economics kill it — the market ask ≈ actual win rate across the ENTIRE τ curve (ask 0.95–0.98 vs actual 0.98–1.0; real_edge ±1% everywhere); no τ where the ask meaningfully lags (unlike 5m’s [60,150)s). The two bigger apparent edges are unreliable: <120s (execution-impractical) and 600–720s (n=290, thin, contradicts prior tape). 15m signal ≠ edge, market efficient — don’t build a 15m money lane; keep the observe container banking data (early-τ pocket is the only maybe). — crypto-shortterm-15m-observe-nogo-2026-07-06
  • BNB ARMED LIVE 2026-07-05 13:22 CEST — before the actual-label / calibration-transfer confirmation. After the ~1h dry period bnb_5m_algo_up/_down were set live/enabled in pmv2_autotrade_sources (BNB promoted past observe-only-by-construction — the emit {Btc,Eth,Sol,Xrp,Doge} allowlist must have been extended). Early live: ~breakeven, ~21 resolved. EARLY ADVERSE-FILL HINT (watch): in 17–19 CEST BNB signal won 10/11 (91%) but fills won only 7/11 (64%), filled at 0.668 (cheapest entry) = the btc-up adverse-fill signature — if it holds, BNB has the same execution problem. Attribution caveat: the wallets carry some manual BNB trades (Balu), so a live BNB lane shares the wallet with manual BNB → muddied attribution. Documented as an ARMED-LIVE update on the BNB note. — crypto-shortterm-bnb-expansion-eval-2026-07-05, crypto-shortterm-lane-state-adverse-fill-2026-07-05
  • HARNESS ROBUSTNESS FIX — Chainlink eth_call had no HTTP timeout (backtest hung). ingest::chainlink::eth_call_result made its eth_call with NO reqwest timeout → the backtest_run on-chain Chainlink round-walk HUNG indefinitely on the free publicnode RPC. Fixed on the analysis worktree: reqwest 12s timeout + 4 retries with backoff (makes any harness Chainlink walk reliable — worth upstreaming to main). Remaining slowness: the walk uses a fresh client per call (no connection reuse) → slow for long ranges (a week of BTC = too many sequential RPC calls); connection pooling is the next optimization if long backtests are needed. — crypto-shortterm-chainlink-harness-timeout-fix-2026-07-06

2026-07-05

  • BNB OBSERVE LANE IS LIVE (milestone) — code merged + observe container rolled + recording VERIFIED end-to-end. @deploy rebase-merged feat/bnb-observe to main (commit 336bebea, “PR #1”) and rolled the observe-only ASSET=bnb container. Recording VERIFIED from the crypto_shortterm read path: asset='bnb' = 786 rows / 4 windows (07:55–08:11Z), avg K 575.08 (matches the on-chain Chainlink BNB/USD ~$571–575 I verified → aggregator 0x82a6c4AF830caa6c97bb504425f6A66165C2c26e correctly wired), avg p_fair 0.568 with ZERO nulls (model computing on BNB), book asks 0.53 up / 0.41 down captured. Observe-only confirmed (no bnb emit container, by construction) → the code change is validated live and the calibration-transfer clock has started. NEXT = bank ~a few days then run the frozen-BTC-calibration ~2 SE transfer check vs the 1009-window label set (data/bnb5m_actual_winners_2026-07-05.csv) = the BNB go/no-go. babylon pmv2 #1052 (deploy) / #1053 (crypto confirm). — crypto-shortterm-bnb-expansion-eval-2026-07-05
  • btc-up DISARM CONFIRMED by @deploy (babylon pmv2 #1050). crypto_5m_algo_up = dry_run in DB, PM rehearses only (DryRehearsed), last REAL submission 2026-07-04 20:01Z, all other lanes live — matches my independent read (live btc-up last fired 19:58:03Z). Caveat re-posted #1051: because btc-up’s loss is ADVERSE FILL, not signal, the DryRehearsed btc-up P&L will look profitable = a MIRAGE (dry fills at the intended price, never models the adverse fill) → do NOT re-arm btc-up on dry-rehearsal results; any re-arm needs a real fill / execution fix. — crypto-shortterm-lane-state-adverse-fill-2026-07-05
  • LIVE-ALGO PERFORMANCE DASHBOARD built (per-lane W/L + net P&L). Data source = pmv2_autotrade_orders.realized_pnl_usd (CLOB settlement, net of fees). Two reusable scripts, currently in the SESSION SCRATCHPAD (ephemeral — flag to persist to repo scripts/ if kept): dashboard_data.py (pulls overall/daily/per-lane/per-asset metrics as JSON from the live DB via duckdb+postgres READ-ONLY) + build_dashboard.py (re-bakes the JSON into a self-contained dark-themed HTML). Deployed as a claude.ai artifact = SNAPSHOT, NOT live (Artifacts have no network access → refresh = re-run build_dashboard.py + redeploy). Snapshot at build: booked net ~+307; xrp/doge both-sides green (~78%), sol-down/eth-down the soft down-side cluster, btc-up the −$171 drag now dry. — crypto-shortterm-perf-dashboard-2026-07-05, crypto-shortterm-lane-state-adverse-fill-2026-07-05
  • BNB 5m EXPANSION GROUNDWORK — feeds + market verified, observe-only code shipped, calibration-transfer is the open go/no-go. Prompted by the historian noticing BNB passes every documented gate yet was skipped by inheritance (the Asset enum copied cohort_algo’s 4 majors + doge), not by decision. (1) The 5m Polymarket universe is EXACTLY 6 assets — btc/eth/sol/xrp/doge + bnb, nothing else: probed 38 candidate coins (all top-30 + major memecoins) via Gamma events?slug={asset}-updown-5m-{ts} (series pattern {asset}-up-or-down-5m; the tag_id endpoint is a dead-end schema blob), only these 6 have active 5m markets — matches the harness spec’s candidate list → BNB is the ONLY untapped 5m asset (next breadth axis is 15m, not more 5m coins). (2) Feeds verified: Binance bnbusdt exists; Polygon Chainlink BNB/USD aggregator 0x82a6c4AF830caa6c97bb504425f6A66165C2c26e VERIFIED ON-CHAIN (eth_call description() = “BNB / USD”, ~128k** (just below doge 153k, our smallest lane); deeper/more continuous than expected — labels saved `data/bnb5m_actual_winners_2026-07-05.csv`. **(4) Code shipped — `feat/bnb-observe` @ `5673d6e` (pushed, NOT merged):** adds `Asset::Bnb` observe-only + **CRITICAL SAFETY HARDENING** — `emit_supported` was gating only on `window_secs==300`, which would NOT have refused a new 5m asset like bnb (unlike the 15m series); changed to require an explicit validated-asset allowlist `{Btc,Eth,Sol,Xrp,Doge}` AND `window_secs==300` → **bnb is observe-only by construction (emit hard-refuses at boot) and NO future 5m asset auto-arms on being added to the enum** (must be explicitly promoted after passing its gate); test `emit_arms_only_validated_5m_assets`; all 4 gates clean. **(5) Coordination — babylon `pmv2` #1049:** asked @deploy to merge + roll ONE observe-only `ASSET=bnb` container (no migration/NATS/PM rows), explicitly NOT arm. **(6) NEXT / OPEN — the go/no-go = does the frozen BTC calibration map transfer to BNB within ~2 SE** (same test that cleared the 4 alts); needs BNB `p_fair` (only from a live observe container) → deploy the container (few days) then one-query check vs the 1009-window labels, OR a Binance+Chainlink harness backfill for a faster read. BNB expected doge-sized (~5–30/day) — small but near-free if calibration transfers. Speculative plus: BNB is less efficient than btc, so our model’s disagreement may be LESS adversely-selected than btc’s (ties to the btc-up adverse-fill finding). — crypto-shortterm-bnb-expansion-eval-2026-07-05, crypto-shortterm-multi-asset-status-babylon-2026-07-02, crypto-shortterm-lane-state-adverse-fill-2026-07-05, crypto-shortterm-algo-accuracy-audit-2026-07-02
  • OPERATIONAL STATE + btc-up = ADVERSE FILL, not a broken signal (REVISES the earlier “adverse selection at signal level” framing incl. babylon pmv2 #1048). State (2026-07-05): btc-up (crypto_5m_algo_up) set to DRY 2026-07-04 ~20:13 UTC (Andras, #1048) — NOT a full kill: live BUYS off but emit STILL FIRES btc-upmode='dry_run' rows in pmv2_autotrade_orders (filled_size=0, realized_pnl_usd=NULL; intent logged, no fill/no P&L) → live btc-up losses stopped, signal continues for shadow. sol-down (sol_5m_algo_down) LIVE, ON WATCH — pre-registered rule: re-eval on its first FULL day (Jul 5 complete), full-day WR <70% → confirmed decay → dry-shadow, ≥70% → thin-but-stable → leave (small mag ~−$20–40/day; full days Jul 3=72%/Jul 4=71%; Jul 5 partial 22 fills 63.6% too noisy). All other 8 lanes live/armed unchanged. ROOT CAUSE (decisive test, Jul 3 actual CLOB winners CSV, btc-up fires): the 8 we FILLED won only 25%; the 14 we FAILED to fill (FAK no-match, ask ran away) would have won 100% — the btc-up SIGNAL is fine (~73–84% actual on fired windows); we systematically FILL the losers (ask stays cheap = market knows it loses) and MISS the winners (ask runs away). So btc-up’s live loss is EXECUTION / adverse fill, NOT signal. BINDING: dry btc-up will LOOK profitable because dry ignores adverse fill → DO NOT re-arm on dry results — it’s a mirage; the real fix is execution-side (don’t fill the toxic cheap-ask fires) and needs the fill harness (@simulator). Caveat n=22, one day; a sharper filled-vs-failed-labeled version of the disagreement-ceiling execution-toxic finding + the FAK-reprice thread. Contrast — sol-down is the OPPOSITE: fill is fine (filled 71% vs failed 80%, ~representative), signal marginally below breakeven (71% vs ~74% needed at 0.72 entry) = thin/decaying signal, NOT adverse fill → different disease, hence watch not kill. — crypto-shortterm-lane-state-adverse-fill-2026-07-05, crypto-shortterm-disagreement-ceiling-nogo-2026-07-03, crypto-shortterm-fak-nomatch-root-cause-2026-07-02
  • DATA FACT CONFIRMED — realized_pnl_usd IS CLOB settlement truth (settles the label question). Spot-check 2026-07-05: 24 of 24 sampled windows, CLOB tokens[].winner matches the SIGN of realized_pnl_usd exactly (0 mismatch)realized_pnl_usd = the actual CLOB market resolution, net of fees (fee baked into cost_basis). The ~9–15% near-tie noise applies ONLY to resolved_up (our on-chain settle≥K proxy), NOT to realized_pnl_usd. BINDING: live P&L tables built on realized_pnl_usd are label-truth-grade — no CLOB backfill needed for them (backfill only helps on-chain/resolved_up-based analyses like the dry-btc-up shadow and signal parquets). Confirms + sharpens the pnl-attribution corrections; recorded in the data-sources note. — crypto-shortterm-lane-state-adverse-fill-2026-07-05, crypto-shortterm-data-sources, crypto-shortterm-pnl-attribution-corrections-2026-07-03

2026-07-03

  • DISAGREEMENT-CEILING SHAVE → OOS NO-GO (directional) — the disagreement tail is EXECUTION-toxic, not selection-toxic. Live attribution (Jul 2–3, realized_pnl_usd) found ~60% of fills are high-disagreement flow (ask ~0.62 vs p_side ~0.865) realizing 0.64 win vs 0.85 in the agreement zone = the entire booked loss; hypothesis = a SELECTION ceiling (skip ask<~0.70 / modeled edge ≥~0.07). Ran the shave through the refit/exit walk-forward rig on BTC (21-day actual-resolution parquet, won_up via extract_refit_btc.sh + winners CSV, NOT resolved_up; day-folds + clustered bootstrap + deflated_threshold(6,1.64)=3.53): NO-GO on all 6 grid combos. Baseline at-ask EV +28.1 / 4,965 OOS windows / 19 days (matches the refit study’s universe); every ceiling LOWERS at-ask EV (tightest max_edge<0.05 → −7.3), best Δz −1.45 (need +3.53), no majority-positive folds — a DIRECTIONAL rejection. Why: low-ask fires pay more per win (buy 0.62 → win pays 0.38) → +EV at the ask, a ceiling deletes good flow; the live loss is EXECUTION — the tail fills 3.3c ABOVE the emit-ask (break-even limit df6f500 chasing the reprice = ~2.3c/share drag) vs 1.7c BELOW on the core, plus depressed FILLED win 0.639 vs 0.674 signal (adverse fill selection); ties to the FAK-reprice finding (ask lifts within ~0.7s on 19–38% of fires, DOWN worse). BINDING: edge-band fire selection is dead BOTH ways (FLOOR EMIT_MIN_EDGE↑ already wrong; CEILING now OOS-rejected — every band ~breakeven at the ask); the lever is the limit/fill path (cap fire_limit near ask on low-ask fires, don’t chase), NOT the selector; the at-ask harness CANNOT validate an execution fix (needs a fill/repricing model or live shadow A/B). Branch feat/disagree-study (worktree; refitlib first_fires_ceil + bin/disagree_study.rs, 4 gates clean, NOT committed pending review); report docs/superpowers/disagree-study-btc-report.md; spec docs/superpowers/specs/2026-07-03-disagreement-ceiling-design.md (→NO-GO); babylon pmv2 1046. — crypto-shortterm-disagreement-ceiling-nogo-2026-07-03, crypto-shortterm-pnl-attribution-corrections-2026-07-03, crypto-shortterm-fak-nomatch-root-cause-2026-07-02
  • P&L ATTRIBUTION — two corrections (live P&L session + babylon pmv2 #1032–#1036). (1) Data trap: crypto_shortterm.outcomes.window_id is NOT unique — it is shared across all 5 assets for the same 5-min bucket (each asset resolves independently; ~9,643 rows / ~5,995 distinct window_ids; (asset, window_id) is the unique key, with duplicate rows per pair). A window_id-only join collapses all assets → coin-flip labels (>50% of merged windows self-contradicted on resolved_up). ALWAYS join on (asset, window_id); dedupe bool_or(resolved_up) GROUP BY asset, window_id; asset = crypto_5m_algo_*→btc else split_part(source_id,'_',1). (2) #1031 is ZERO-DELTA (REVISES prior handover/LOG): the “PM realized_pnl_usd biased NEGATIVE / 7 false losses” claim is REFUTED — @deploy CLOB-re-checked all 7 (tokens[].winner = payout oracle = ground truth), we LOST all 7; PM’s −cost_basis booking was CORRECT; the artifact was a flipped token-side cross-ref in an earlier crypto-side session, NOT a PM bug. Root mechanism: our outcomes.resolved_up (on-chain Chainlink settle≥K) over-books near-tie wins vs the market’s Chainlink Data Streams settlement — on the 2026-07-03 doubled-vol day ~15% of resolved fills disagreed resolved_uprealized_pnl_usd, concentrated entirely in sub-5bps near-ties (disputed median |settle−K| 0.9–4.8bps vs agreed-win 8.9bps). BINDING: realized_pnl_usd (CLOB/winner-flag) is the trustworthy P&L label; resolved_up is a trading GATE, not settlement truth — never for P&L/sizing/certainty tiers (reconstructing from resolved_up is optimistic: showed Jul-3 +88). Reinforces the Data Streams licensing track as the real label fix. Coverage: realized_pnl_usd only from 2026-07-02 on (no pre-tweak settlement-grade P&L); the two booked days net ≈ −34, Jul-3 −$88). — crypto-shortterm-pnl-attribution-corrections-2026-07-03, crypto-shortterm-data-sources, crypto-shortterm-algo-accuracy-audit-2026-07-02
  • VALIDATION MATRIX FULLY CLOSED — both pending cells of the competitor dissection resolved same-day (§C8 mechanism NAMED + on-chain leg ALL PASS). §C8 mechanism named and reproduced: survivorship by activity-dependent tape truncation (outcome-correlated coverage-depth censoring) — the original pipeline re-run on the frozen tapes reproduces the 242-0 output verbatim (242 fills/55 markets/242-242 winner-side; tape byte-identical). Stage A: a market is tape-visible only if ≥1 fill survives the last-~1000-print cap, and coverage depth tracks outcome — his losses are late-flip volume storms (coverage dies ~70s before close; entries ~150s out) while wins cruise (coverage ≥127s) → 0/56 lost vs 55/129 won markets survive (zero-error prediction); Stage B: survivors tautologically winner-side (same 55 markets’ late slice only 71% winner-side; 242/242 by chance ≈1e-36). True window record 129W/56L (69.7%); tape credited +6,998. C7-vs-C8 “contradiction” dissolves — different conditionals (fill-level on busy markets skews losing vs market-visibility over all markets skews winning), both correct → name the conditional before comparing rates (methods lesson). ON-CHAIN LEG ALL PASS (data/validation_onchain.md): Polygon trading-net **−87,755 (redemptions 114.7k − withdrawals 7.1k) → genuine self-funded net loser; fee formula 1,837/1,837 nonzero fills within 2% (median rel-err −8e-6) + ~Jun-13 cutover independently confirmed; 3 sample days ≤3e-6 USD; leak check clean; last alternative REJECTED — no loss-absorbing paired/sibling wallet (maker counterparties diffuse: max single-maker share 19% on one day, 174–495 distinct makers/day; the one recurring maker is itself an episodic Polymarket proxy). Every finding now has ≥2 independent no-shared-code derivations, blockchain as the final non-Polymarket source. Canon updated to matrix-complete: docs/superpowers/winner-dissection-2026-07-03.md. — crypto-shortterm-competitor-wowitsamazing-refuted-2026-07-03
  • COMPETITOR DISSECTION BULLETPROOF-VALIDATED — pre-registered 3-leg validation (clean-room + adversarial + on-chain, closed above) CONFIRMS every core finding; “zero selection edge” sharpened to “no FEE-CLEARING edge”. At 100% CLOB-truth label coverage (1,621 pre-Jun-8 labels backfilled; label audit 200/200): overall win−price +0.22¢, 95% CI [−1.48,+1.77]¢ — any BTC-5m edge >~1.8¢ excluded, day-clustered; cash components reproduce exactly (net 90k across three accountings — cleanroom’s deeper paging found ~289 extra fills, loss slightly deeper); fee identity confirmed + fee CUTOVER ~2026-06-13 discovered (zero-fee subset = pre-fee era); effective fee 1.72% after ~16% taker-rebate clawback (+22k, pre-window dead). STRATEGY IMPACTS: live net-edge gate’s fee term empirically validated on ~190k fills; EMIT_MAX_ASK 0.85 + edge-vs-ask independently re-validated from the book side (book-calibration audit: buyable zone efficiently priced, favorite win exceeds ask by ≈ the fee in every bin; 0.95+ “overpricing” = stale-wide-quote artifact, realized 0.468 @ ask≈0.99; selector kill test: no slice of his flow beats the ask OOS); size_hint certainty tier must re-base on our own fills (+ actual-resolution calibrations per the refit study); NEW follow-up: check whether OUR wallet accrues taker rebates (free EV if yes); sizing week-read/refit/15m queue unchanged. METHODS (reusable): three headline claims died on independent reproduction (v1 15m “+7.5¢” band = selection lookahead via whole-window share-weighted cushion; blended-price “+4.5¢” = aggregation artifact → −1.1¢ at matched cushion; “58% vs 55.3¢” = mixed weighting) — the pre-registered 3-leg design caught every one; frozen post-settlement tapes are row-exact reproducible (10/10) → freeze tapes at study time. INCIDENT: dissect_discriminator.py briefly clobbered data/0xd02b_activity_full.csv with an 8h slice (restored from raw JSON; script repointed to 0xd02b_activity_recent8h.csv); two validators read the stub → check file mtime when a claim contradicts a recent measurement. Canon: docs/superpowers/winner-dissection-2026-07-03.md (+ erratum prepended to competitor-analysis-2026-07-02.md, HANDOVER.md updated, data/validation_{cleanroom,adversarial}.md + all data/*findings*.md). — crypto-shortterm-competitor-wowitsamazing-refuted-2026-07-03, crypto-shortterm-polymarket-data-api-gotchas, crypto-shortterm-algo-accuracy-audit-2026-07-02
  • data-api gotchas note extended with 4 new reusable gotchas (hardened by the validation legs): capped /trades?market= tapes unusable for per-wallet PnL in BOTH directions (1000-print recency cap + ~32% maker-invisible → ~9% taker-visible; 29%-vs-54% fill-level skew; per-wallet-z methodology retracted); fee-era cutover ~2026-06-13 (split fee analyses by era); rebate program → effective fee 1.72% not 2.06% (+ the our-wallet rebate follow-up); frozen post-settlement tapes row-exact reproducible (10/10) + the shared-file clobber/mtime lesson. — crypto-shortterm-polymarket-data-api-gotchas
  • “CERTAIN WINNER” COMPETITOR REFUTED — the sole skill-verified wallet (0xd02b…/wowitsamazing/Novel-Plow-Toot) has ZERO selection edge and is a NET LOSER over its 38-day life. Full lifetime harvest (198,988 activity rows, 2026-05-26 → 07-03) overturns the 24h-tape headline (z=4.29, +84,112** (nominal +7.6k — every method breakeven-to-losing). True directional win rate 58.1% (dominant-side, CLOB-truth), calibrated to the 55.3¢ he pays in every price×τ cell ⇒ zero edge, pre-fee PnL ≈ 87,540 (per-fill ratio ≈1.0, corr 0.9915; NOT a bid-ask spread); rebates claw back +35,690 (donates the fee, doesn’t extract); 14 arms 24/7, ~14 clips/mkt (median 1s), hold-to-resolution (305 sells lifetime), ~200× scale in 5wk, alts added 06-23. KEY: that fee is the exact 0.07·eff·(1−eff) term in our live net-edge gate (trade_decision_slipped/EMIT_MIN_EDGE) → this wallet empirically VALIDATES our gate on ~190k real fills (zero-edge fee-unaware flow bleeds precisely the fee). Corroborates the same-day oracle-state Phase-0 kill; NOT a threat to our $400 roll. — crypto-shortterm-competitor-wowitsamazing-refuted-2026-07-03, crypto-shortterm-algo-accuracy-audit-2026-07-02
  • Polymarket data-api /activity harvest gotchas (reusable). (1) offset hard-caps at 3000 → time-page via start/end unix secs (newest-first, end inclusive, browser UA, binary-search first-activity ts, parallelize by time-chunk); (2) usdcSize includes the Polymarket taker feeusdcSize is REAL USDC moved, price is the nominal pre-fee price; the gap is the fee 0.07·x·(1−x)·size (per-fill ratio ≈1.0, corr 0.9915), NOT a bid-ask spread → use usdcSize for true PnL; PM’s own lb-api/profit + user-pnl-api/user-pnl compute off nominal/marks (exclude the fee → flatter high-freq takers). REDEEM rows attributable via conditionId/slug (asset=”/outcomeIndex=999); redeem-presence is a valid per-market OUTCOME proxy but OVERSTATES the directional win RATE (two-sided hedging always redeems — use CLOB-truth dominant-side labels for a rate). — crypto-shortterm-polymarket-data-api-gotchas
  • CALIBRATION REFIT STUDY (build-queue item 0) → NO-GO at the deflated bar — frozen calibration.json stays live. Branch feat/calibration-refit @ 2cf12d8 (pushed); report docs/superpowers/refit-study-btc-report.md; repo /Users/levander/coding/pmv2/crypto_algo. BTC OOS 4,965 windows / 19 days, purged+embargoed expanding day folds (exit-study convention), 16 pre-registered variants = {chainlink-σ, binance-σ} × {pooled, τ-bands} × {single, per-side} × {± print-state}; labels = ACTUAL market resolutions; the frozen map scored identically as control. Best variant Δz = −0.32 vs deflated bar 3.99 — nothing close; every refit variant fires ~35% fewer times than frozen and earns statistically-tied-or-less EV. The 2026-07-02 “0.75–0.85 dead zone = biggest calibration prize” is NOT confirmed — a sharper map refuses the zone rather than unlocking it. — crypto-shortterm-algo-accuracy-audit-2026-07-02, crypto-shortterm-exit-loss-cut-study-2026-07-01, Crypto 5m Up-Down — Oracle-Lag Strategy (Canonical Spec)
    • DURABLE FINDING (binds work beyond this study): resolved_up (on-chain settle ≥ K) labels inflate per-bin win rates by 4–7pt exactly in the high-conviction bins — near-tie flips land in those bins’ loss column. The frozen map claims 0.91–0.99 across bins [0.70..1.00); actual-resolution rates are 0.83–0.94 (stable early-vs-late era); frozen p_cal ≥ 0.95 actually wins ~0.92. BINDING: the certainty-tier in the v2 size_hint spec and any conviction-keyed sizing MUST be built on actual-resolution calibrations, never frozen p_cal; break-even limit-margin bookkeeping is optimistic by the same inflation.
    • Mechanism of the NO-GO: honest (refit) maps yield lower p_cal → the net-edge gate then demands cheaper asks → they refuse the thin-edge 0.75–0.85 fills the frozen map’s inflation takes; those fills realized ≈ +1.2c/fill (indistinguishable from zero at the day-cluster level). A sharper map does not “unlock” the dead zone — it declines it.
    • Variant verdicts: H2 (per-side maps on p_fair bins) proven structurally vacuous — an Up/Down split selects exactly the rows each p_fair bin already isolates (bit-identical results); a real per-side idea needs a different input (p_side at fire) + a new spec. H4 (print-state conditioning) actively harmful under the live rule (all Δz −0.8 to −2.7) — consistent with the oracle-state Phase-0 kill (market already prices print state). H3 (binance-σ, 10s EWMA λ=0.94 from binance_ticks) = best runner-up (Δz −0.32, better log-loss than frozen) — the only direction worth a future study.
    • Data assets (now in repo data/): btc5m_actual_winners_2026-07-03.csv rebuilt clean (5,925 windows; two concurrent backfill runs had interleaved one corrupted CSV row → dedupe now validates field shape); alt winners {eth,sol,xrp,doge} ~830 windows each (3 days). Alts confirmation N/A on a NO-GO.
    • Ops: refit worktree wt-refit (session 899069e2 scratchpad) with feat/exit-study merged in (merge 2396c1a — exitlib fold machinery; semantic conflicts from the 15m window_secs API fixed; inherited clippy lints cleared); extraction via duckdb→parquet scripts/extract_refit_btc.sh; refit_study bin deterministic (seeded bootstrap 2000×, seed 42).
  • BTC 5m actual-winners backfill COMPLETE — ground-truth winners dataset for the exit-study. scripts/backfill_actual_winners.py (Gamma API, in crypto_algo): two overlapping background runs both finished exit 0 and agreed on every overlapping window (0 winner conflicts); deduped to 5,931 unique windows (window_id, condition_id, winner), persisted to data/btc5m_actual_winners_2026-07-03.csv (replaces the earlier 4,698-row partial). Ground truth for the exit-study work on branch feat/exit-study. — crypto-shortterm-exit-loss-cut-study-2026-07-01
  • ORACLE-STATE PHASE-0 RUN → NAIVE DETECTOR KILLED (pre-registered criterion did its job). Grid: 4 ref-times × 5 cushion bands, 18,771 evals / 4,693 actual-resolution-labeled windows / 19 days, 14/5 chronological split. Detector precision ≈ quoted ask in every cell (market already prices on-chain print state; 90s/5-10bps: 0.820 vs ask 0.833); best cell +0.034/sh t=1.75 (noisiest bin); several cells NEGATIVE after taker costs. Competitor’s edge = selection WITHIN cells (his ~0.98 realized vs ~0.80 cell-level at identical asks; wallet z=4.29 survives, naive replication does not; Data Streams = prime discriminator suspect). SURVIVES: actual-resolution label fix (in production); print-state → candidate INPUT for calibration refit (our p_fair under-prices these windows the market prices right — a model gap, not free money); Data Streams research track. No tier ships. **Same morning: post-fix truth-labeled P&L +75.61 gap explained to the cent; #1031 filed with row ids; PM column biased negative until fixed). [REFUTED 2026-07-03 — #1031 is ZERO-DELTA: @deploy CLOB-re-checked all 7 (tokens[].winner = payout oracle = truth), we LOST all 7; PM’s −cost_basis booking was CORRECT; the “false losses” were a flipped token-side cross-ref in an earlier crypto-side session, NOT a PM bug. realized_pnl_usd is the trustworthy label; resolved_up over-books near-ties and is a gate, not settlement truth → crypto-shortterm-pnl-attribution-corrections-2026-07-03]crypto-shortterm-algo-accuracy-audit-2026-07-02
  • ORACLE-STATE TIER DISCOVERED — the competitor’s 100% record solved. Joined 0xd02b6d91…’s entry timestamps to our per-second recordings: at his entries our model saw p_fair 0.639 avg at ~5bps spot cushions (coin flips by diffusion) while he paid 0.795 and won 47/47. Mechanism: sign(last on-chain Chainlink print − K) at his entry predicts the ACTUAL market winner 45/46 (97.8%) despite prints continuing — the oracle print path is far stickier than GBM at 100-200s horizons; he conditions on settlement-oracle state. Resolution source verified from market description = Chainlink Data Streams BTC/USD (data.chain.link/streams/btc-usd) — the same source gap behind our ~9% label noise. Our conviction threshold excludes the region (we never fire there; nothing bleeding — uncollected). Spec written: docs/superpowers/specs/2026-07-03-oracle-state-tier-design.md — H1 print-sign band detector (on-chain proxy, buildable with ZERO new dependencies), H2 parametric flip-prob, H3 economics (ask>0.85 cap interplay); Phase-0 runnable now on existing cl_reports + actual-resolution labels; then shadow → gate → dry → arm; size_hint certainty multiplier belongs to this tier. Parallel: research Data Streams licensing. Risks pre-registered: selection-transfer (his picks ≠ unconditional), crowding, cap interplay. — crypto-shortterm-algo-accuracy-audit-2026-07-02
  • 15m recording VERIFIED live (btc15/eth15 ~1/s, 9 windows each in first 2h, books ~98%, flow 100%; all 12 containers on 860dddc0; babylon #1014 resolved).

2026-07-02

  • 15m BENCHING STARTED (Andras go): main = 85390aa, observe-only recording for btc15/eth15. window_secs moved into AssetConfig (Btc15/Eth15 variants @900s, slugs btc-updown-15m/eth-updown-15m, labels btc15/eth15); floor_window/make_window/polymarket discovery parameterized (polymarket.rs had a hidden second WINDOW_SECS for the candidate scan); emit hard-refuses non-300s assets at boot — no 15m money path can exist until a fresh 15m calibration passes its own full gate. Zero migration (label-scoped). Deploy task filed for 2 new observe containers + roll the 10 existing to 85390aa (behaviorally identical). 15m path: ~2-3wk data accrual → harness adaptation for the early win-rate read → fresh calibration + deflated-bar gate (band DERIVED, not copied — competitor buys ~500s-to-close; 600-900s and pre-window are tape-confirmed negative) → dry → arm. — crypto-shortterm-algo-accuracy-audit-2026-07-02
  • HEAD-TO-HEAD vs the sole skill-verified wallet (0xd02b6d91…), fires joined on condition_id. 38 shared BTC-5m markets (his 55 vs our 227 — he’s 4× more selective); 36/38 same side, both won all 36; the 2 disagreements were our extreme-disagreement fires (ask≈0.25) — he took the market’s side at 0.84 and won both (independent confirmation of the pred−ask>0.20 fade → refit shave). His shape: 5m certainty tier (242 trades, 100% wins, entry 0.797, 6.2k). Not a better model — a higher trigger, bigger clips, and an interval we don’t trade; our probability tier has equal-or-better ROT. Feeds: size_hint certainty tier (clip ∝ conviction), 15m expansion case, fade shave. — crypto-shortterm-algo-accuracy-audit-2026-07-02
  • COMPETITOR ANALYSIS (6-agent workflow, 379,732 trades / 12,715 wallets, 24h BTC 5m+15m tape). Headlines: the tape independently re-derives our lane as the market optimum (5m 0.65-0.85 @ 60-150s = +28.4% ROI t=15) while the CROWD in the same raw band is net-negative → our p_cal gate is the edge, the leaderboard is not a strategy menu. Only 1/12,715 wallets survives the selection-adjusted skill bar (z=4.29, +55-115 clips). Ranked: #1 size 2-4× (aligns with week-gated plan), #2 15m expansion (new workstream: WINDOW_SECS param + fresh 15m calibration + gate), #3 τ 30-60s widen (re-derive on Phase-0 data, tape censored), #4 cap 0.90 only after sizing (0.90-0.97@final-30s = −14% trap), #5 selection beats band-widening (65% of winner edge = same-price selection → validates calibration-refit priority). Dead ends confirmed: mid-band exits, sub-0.35 longshots, sub-60s entries, post-close sniping (−$15.1k aggregate), mint-and-sell, churn. DISSENT: workflow’s “hold DOWN” is a one-day crowd-tape artifact — our ledger shows DOWN lanes are our best; ignored. Tape caveat: 1000-print cap censors 5m to ~last 104s. — crypto-shortterm-algo-accuracy-audit-2026-07-02
  • CROSS-ASSET WITHIN-WINDOW LOSS CORRELATION — confirmed & quantified (operator noticed “when we lose we usually lose all 3”). Clean-era, all 10 lanes: outcomes across lanes in the same 5-min window are strongly positively correlated (same BTC-led move fires everything; one late reversal takes the group): 5-lane windows all-win 64% vs 50% under independence; losing multi-lane windows carry ~1.8 simultaneous losses (indep 1.1); literal all-lose = 0.9% of ≥3-lane windows (8 stakes). Concurrency EV flat-positive (win 0.87-0.91, +0.05-0.11/share in every 1→5-lane bucket) → correlation is VARIANCE not edge → no entry filter (skipping correlated fires skips correlated wins). Consequences: (1) sizing risk-unit = the WINDOW, not the fill (4-5-lane windows = 36% of fire windows ≈ one 400 bankroll (at 3×+alts one group-loss window ≈ 327** (fees in; wallet-level −1.04/fill, best bucket) → EMIT_MIN_ASK REJECTED; true dead zone = fills 0.75-0.85 (n=441, 80.0% win ≈ breakeven +$0.02/fill) → top calibration-refit target; post-fix era n=77 ≈ breakeven ±noise, week read governs. — crypto-shortterm-algo-accuracy-audit-2026-07-02
  • LOSS-ELIMINATION DISCOVERY SWEEP — VERDICT: residual losses are irreducible and fully priced; do NOT build loss-avoidance machinery. Tested every remaining mechanism on the clean era (n=2,866 fires, 341 losers, fee-inclusive net EV per bucket): (1) fire-time loss prediction is NULL — loser rate 11-14% in EVERY slice: s_hat mom30/mom60 (signed), side-book spread, best-ask size, Chainlink round age, BTC same-window agreement (alts), and composites; only −EV pocket = mom30<−5bps hard-retrace (n=60, 2% of fires, ~0.8 shares — negligible; may shrink further with per-event firing); (2) loss clustering NULL — P(win|prior fire lost)=0.882=P(win|prior won), no direction interaction → cooldown-after-loss dead; (3) exits already NO-GO ×2 (study), and Polymarket has no conditional orders → no resting-order stop-loss exists. After the edge gate (which already removed the only systematically-losing 38%), remaining losses = the priced risk premium of buying ~0.88-prob at ~0.80. The sweep’s actual alpha = EV conditioners for SIZING, not filtering: (a) Chainlink-round staleness: fresh ≤10s EV 0.044/WR 0.862 vs stale >30s EV 0.090/WR 0.898 (Binance lead worth more pre-print); (b) BTC-disagreement premium on alt fires: EV 0.153 (disagree) vs 0.031 (agree) at IDENTICAL 0.876 WR — agreement is priced in, disagreement is the private edge; (c) momentum extremes mildly adverse (calm-with-move best). Sizing tilt blocked on the #62 Directive (no size field) → spec size_hint into the pmv2-contracts v2 cutover; per-fire sizing study then needs exit-study-grade rigor (day-clustered, deflated, OOS). The legit “smart model” track remains the refit-era calibration work (Binance-tick σ, τ/side-conditional maps) — improves p_cal sharpness feeding the edge gate, NOT a loss classifier (data rejects one). — crypto-shortterm-algo-accuracy-audit-2026-07-02
  • PRODUCER FIXES IMPLEMENTED — feat/emit-edge-gate (4 commits, gates green, PENDING merge + redeploy). All 10 lanes now armed live (Andras), so the audit’s findings were implemented same-day, inline (spawned subagents kept dying — terminal-pane crash issue — so after two dead implementers the work was done directly in an isolated worktree off main): (1) net-edge fire gate restoredobserve::net_buy_edge(p_side, ask) = p_side − (ask+0.01) − fee must exceed EMIT_MIN_EDGE (default 0 = the exact walkforward.rs validated rule; re-validated fee-inclusive on n=2,686 fires: gated policy = 62% of fires, net EV/share +83%, net TOTAL +14%, dropped fires −21.2 shares net); (2) EMIT_MAX_ASK defaults 0.85 (was unset everywhere; 510/1,146 fires in the last 24h above 0.85, max 0.979; ask>0.85 subset net −EV); (3) per-event fire evaluation (removed the 1s due_obs quantization — up to 1s self-inflicted latency ahead of the 250ms itode delay; per-window dedup unchanged so same fire, just earlier; spammy warns throttled to 1/s); (4) Binance WS 30s idle-watchdog (tokio::time::timeout around ws.next() → reconnect loop fires on half-open connections; kills the #832 silent-stall class for observe+emit). Branch also repairs pre-existing rustfmt violations in 6 untouched files (main wasn’t fmt-clean); functional diff = observe/lib.rs+emit.rs+ingest/binance.rs. HAZARD found: shadow-mode emit still INSERTs into live emitted_signals (insert precedes the shadow check) → claims windows, suppresses live fires; never run emit locally against the live DB. PM retry spec v1 confirmed to @positionmanager (#948: k=0.03, ceiling 0.88, N=2, telemetry join on origin_ref; no EMIT_PRICE_BUFFER widen — superseded by these fixes). — crypto-shortterm-algo-accuracy-audit-2026-07-02, crypto-shortterm-fak-nomatch-root-cause-2026-07-02
  • ALGO ACCURACY AUDIT (main artifact) — live emit dropped the validated edge-vs-ask trade condition (CRITICAL fidelity gap). Live crypto_shortterm DB, clean both-sides era Jun 28–Jul 2, emitted_signals × outcomes. The Phase-0-gate-clearing rule (walkforward.rs:32 trade_decision_slipped: trade only when p_cal − (ask+0.01 slip) − fee(0.07·p·(1−p)) > 0) is absent live — emit.rs fires on is_regime(τ,p_cal) alone, never comparing p_cal to the ask. BTC n=920: 26% of fires have pred−ask ≤ 0 (avg ask 0.918, gross EV +0.002/share ≈ −EV net of fees); EV monotone in disagreement (+0.03 → +0.28/share at Δ>0.20, n=148, WR 0.818). Simulated gate pred−ask≥0.03: 920→521 fires, EV/fire 0.0905→0.1446 (+60%), gross total ≈ equal net of fees; ≥0.05: 440 fires, 0.1715 (+90%). OOS confirmation on the 4 alt dry lanes (hypothesis formed on BTC only): 1756→942 fires, 0.0704→0.138/share, total 123.7→130.0 HIGHER (dropped fires −0.0078/share gross). Action = restore the validated rule in emit (fidelity fix, NOT a new regime — no re-gate needed); halves FAK exposure/fee drag/spend-cap burn. Also: EMIT_MAX_ASK unset on ALL 10 lanes — live UP lane fired at asks up to 0.979, 43% of last-5d fires above the backtested 0.85 cap (−EV pocket ask>0.90 per a27b852); immediate deploy config fix. Calibration audit (frozen 20-bin, folded, n=920): healthy overall except (a) τ60–90s under-confident (pred .866→realized .936, z≈2.9; 12% of fires — 74% fire at first τ<150s crossing), (b) DOWN .855-bucket under-confident (z≈+2.5) / DOWN ≥.965 possibly over-confident (realized .886, n=44), (c) lowest-σ quartile over-confident (−3.8pt — EWMA λ=0.94 σ over 10s-decimated Chainlink prints lags; Binance-tick σ would be richer), (d) extreme disagreement (pred−ask>0.15) fades −2/−5pt → haircut sizing, don’t skip. Refit-era items (τ-conditional/per-side/σ-conditional) all require refit + walk-forward re-gate. Cross-asset transfer: BTC’s frozen calibration on eth/sol/xrp/doge (n≈430–445 each) within ~2 SE everywhere — the oracle-lag edge generalizes. — crypto-shortterm-algo-accuracy-audit-2026-07-02
  • FAK no-match root cause QUANTIFIED (babylon #942, answers PM #917). Joined emitted_signals × pm_book_snaps (48h, all 10 lanes, ~2,200 fires): ask essentially never empty (1/~2200) — no-match = ask repriced ABOVE our directive limit within ~0.7s (UP 19–25%, DOWN 29–38% of fires; avg move +1.4–1.9c UP vs +2.4–3.1c DOWN → explains PM’s DOWN-skewed counts). Binding worst-price = OUR limit_price = ask_at_emit + EMIT_PRICE_BUFFER (0.02) — why PM’s slippage_cap_bps 200→500 widening did nothing. WR conditional on repricing ≈ unconditional (87–88%) → recommended PM spec: re-priced FAK retry (+5–7c/share on recovered fills) with absolute ceiling ~0.88 + valid_until respect; same-price 1s retry mostly won’t fill (~2/3 still above old limit at +1–3s). Latency stack: emit fire-eval 1s-quantized (due_obs gate) + PM ~100ms + Polymarket mandated ~250ms taker delay — sub-second fire evaluation is the producer lever. — crypto-shortterm-fak-nomatch-root-cause-2026-07-02
  • Exit/loss-cut study: definitive NO-GO — confirmed by two independent runs. This session’s run (BTC-only era Jun 8–Jul 1, ~1.42M rows): Phase-0 PASSED (held=2530, losers=316, adverse_days=19, density 0.9998 — NOT underpowered); Phase-1B best rule −0.77¢/bet, z=−2.33, couldn’t-fill 63%, NO-GO at all ceilings — matching the parallel session’s run (to Jul 3, 1.51M rows: −0.79¢, z=−2.68, couldn’t-fill 52%). Killers: exit fills fail ~half the time exactly when needed (collapsing bid) + weak loser-lead (θ=0.10: precision .86/recall .13); NOT winner-forfeit. Policy: hold-to-settlement stands; do NOT build exit machinery. Repo report docs/superpowers/exit-study-report.md. Appended the replication to the study note (also indexed it in TOPICS/index — it was created 07-01 by the parallel session but never logged). — crypto-shortterm-exit-loss-cut-study-2026-07-01
  • Multi-asset rollout LIVE end-to-end + babylon coordination. 10 containers (BTC live + eth/sol/xrp/doge × Up/Down dry), pipeline verified by dry order rows. Source naming settled for deploy/PM (#943, resolved #885+#887): literals {asset}_5m_algo_{up,down}, BTC legacy crypto_5m_algo_*. Answered PM #917 with the FAK data (#942); resolved stale URGENT #832 as superseded (#944) — silent-stall (781f743) superseded by f00edad, unreproduced; hypothesis = half-open Binance WS starves the engine silently → TODO idle-watchdog + TODO startup-announce rate-limit (crash-loop telegram flood). Gotchas: EngineState uncertainty defaults are absolute dollars (s_sigma 50.0 / k_sigma 1.0 — wrong scale for alts, affects only the p_lo/p_hi band); babylon MCP now natively loaded (mcp__babylon__*, curl-HTTP = fallback). Index status callout updated (UP-only framing superseded). — crypto-shortterm-multi-asset-status-babylon-2026-07-02, crypto-shortterm
  • Durable data-access patterns appended to the data-sources note. (1) homebrew duckdb reads the live Supabase DB without psql: parse SUPABASE_DB_URL into PG* env vars + PGSSLMODE=require, then INSTALL postgres; LOAD postgres; ATTACH '' AS pg (TYPE postgres, READ_ONLY); SELECT * FROM postgres_query('pg', $q$…$q$) — server-side execution, read-only. (2) emitted_signals.side literal changed eras: 'BUY' (Jun 25–28 v1) then 'Up'/'Down'side='Up' comparisons silently mis-score the legacy era (produced a fake 15% win-rate scare this session); always filter side IN ('Up','Down'). — crypto-shortterm-data-sources

2026-06-28

  • Post-review MONEY-PATH FIX PASS on feat/down-side (the branch implementing the both-sides plan; adds a DOWN side to the LIVE UP-only producer). Repo wowjeeez/pmv2-crypto-algo @ /coding/pmv2/crypto_algo; 3 commits (0d5324f ingest warn-logging, 1aef0ef probe_window, 90808d0 observe core); gates green (build clean, 84 tests / 3 ignored live-smoke, clippy -D warnings clean, fmt clean); NOT pushed, NOT deployed, no live-DB writes. C1 — persist-before-publish invariant (headline): emit now runs the emitted_signals INSERT (ON CONFLICT (window_id) DO NOTHING) BEFORE the NATS publish — the DB row is the dedup authority; decision is a pure unit-tested helper fire_decision(Option<u64>) -> {Publish|SkipDuplicate|SkipError} (None=INSERT err → log+no-publish+retry; Some(0)=already claimed → no publish; Some(1)=claim owned → publish). A publish failure AFTER a successful claim is logged loudly and NEVER retried — the written row is a detectable MISSED trade. Invariant: never a duplicate real-money trade, worst case a detectable miss. In-memory fired_windows/total_fires advance right after the claim, before publish (conservative spend). This inverts/supersedes the old publish-then-mark rule (bc38734, crypto-shortterm-emit-guardrails). BookView refactor: Observation now carries up: BookView, down: BookView (8 fields each) not 16 flat pm_*/down_* fields; book_cols() returns a BookView (killed the positional 8-tuple BookCols transposition hazard); emit picks per-side params via pure fire_params(side, obs, market) — closes the “Down-reads-UP-book” copy-paste bug class (test asserts Down⇒down book/token_down/down_outcome_index/SOURCE_DOWN/“Down”). recorder::Row::Observation stays FLAT (BookView mapped to flat fields in observe/src/main.rs — do NOT push BookView across the recorder/SQL boundary). FireSide deleted → select_side(p_cal) -> Option<Side> (two arms: ≥0.80 Up, ≤0.20 Down, else None); regime gate FROZEN τ∈[60,150)s ∧ |p_cal−0.5|≥0.30. C2 startup fail-fast: probe_schema_column(pool,table,col) aborts boot with “apply migration 0006” if a column is missing (emit→emitted_signals.source, observe→observations.pm_down_ask) instead of failing every INSERT in the loop. DRY/test-integrity: is_regime()/max_price() now shared between emit.rs and the lib tests (a test can’t pass while the real gate regresses); deleted stale UP-only tautology tests + hand-copied is_candidate_test/compute_max_price. Down-book visibility: ingest poll_book’s three error arms now warn! (were debug!, suppressed at prod info) with side+token. probe_window now treats source=crypto_5m_algo_down rows as expected (they store p_cal=P(up) next to the Down ask; p_cal meaning unchanged). Tooling gotcha: the rtk cargo hook mangles cargo clippy --all-targets -- -D warnings (passes -D to rustc as a filename → “multiple input filenames”); run clippy + fmt --check via rtk proxy. Created crypto-shortterm-down-side-money-path-fixes-2026-06-28. — crypto-shortterm-down-side-money-path-fixes-2026-06-28, crypto-shortterm-crypto-5m-algo-both-sides-2026-06-27, crypto-shortterm-emit-guardrails

2026-06-27

  • DECISION + PLAN (code not yet built): graduate from single UP-only latency-test source to a real BOTH-SIDES strategy. Two coupled, pmv2-sources-scoped changes: (1) enable the DOWN side (EMIT_BOTH_SIDES=true) — v1 only polled the Up book; the validated regime |p_cal-0.5|>=0.30 is symmetric, so DOWN = p_cal<=0.20 → buy the Down token at the Down-token book ask (the crux: the real Down order-book ask, NOT the synthetic complement 1-Up_ask); (2) split the source crypto_shortterm_latency_testcrypto_5m_algo_up + crypto_5m_algo_down so @positionmanager can dry/live-arm and daily-cap each direction independently (run UP live while DOWN stays dry). Scope = sources/subjects ONLY — Supabase schema stays crypto_shortterm, services stay shortterm-crypto-algo-*, repo stays pmv2-crypto-algo; full rebrand explicitly OUT. Why it’s safe / reused unchanged: calibration reused as-is (20-bin map already covers the Down-conviction bins, thousands of samples — no re-fit, no gate re-run); subjects derive 100% from order.source (pmv2.order.<source>.entry, no subject literals in our code); single-fire-per-window preserved (a window fires Up XOR Down; emitted_signals keeps window_id + gains a source column); PM is outcome-agnostic (buys whatever token_id we send → only data change is two new pmv2_autotrade_sources rows, zero PM code). Structural work: thread a side through feed→engine→Observation→emit, poll BOTH token books, record the down book — migration 0006 adds 6 down-book observation cols + emitted_signals.source. Plan at docs/superpowers/plans/2026-06-27-crypto-5m-algo-both-sides.md in pmv2-crypto-algo (commit d61983b): 8 TDD code tasks + coordination + staged dry→live (Approach A: live-validated). Coordination (babylon, 2026-06-27): @deploy DM #819 — NKey permits for the two new subjects (parallel-OK) → 0006 + emit redeploy at the code pin → retire old permit; @positionmanager DM #820 — two dry source rows (crypto_5m_algo_up/down, $100 cap each proposed), confirm pmv2.order.> wildcard captures them, confirm outcome-agnostic Down, retire old source; pmv2 channel awareness post #821. Status: PLAN + coordination DONE, code NOT yet implemented; dry→live stays Andras’s explicit per-side switch. Created crypto-shortterm-crypto-5m-algo-both-sides-2026-06-27. — crypto-shortterm-crypto-5m-algo-both-sides-2026-06-27, crypto-shortterm-handover-2026-06-26, crypto-shortterm-phase1-bet-sizing, crypto-shortterm

2026-06-26

  • Loss-filter analysis (PRELIMINARY, n=4) — strategy healthy, one filter worth shadow-testing. Forensic decomposition of the first batch of live UP-token fills: 32 trades on wallet 0xAdAc…aEb2, joined to DB (outcomes/emitted_signals/per-second observations) → 31 BUY entries: 27 wins / 4 losses = 87.1% (12.9% loss < ~15% priced → no obvious leak). The 4 losses are 3 distinct phenomena, not one: 2 rollover-from-peak (…487500 spiked −61→+160 then bought mid-collapse at mom30 −86; …490500 peaked +96, bought a dead-cat bounce) = Andras’s “climbed too fast then fell back” hypothesis CONFIRMED + the only cleanly filterable ones; 1 latency/thin (…466500: emit in-regime τ≈62 + ~3.3s fill latency → fill at τ=59); 1 irreducible coinflip (…473700: settled $2 below strike, three wins had thinner cushions). One filter survives EV scrutiny: shat_mom_30 ≥ 0 — skips 2/4 losses, costs 2/27 wins; the case is EV not win-rate (mom30<0 bucket is 2W/2L; momentum doesn’t separate W/L, but losses were expensive fills 0.75/0.76 vs sacrificed wins cheap 0.90/0.93 → net ≈ +1.34 stakes). It’s a physical-state filter (corr w/ pm_ask 0.17) so it does NOT touch the validated disagreement edge. Bonus lever: raise emit τ-floor to ~τ65 for ~3.6s latency headroom (but second-granular + Polygon block time → can’t measure the 128ms REST budget; needs CLOB match timestamps). Rejected (confirmed): pm_ask≥0.80/conv_fair_ask agreement gates delete the disagreement edge (+44 best bet at ask 0.56 dies — already correctly rejected); also overshoot-ratio/sigma/depth/dist≥60 (overfit on n=2). Recommendation: default = do nothing; if acting, shadow-log only (no live bet-dropping); promote a live filter only after ≥15–20 losses with the bucket persistently EV-negative; re-run at n≥20. Method: trades via data-api.polymarket.com/trades, joined window_id=floor(ts/300)*300, entry-available vs hindsight features, 10-agent Claude workflow (5 lenses + 4 per-loss forensics + adversarial overfit synthesis). Created crypto-shortterm-loss-filter-analysis-2026-06-26. — crypto-shortterm-loss-filter-analysis-2026-06-26
  • Phase-1 sizing decision REVERSED — Andras changed his mind; no bet increase. Reverted the (now-superseded) double-bet decision below: the EMIT_SIZE_MULTIPLIER knob was committed (4587df1) then fully reverted (9afe055), directive.rs restored to the venue-floor size_shares. Net effect = none: the knob was a no-op at its default and the activating env/cap DMs to @deploy/@positionmanager never sent (Bash classifier was down at the time), so production never changed — bet stays at the venue floor (~5sh), cap stays $100. The two durable findings survive (size_shares is authoritative via pricing.rs:217; the cap is a reject-GATE not a resizer via reserve.rs:146), kept in crypto-shortterm-phase1-bet-sizing (rewritten as a size-mechanics reference). — crypto-shortterm-phase1-bet-sizing
  • [SUPERSEDED — see entry above] Phase-1 sizing decision — double the bet 5→10 shares + raise the daily cap 200 (Andras). For the live latency test: doubling under a fixed cap would halve fills (~20–30 → ~10–15), so the cap bump is paired to keep ~20–30 latency samples. Per-trade size and the daily cap are independent levers. Still dry until Andras’s explicit live flip (global+source→live; PM min(global,source)) — this only sets behavior for when armed. New env knob EMIT_SIZE_MULTIPLIER (crypto_algo/crates/observe/src/directive.rs, build_entry): reads f64, multiplies the venue-floor share count via pure helper compute_size_shares(min_size, multiplier); default/unset/non-finite/<1.0/unparseable → 1.0 = byte-identical to prior; deploy sets =2. Floor-safety: env .filter(is_finite && >=1.0) + helper .max(1.0) jointly guarantee a bad env can never emit below orderMinSize. Joins the EMIT_* family (EMIT_MAX_FIRES, EMIT_BOTH_SIDES). Key finding (resolves an old ambiguity): @positionmanager honors the producer’s EntryOrder.size_shares directlyposition_manager/crates/position_manager/src/pricing.rs:217 does size_usd = limit_price × size_shares, NO config override — superseding the 62-Directive-era note that PM sized from its own base-size config and treated our size as just a venue floor; so producer-side size_shares IS the authoritative per-trade lever (units = outcome-token shares, not USD). The cap is a GATE not a resizer: pmv2_autotrade_sources.daily_spend_cap_usd (PM-side DB, @positionmanager-owned) rejects over-cap orders (reserve.rs:146ReserveOutcome::CapSrcDaily), it does not shrink them — hence the paired bump. Status: code built + gated green (build/test/clippy/fmt), NOT yet committed (held for Andras’s commit/push go; @deploy redeploys from a pushed main pin). Rollout: (1) commit+push to main, (2) @deploy sets EMIT_SIZE_MULTIPLIER=2 in CRYPTO_SHORTTERM_ENV + redeploys emit, (3) @positionmanager bumps daily_spend_cap_usd 100→200 for crypto_shortterm_latency_test. No effect until the live arm. Created crypto-shortterm-phase1-bet-sizing; corrected the superseded per-fill-size claim in crypto-shortterm-phase1-execution-decision (×2) + cap now a gate / $200; forward pointer added to crypto-shortterm-handover-2026-06-26 item 1. — crypto-shortterm-phase1-bet-sizing, crypto-shortterm-phase1-execution-decision, crypto-shortterm-handover-2026-06-26
  • Handover snapshot captured — producer LIVE, dry-validate PASSED end-to-end, latency STILL untested. Mirrored the repo HANDOVER.md into the vault. Major deltas since last sync: (1) emit is LIVE on the Zurich VM and dry-validate passed end-to-end — emit fires → publishes a pmv2-contracts schema_2 EntryOrder to NATS pmv2.order.crypto_shortterm_latency_test.entry@positionmanager consumes on stream PMV2_ORDERS → eff=Dry → DryRehearsed (build+sign, NO POST, no money) → telegram notify; EMIT_MAX_FIRES raised 50→1000 (PM daily_spend_cap_usd=$100 is the real live bound). (2) VM moved Frankfurt→Zurich to clear a Polymarket order GEOLOCK — critical correction: the autotrade “didn’t work” because of the geolock, NOT latency, so the ~128ms latency-capturability is STILL UNTESTED; Andras’s manual Zurich trades = net profit (validates the signal, small sample). (3) v1 = UP-only by design (EMIT_BOTH_SIDES=false): emit only polls the Up-token book; Down would need the No-book/synthetic complement (deliberately avoided for a clean latency test) → leaves ~half the symmetric edge on the table (deferred upgrade). (4) Architecture now v2 / Pattern B: producer-only, repo wowjeeez/pmv2-crypto-algo, code /coding/pmv2/crypto_algo, GAR apps/pmv2-crypto-algo, source_id crypto_shortterm_latency_test, subject pmv2.order.crypto_shortterm_latency_test.entry, executor @positionmanager (not @polymarket); main 36592cd+; gate CI needs PMV2_REPO_MASTER_KEY for the private pmv2-contracts dep (rev 1f05fff). Durable gotchas: u256 token ids (never u128); Gamma orderPriceMinTickSize/orderMinSize; EMIT_MAX_FIRES hydrates from emitted_signals on boot; PM mode = min(global,source) ignores enabled (live = both flips); babylon MCP flaky → HTTP fallback. Open items: latency capturability (now testable, Andras’s $100 switch, ~20–30 fills), Down side, reconcile emitted_signals (58+ fires) vs outcomes vs manual fills. Refreshed the index Status/mermaid/Next-step/Related. Created crypto-shortterm-handover-2026-06-26. — crypto-shortterm-handover-2026-06-26, crypto-shortterm

2026-06-25

  • apps-shortterm-crypto-algo-emit Pattern A → B residue: EMIT_MAX_FIRES cap cleared (~19:25Z). Container had been spinning EMIT_MAX_FIRES reached, signal emission paused max_fires=50 ~22h at 1Hz with zero JS publish attempts. Real cause: 50 Pattern-A rows in crypto_shortterm.emitted_signals (never truncated on cutover) + compile-time default EMIT_MAX_FIRES=50 hydrated the BLOCKER-3 in-process counter to 50 at boot → paused → short-circuit before publish. Critical insight: TRUNCATE alone does NOT clear the cap — the BLOCKER-3 hydration is one-shot at boot. Must sudo systemctl restart apps-shortterm-crypto-algo-emit.service; boot log INFO emit: guardrail hydrated from DB hydrated=0 confirms reset. Producer (cryptoshort) first hypothesized JetStream subject-capture failure (DM #701); falsified via in-nats-container wget /jsz?streams=true&config=true showing PMV2_ORDERS subjects:["pmv2.order.>"] wildcard DOES capture. cohort_publisher NKey permits verified (PUB pmv2.order.crypto_shortterm_latency_test.> + $JS.API.> + $JS.ACK.>, SUB _INBOX.>). Approval chain: producer DM #705 (“GO with option A — truncate”) → producer-DB mutation OUTSIDE Andras #480 standing-auth carve-out → surfaced + operator-approved before TRUNCATE. Operational corollary added to crypto-shortterm-emit-guardrails BLOCKER-3 section; gotcha #38 in levandor-infra captures the general “in-process DB-hydrated counters re-read only on boot” rule — crypto-emit-max-fires-cap-clear-2026-06-25, crypto-shortterm-emit-guardrails

2026-06-22

  • Phase-1 execution decision documented — Directive producer + staged 100 test: dry-validate (free) → tiny 5-fill ~100(backstop). emit deployed for dry-validate (#264, both services active, sourcemode=dry_run`, migrations 0003/0004/0005 applied). Created crypto-shortterm-phase1-execution-decision. — crypto-shortterm-phase1-execution-decision, crypto-shortterm-phase1-latency-budget
  • Repo rename + relocation — pmv2 dir-per-producer layout. GitHub repo wowjeeez/crypto-fetchwowjeeez/pmv2-crypto-algo (GitHub redirects old→new; local remote updated + verified, origin/main=7bfe3a0); code /Users/levander/coding/crypto_fetch/Users/levander/coding/pmv2/crypto_algo (git history + .env/.mcp.json (babylon handle crypto) moved, builds clean). pmv2 re-arch = one self-contained repo per component: crypto_algo (this, a producer), cohort_algo + us_weather_algo/as_weather_algo (producers), position_manager (shared executor), telegram_connector (notifications). Updated the index + every note citing the old name/path. Created crypto-shortterm-repo-rename-pmv2-rearch. — crypto-shortterm-repo-rename-pmv2-rearch
  • Root-cause fix: Polymarket token ID u128 overflow silently zeroed the emit market-cache (commit 7bfe3a0). try_build_market validated token IDs via token.parse::<u128>(); Polymarket tokens are 256-bit (77-digit) → overflow → parse failure → Ok(None) on every poll → emit fired 0× in ~2 days of dry-validate. Found via a live Gamma probe (no VM logs). Fix: digit-string check !s.is_empty() && s.bytes().all(|b| b.is_ascii_digit()); also corrected two GammaMarket serde renames (minSizeorderMinSize, tickSizeorderPriceMinTickSize — real tick 0.01, was defaulting 0.001 = 10× off); condition_id-absent hypothesis REFUTED (Gamma supplies a real 0x); 3 regression tests added. Separately: observe went ~5h+ stale (latest window lagging), a distinct feeds issue likely the shared Alchemy creds drain (@polymarket #269) — must also clear or emit has no live windows. Pending @deploy. — crypto-shortterm-polymarket-token-overflow-fix
  • DEDUPE: merged the two overlapping guardrail notes into one canonical crypto-shortterm-emit-guardrails. Folded the unique Files-Changed table, Test-Results, and DB-vs-NATS-failure-semantics callout from crypto-shortterm-emit-safety-guardrails.md into the canonical note, then deleted the redundant note and repointed all wikilinks (phase1-emit ×2, prelive-pass2, TOPICS). — crypto-shortterm-emit-guardrails
  • Index brought current to Phase 1. crypto-shortterm.md For-Agents block, canonical-artifacts, Status (now “Phase 1 — live latency test in progress”; Phase-0 offline gate PASSED; staged $100 test; @deploy + @polymarket #285 blockers; dry-validate 0-fire bug fixed), mermaid “we are here” note (Phase0→Phase1), Next-step, and Related all updated. — crypto-shortterm

2026-06-20

  • Phase-1 latency budget analysis (latency_budget.rs, commit 2f5fb4d): publish→fill target = 128ms (p90 adverse drift), sensitivity HIGH. 143,228 consecutive obs pairs in the candidate regime (τ∈[60,150)s, c≥0.30); 29.2% of 1s steps already exceed 0.5¢ ask move; adverse drift at p90 = 3.91 ¢/s → 128ms budget; p95 = 5.84 ¢/s → 86ms; targeting <50ms strongly recommended. Caveat: 1s poll cadence means intra-candle drift is unobserved — true budget may be tighter. — crypto-shortterm-phase1-latency-budget

2026-06-19

  • Pre-live Review Pass 2 — Run+Integrate Fixes (commit 89e354e): 1 BLOCKER + 2 HIGH + 3 MEDIUM + 2 LOW, all fixed; 7 new tests; fmt/build/clippy -D warnings clean. BLOCKER: migrate binary replaced runtime fs::read_to_string path-search loop with include_str! const (MIGRATIONS: &[(&str, &str)]) for all five migrations — no CWD dependency, safe for systemd. HIGH: NATS NKey auth wired (async_nats::ConnectOptions::with_nkey(seed).connect(url); hard error at startup if URL set but seed absent); Binance staleness gate clock-source fixed (pp.as_of = Utc::now() after parse_trade, local-clock-to-local-clock compare; also .unwrap() on timestamp_millis_opt.single().unwrap_or_else(Utc::now)). MEDIUM: discover_active_market() removed from hot event loop — background 30s poll via tokio::sync::watch, fire path reads market_rx.borrow().clone() (zero HTTP latency); outcome labels confirmed ["Up", "Down"], .trim() + .to_lowercase() + loud warn! on no-match + token_down u128 validation; User-Agent + 5s timeout on poll_book reqwest client. LOW: EnvFilter::new(filter) with "info" fallback (no set_var); ON CONFLICT (window_id) DO NOTHING confirmed correct. Live label confirmed: btc-updown-5m-1781901900["Up", "Down"]. — crypto-shortterm-emit-prelive-pass2
  • Pre-live money-safety code review + hardening of the emit binary (commit bc38734, review on schema-align pin c12c830) — 3 BLOCKERS + 3 HIGH + several medium/low, all fixed; 21 tests green, gates clean. BLOCKER fixes: condition_id must start with “0x” else skip the fire (return None) — a slug is not a valid 0x id so pmv2 dedup/settlement would break (wrong-instrument guard); YES/UP token selected by matching the outcomes label not array index clobTokenIds[0] + validate id parses as u128 (a Gamma reorder would silently BUY the wrong outcome); single-fire set + EMIT_MAX_FIRES cap hydrated from emitted_signals on startup (were in-memory only → crash/deploy mid-run could re-fire or exceed the cap). HIGH: stale/zero/NaN ask drove bad max_price (to_f64→0.0) → PM book staleness gate (5s) + binance staleness gate (30s) + finite/>0 ask guard; NATS publish failure now continues (only mark fired after publish AND db-insert succeed, else retry next tick, idempotent via Nats-Msg-Id) — previously a failed publish still marked fired → signal silently lost. MEDIUM: valid_until_ts lead guard (≥now+10s), discover HTTP timeout 5s, calibration p clamped + NaN check, market-race assert (book window_id==directive window_id). LOW: calibration baked in via include_str!, DB INSERT duplicate binding fixed + emitted_signals schema reconciled (migrations 0004+0005). Pattern: build→review→fix is the standing pre-live discipline for any money-path code; a final delta review is planned right before the live flip. Remaining gates to a live latency test: Andras merges pmv2 #62 + source migration → deploy polymarket-fetch → deploy crypto-shortterm @ bc38734 + add emit service → shadow-validate → Andras flips live with tiny caps. — crypto-shortterm-phase1-emit, crypto-shortterm-emit-guardrails
  • Directive struct aligned to pmv2 confirmed schema (commit c12c830). Removed provisional fields (token_id, side, size, emit_ts); added id ("clt-{window_id}", idempotent retry key), source ("crypto_shortterm_latency_test"), schema_version=1, action="BUY", condition_id (0x hex), yes_token_id (String not i64 — avoids overflow), tick_size, min_size, neg_risk, max_price (tick-rounded/clamped formula), valid_until_ts (default = win_open+300, EMIT_VALID_SECS env override = latency signal), bucket_label (optional). NATS publish now sets Nats-Msg-Id = directive.id (pmv2 dedup key). UP-only v1: fires on p_cal>0.5 only; EMIT_BOTH_SIDES=false const gates DOWN path. polymarket.rs now extracts condition_id, neg_risk, tick_size, min_size from Gamma API (fallbacks: DEFAULT_TICK_SIZE=0.001, DEFAULT_MIN_SIZE=5.0). Migration 0004 adds id+yes_token_id columns, makes size nullable. — crypto-shortterm-phase1-emit
  • Phase-1 live signal emitter built and committed (SHA 5e92c7f). DRY refactor: EngineState, Observation, floor_window, make_window, to_f64, constants extracted from observe/main.rs into observe/src/lib.rs; both observe and emit binaries reuse the same engine. New files: lib.rs, calibration.rs (CalibrationMap), directive.rs (Directive + to_json), bin/emit.rs, migrations/0003_emitted_signals.sql, cal_export.rs. Regime: τ∈[60,150)s AND |p_cal−0.5|≥0.30 (winning cell from selectivity). Shadow mode (no NATS creds needed): logs directives when NATS_URL absent, publishes to "cohort.signals" when set (async-nats = "0.37"). Single-fire-per-window HashSet + EMIT_MAX_FIRES cap (default 50). Records every fire to crypto_shortterm.emitted_signals. PROVISIONAL TODO: Directive.condition_id = market.slug — must swap to confirmed Polymarket CLOB condition ID in directive.rs before go-live. calibration.json not yet generated — must run cal_export against DB first. 8 new unit tests; 0 failures; build/clippy/fmt clean. — crypto-shortterm-phase1-emit
  • Audit of depthaware.rs window-count discrepancy (commit a3e0468): 524 (header) vs 770 (decay table) — NO BUG, different populations. 524 = raw p_fair candidates over ALL folds incl. train-only fold 0; 770 = calibrated p_cal candidates over OOS folds 1–5 only. Test fold slices are non-overlapping; N=770 in t-stat denominator is legitimate; t-stats unchanged; SURVIVES verdict stands. ~47% increase is a calibration effect. Fixes: split n_candidate_windows into n_raw_candidate_windows/n_oos_candidate_windows, replaced hardcoded "523" in build_verdict with live n_windows, replaced 31_536_000.0 magic number with types::SECS_PER_YEAR (no numeric change). Clippy/fmt/build clean. — crypto-shortterm-phase0-depthaware
  • Phase-0 depth-aware selectivity re-run (bin depthaware.rs, commit bba5739) — DECISIVE OFFLINE GATE: SURVIVES. 524 candidate-depth windows (τ∈[60,150)s, c≥0.30), Bonferroni bar t≥2.99. Edge clears the bar at all sizes up to 2000 shares: S=50 → t=5.52/5.01¢/99.8%; S=2000 → t=3.01/2.61¢/86.8%; ∞ uncapped → t=5.67. Decay shape SANE/DECLINING (¢/share monotonically falls as S grows) — not a top-of-book artifact. Capacity ceiling 2000 shares. Prior t=4.61 result (crypto-shortterm-phase0-selectivity) was NOT a fill-realism mirage. Latency buffer +0.5¢ sensitivity: verdict unchanged. T19 status: SURVIVES offline depth-cap test; remaining gate = funded real-fill test to bound latency cost. Report docs/superpowers/phase0-depthaware-report.md. — crypto-shortterm-phase0-depthaware

2026-06-18

  • Phase-0 DEPTH PROFILE of the new book-depth columns (added in deploy pin 69b84b1; bin crates/analysis/src/bin/depthprofile.rs, commit 35aa202) — the #1 risk (thin-book/0.98-spread/unfillable mirage) is WEAKENED, T19 stays CONDITIONAL. Run over 3 days of depth-populated data: 159,113 rows / 761 windows; candidate regime (τ∈[60,150)s, |p_fair−0.5|≥0.30 = the crypto-shortterm-phase0-selectivity winning cell) = 12,197 rows / 360 windows. HEADLINE: in the candidate regime the take-side TOUCH median = 658 shares (5,109); take-side 3¢ cumulative depth median = 6,171 (~1 each). So the books are NOT empty where the edge lives → the ~7%/trade backtest edge is not obviously unfillable due to empty books; the long-standing 0.98-spread mirage hypothesis (from crypto-shortterm-phase0-codebase-complete, was 2 windows) is substantially weakened, now superseded by 360 windows of recorded depth. STILL NOT A GO — three skeptic caveats: (1) capacity ceiling ~$432/trade (6,171×0.07) is CIRCULAR — multiplies real depth by the still-UNVERIFIED 7¢/share edge; depth is now an independent fact, the edge is not; (2) depth = size at the 1s snapshot, says nothing about survival ~150ms after signal (latency unverified, needs real fills); (3) sanity watch — depth-aware re-run edge should DECAY sensibly as size grows; a flat curve smells like an artifact. Next gate: depth-aware selectivity re-run (walk-forward, cap each fill to recorded best_*_sz, plot edge-vs-size decay, corrected significance bar). 360 candidate windows now vs 834 in the original power calc → preliminary pass possible now, decisive run ~1 week at ~800 windows. Only if the edge survives realistic sizing → small real-fill test on a funded account in τ60–150s/c≥0.30; snapshot paper-lane still rejected. Report docs/superpowers/phase0-depthprofile-report.md. — crypto-shortterm-phase0-depthprofile

2026-06-12

  • Phase-0 selectivity/regime sweep (bin selectivity.rs, commit 2d52628) on the same K=6 purged/embargoed walk-forward — T19 = CONDITIONAL pass. 7 a-priori economically-motivated cells (conviction |p_cal−0.5|∈{0.10,0.20,0.30} × τ-band∈{[60,150)s,[150,300)s} + baseline), Bonferroni bars t≥2.69 (K=7) and t≥2.97 (broad N≈17). WINNER τ∈[60,150)s c≥0.30: t=4.61, 5/5 OOS folds positive, ~34k trades, 834 windows, mean 2.21/window, Sharpe 0.16 — PASSES both. Also pass c≥0.20 (4.10), c≥0.10 (4.04). FAIL: baseline all-τ≥60s (2.51) and ALL τ∈[150,300)s (<1.8). Edge lives exclusively in the last 60–150s on clearly-ITM/OTM contracts. Selectivity raised t 1.46→4.61 by variance reduction — mean PnL actually FELL (4.15→2.21/win), so SIGN/persistence is robust OOS. CAVEAT: magnitude is a mirage — ~5.5¢/~80¢ ≈ ~7%/trade is implausible for real binary fills ⇒ 1s top-of-book ≠ fillable price (thin-book/snapshot optimism, the recurring #1 risk); unresolvable offline (observations lack book DEPTH; every PnL assumes infinite top-of-book fill). Next: record book sizes → re-run depth-aware → only then a tiny real-fill test in τ60–150s/c≥0.30; snapshot paper-lane stays rejected. Report docs/superpowers/phase0-selectivity-report.md. — crypto-shortterm-phase0-selectivity
  • Phase-0 purged walk-forward (K=6, 750 OOS windows, τ≥60s, 1% slip): t-stat 1.46 < 3.0 — EDGE FRAGILE / LIKELY OVERFIT; 5/5 folds positive (consistent direction) but N insufficient to clear the multiple-testing bar; T19 OPEN; do NOT build live machinery. (commit 476cab1) — crypto-shortterm-phase0-walkforward
  • CORRECTION (verdict overstated, now fixed): T19 is NOT passed/resolved — it remains OPEN, status: offline edge confirmed (promising, de-risked), live-capturability UNCONFIRMED. validate.rs (commit 4b3176d, 198,733 rows / 243 test windows) DID de-risk the leakage concern: ≥60s all-trades baseline = −3.21%, so the earlier +2.08% baseline WAS 0–60s near-expiry leakage (concern resolved). ≥60s gated strategy nets +1.61%/trade at 0 slip, +0.81% at 1% slip on ~44k trades; ~54% of PnL is outside the leaky zone — a real OFFLINE forecast edge survives the debunk. But this is still an offline backtest with optimistic fills (top-of-book, no depth/capacity — sizes unrecorded; book ~1s stale vs P_fair), which for a latency-style edge structurally overstates live-grabbable PnL; the edge is slippage-sensitive (halves at 1% cost); and it’s a single train/test split, not the spec’s purged/embargoed walk-forward + deflated Sharpe + t≈3. Remaining gates before any GO: full purged walk-forward + deflated Sharpe/t≈3; capacity/depth read (needs recording book sizes); and decisively a forward paper-trade (real-time, real fills, no money). Updated crypto-shortterm-phase0-validation (verdict callout), crypto-shortterm-open-questions, TOPICS — verdict changed from “PASSED/resolved” to “OPEN / validation-in-progress”.
  • T19 EXIT GATE PASSED — validate.rs confirms GENUINE EDGE — verdict OVERSTATED, corrected by the entry above; T19 is OPEN. (commit 4b3176d, 198,733 rows / 243 test windows). ≥60s gated strategy: +1.61%/trade at 0 slip, +0.81% at 1% slip (n=43,961). ≥60s forced baseline (raw p, no gate): −3.21%. 0–60s zone = 45.6% of total PnL, confirming the original +2.08% all-trades baseline was leakage-contaminated. Created crypto-shortterm-phase0-validation.

2026-06-11

  • CORRECTION (overturns the entry below): T19 is NOT passed — status OPEN / validation-in-progress. The prior write-up’s verdict (“PASSED”, +7.28%/+19.46% confirmed edge) was wrong/overconfident. Numbers are factually accurate (calibrated test +7.28% overall / +19.46% in 0–30s / hit 54.4% / Sharpe 0.17; raw-forced baseline +2.08%; calibration lifts hit-rate 30%→54%), but reframed to PROMISING but NOT confirmed; likely contaminated. 🚩 The +2.08% all-trades baseline should be ≈ −fee (negative) for a no-edge strategy — positive ⇒ systematic bias, most likely near-expiry outcome leakage (as τ→0, P_fair’s σ√τ denom → 0 so P_fair collapses onto the already-determined outcome). Edge concentrates in the 0–30s leakage zone and decays with τ — consistent with leakage (also partly consistent with the genuine fast-feed-vs-lagging-book thesis; not yet separated). Fills are optimistic (top-of-book @1s, no depth/slippage; near-expiry depth thin → binary-near-expiry mirage). ESTABLISHED: directional signal is real (monotonic calibration). NOT established: a tradeable after-cost edge. Open gate before any GO: explain/kill the +2% baseline; re-run excluding last ~60s; realistic-fill model (depth cap + slippage); purged/embargoed walk-forward + deflated Sharpe + t≈3; then forward paper-trade before capital — crypto-shortterm-phase0-edge-analysis
  • Phase-0 edge analysis live results: T19 exit gate PASSED (+7.28% avg net PnL, Sharpe 0.1746; best τ-bucket 0–30s: +19.46%, Sharpe 0.4962) — verdict overturned by the correction above; T19 is OPEN. window_id is TEXT not UUID in live schema — crypto-shortterm-phase0-edge-analysis

2026-06-08

  • MILESTONE: Phase-0 codebase COMPLETE — analysis report-gen runs + first (thin) data (codebase commit 97f238d; status stays Phase 0 but advances “observe pipeline live” → “codebase complete; report-gen runs; first data in; awaiting @deploy continuous run for calibration + a stable basis/spread read before go/no-go”). The full stack — types/fairvalue/market_state/clock/ingest/recorder/state/observe(live)/analysis/migrate — is built, tested, and committed; the analysis report-gen now runs against crypto_shortterm and writes docs/superpowers/phase0-report.md. Source moved to a new private repo wowjeeez/crypto-fetch pinned at a11ce63; the continuous run is coordinated with the @deploy agent (build-on-VM on polymarket-infra, Frankfurt) and awaits the @deploy run. FIRST PRELIMINARY DATA — THIN, 2 windows, 0 outcomes, calibration pending, do NOT over-conclude: (1) **Binance↔Chainlink basis ≈ -63k BTC — confirms the settlement-vs-predictor basis is real and non-trivial (already in the confidence band) and validates the Rev-2 reframe (settlement = Chainlink, Binance = predictor); (2) WATCH / TOP RISK — Polymarket book spread ≈ 0.98 on the sampled windows (bid ~0.01 / ask ~0.99 = effectively empty/untradeable books) — if it holds it challenges the tradeable-edge premise (can’t profit through a 98¢ spread) and echoes the earlier ~150–250 touch-depth finding**; flagged as the **#1 risk for the Phase-0 exit gate** (could be a 2-window sampling artifact — needs many windows). **REMAINING (both data-gated on the continuous run):** **T19 Phase-0 exit/gate decision** + the funded **5 fee-at-match CONFIRM. Next op step: @deploy deploys observe → data accrues across windows → re-run analysis (calibration + stable basis/spread) → go/no-go. Created crypto-shortterm-phase0-codebase-complete; touched crypto-shortterm, crypto-shortterm-strategy-design, crypto-shortterm-backtesting, crypto-shortterm-open-questions, LOG, TOPICS.
  • MILESTONE: Phase-0 observe-only pipeline BUILT and VERIFIED LIVE end-to-end (commit f50b084; status moved “logic built, blocked on access” → “observe pipeline live & verified; accruing data”). The observe binary ran all four feeds and wrote to the isolated crypto_shortterm Supabase schema for ~5 min: binance_ticks 3,805, cl_reports 28, pm_book_snaps 728, observations 1,340 (K=63,596 set at a witnessed window open, p_fair evolving correctly toward expiry), outcomes 0 (needs a full window to close during a sustained run). Live architecture: Binance @trade WS → predictor (Ŝ/basis/σ); Chainlink BTC/USD read on-chain on Polygon via Alchemy RPC (~1s poll) → settlement reference (K at open, settle at close, basis) — NOT the paid Data Streams product, which resolves the prior access blocker; Polymarket 5-min market via Gamma (btc-updown-5m-<300s-aligned-ts>) + CLOB /book ~1s; clock offset via Binance server-time (lenient 5s observe / strict 0.25s Phase-1); async recorder → crypto_shortterm. Three findings: (1) CONFIRMED Polymarket 5-min BTC Up/Down exists (series btc-up-or-down-5m, 300s windows) — an earlier build wrongly pivoted to 15-min via a 900s timestamp-floor bug, corrected to 300s; (2) recorder throughput — raw ticks (~50/s) saturate per-row async inserts even at channel cap 20k → spill; throttled to 1/s for Phase 0, Phase 1 / finer basis/σ needs batched/COPY (overflow may be at /tmp/crypto_shortterm_spill.jsonl); (3) K-correctness — K is only set for windows whose open the process witnessed, so meaningful data needs the loop to run continuously across clean window boundaries. Next (operational): run observe continuously to accrue observations + outcomes across many windows, build/run the analysis report (reliability curve, basis-by-τ, depth-by-τ), make the Phase-0 exit/gate decision, then Phase 1. Created crypto-shortterm-phase0-live-milestone; touched crypto-shortterm, crypto-shortterm-data-sources, crypto-shortterm-strategy-design, crypto-shortterm-open-questions, LOG, TOPICS.
  • Phase 0 reached a tested milestone — observe-only pipeline LOGIC built, unit-tested, committed (status moved planned → Phase 0 in progress, blocked on access). In /Users/levander/coding/crypto_fetch: 10 crates (types, fairvalue, market_state, clock, ingest + Binance/Chainlink/Polymarket parsers, recorder, state, observe, analysis), 27 tests passing, clippy + fmt clean, ~20 commits, migration SQL written. Remaining Phase-0 work is BLOCKED on external access: Supabase DATABASE_URL, Chainlink BTC/USD Data Streams access, a run environment, and (deferrable) a funded account for the fee-at-match confirmation. Remaining run-dependent code: live WS connect loops, recorder Supabase writer, SNTP + observe live wiring, analysis report-gen, exit report. Captured design refinements found during the build: (1) band nuance — the band only widens near-the-money near expiry; when clearly ITM/OTM near expiry ∂P_fair/∂Ŝ→0 so it correctly narrows (the high-confidence tradeable regime), so “size up near expiry” is now stated as “size up near expiry when clearly ITM/OTM; near-the-money-near-expiry is the danger zone” — and the pessimistic-band sizing does this separation with no hand-tuned cutoff; (2) σ warm-up — observe engine falls back to a prior σ=0.6 until the EWMA has its 30 warmup samples; this prior must be validated against real measured vol in Phase 0; (3) minor — two plan test fixtures were mis-specified and corrected (geometric vs arithmetic symmetry for the log-moneyness P_fair; a calibration test using uncalibrated data), logic unaffected. Linked the Phase-0 plan + docs/superpowers/phase0-confirmations.md. Touched crypto-shortterm, crypto-shortterm-strategy-design, crypto-shortterm-risk-management, LOG, TOPICS.
  • Updated the whole project to the approved Rev 2 design + verified venue facts. Core reframe: settlement is the Chainlink BTC/USD DON-median stream, not Binance — the edge is “predict the next Chainlink print from the faster Binance feed and trade Polymarket’s lag to the oracle” (P_fair = Φ(ln(Ŝ/K)/(σ√τ)), K/Ŝ from Chainlink, Binance = predictor, Binance↔Chainlink basis = core risk, largest near expiry; close ≥ open → Up, ties → Up). Captured verified venue facts (2026-06-07): taker fee = C·0.07·p(1−p) peaks ~2.5% at p=0.5 and is tiny near 0/1 → near-expiry primary regime, mid-window p≈0.5 fee-killed; maker 0 + 20% rebate; UMA optimistic oracle ≥2h challenge → capital locks ≥2h unless you sell pre-resolution; touch depth ~18.9M/day competition. Key design corrections: τ→0 blow-up (∂P_fair/∂Ŝ→∞) → band widens near expiry, size on pessimistic edge + τ-floor; net-of-fee portfolio Kelly ≤0.5× across correlated overlapping windows; synchronous WAL + reconciliation = financial source of truth (async recorder = telemetry); adverse-flow veto on takers too; causally-typed core + trial ledger + exec-sim/dry-run/canary. Set status planning → design approved / Phase 0 planned; linked the codebase spec + Phase-0 plan paths. Touched crypto-shortterm, crypto-shortterm-strategy-design, crypto-shortterm-data-sources, crypto-shortterm-risk-management, crypto-shortterm-backtesting, crypto-shortterm-open-questions, LOG, TOPICS.

2026-06-07