How this research governs itself
Self-imposed model-governance discipline in the style of SR 11-7. MCH Advisory Services is an independent research publisher, not a registered investment adviser; nothing on this page is investment advice or a regulatory attestation. Generated 2026-08-25T04:00:08Z.
Verification infrastructure
1 re-planted defects · 0 survivors. real shipped defects re-planted into the code; a mutation no suite catches means the guard is certified by nothing, and the harness fails
all BLOCKING invariants hold. properties of the PUBLISHED book, checked every run: blocking stays green, known debts are counted (never silent), register-held pages are exempt-but-printed
Exception register
2026-08-25858 = 845 fresh + 0 caveat + 13 exceptions — estate = fresh + caveat + exceptions — every withheld name is named with its reason, never silently dropped
| Name | Class | Reason |
|---|---|---|
| ACI | — | — |
| AES | — | — |
| AVB | — | — |
| EA | — | — |
| EQR | — | — |
| LITE | — | — |
| MNST | — | — |
| NSA | — | — |
| SLAB | — | — |
| TECH | — | — |
| TMHC | — | — |
| TXNM | — | — |
| WDC | — | — |
Append-only ledgers
Forecast distributions — per-name forecast distributions, hash-chained monthly segments, captured at publish; interim and matured cohorts never pooled✓ forecast_ledger: 13686 rows across 1 segment(s), 0 chain break(s)
Options pre-registrations — every published options structure pre-registered before outcome; v1/v2 cohorts never pooled; duplicate baseline moves only with the causing row named at the constant
entries 15585 · structures 44274 · distinct trades 42848 · distinct contract positions 19423
duplicates collapsed at settlement : 1426 (baseline 1426)
last pre-registration : 2026-08-25 (0d ago)
growth check: ✓ OK
Attribution — last run 2026-07-31 under AM-029+AM-031: 59 names attributed, 47 within their own noise bands, shrugged share 15.3%. a decomposition that mostly shrugs beyond quantified noise REFUSES to emit; refused runs append nothing
Model-change proposals
propose-onlythe learning engine is PROPOSE-ONLY, structurally: its single write target is this ledger (asserted on the AST, re-planted by the mutation harness); auto-deploy of learned changes was refused permanently
| Proposed | Kind | Cohort | n | Evidence |
|---|---|---|---|---|
| 2026-08-10 | expectation-bias | sector:Energy | 37 | 76% of 37 names err the same way; mean error -0.029 (interim, mark-to-date) |
| 2026-08-10 | expectation-bias | sector:Information Technology | 106 | 86% of 106 names err the same way; mean error +0.062 (interim, mark-to-date) |
| 2026-08-10 | expectation-bias | sector:Materials | 52 | 79% of 52 names err the same way; mean error +0.041 (interim, mark-to-date) |
| 2026-08-10 | kill-switch-telemetry | estate | 33 | 9/33 names with price-observable triggers have at least one FIRED trigger (coverage, not accuracy — accuracy waits for maturation) |
Pre-registration amendments
57appended, never edited; corrections are NEW amendments that name what they correct.
AM-059 · 2026-08-20 — Report presentation overhaul — canonical value layer, editorial QA, progressive disclosure
- owner instruction
- AM-059 approved — given verbatim in-session 2026-08-20, against the sign-off request (Desktop/AM-059_SignOff_Request.html); no additional instruction wording supplied.
- what this authorises
- 0
- canonical_triangulation as THE fair-value derivation for report + JSON + margin_of_safety (retires the json unweighted mean)
- 1
- valuation_confidence {level, basis} — advisory, drives display precision ONLY
- 2
- QA-313/314/315 (WARNING → wave-3 flip) and QA-518..523 (editorial engine; QA-523 blocking)
- 3
- one canonical external rating + subordinate 5-tier line (owner Option A)
- 4
- computed weight adjectives; options quote_status vocabulary; Research Conviction definition + technical_trend caveat (display-only)
- 5
- Research Provenance section; Evidence row (citation_coverage_v2 load-bearing subset)
- 6
- coarse IC headline rows under LOW confidence (~$N (≈ ±M% vs spot); nearest $1/1pp)
- 7
- hub wave 2: 9-part baker grouping, typed decision card, part-9 details, TEMPLATE_VERSION in the content hash
- registered definitions
- valuation confidence
- high: >=4 survivors and dispersion<=0.25 and dcf in blend and sign agrees; medium: >=3 and <=0.55; else low (basis names the failing rule)
- restatement budgets
- [object Object]
- banned phrases n
- 10
- length ceilings
- [object Object]
- float leak pattern
- \d+\.\d{6,}|\b\d+(?:\.\d+)?e-\d+\b
- coarse granularity
- nearest $1 and nearest 1pp — fixed by QA-502 half-unit and QA-511 1.5pp arithmetic
- coarse surfaces
- IC Summary: Triangulated fair value row · IC Summary: 12-mo scenario PWEV row
- quote status default
- eod_indicative (AV chains are previous-close marks)
- flip schedule
- [object Object]
- evidence gate
- form
- AM-035 (frozen-book impact BEFORE the publishing run)
- summary
- [object Object]
- top movers
- [object Object] · [object Object] · [object Object] · [object Object] · [object Object] · [object Object] · [object Object] · [object Object] · [object Object] · [object Object] · [object Object] · [object Object] · [object Object] · [object Object] · [object Object]
- stated expectation before running
- zero rating flips (verified); MoS column moves per table; MAS C_value re-ranks at first nightly; nullity unchanged
- known input defects carried
- 0
- AM-055 SBC off (disclosed)
- 1
- D-45 raw benchmark closes
- 2
- ~63 authored names carry a banned phrase (QA-519 stays WARNING)
- 3
- 169 published files carry pre-A2 float artifacts (clear at re-emit)
- defects raised
- D-57
- harness same-length-mutation bytecode blind spot (fixed)
- D-58
- ev_ebitda truthiness guard dropped peer on ~700 names (fixed)
- never this
- 0
- valuation_confidence never feeds a rating, stance, size, gate or MAS input
- 1
- coarse forms never reach the Quick-Facts table, valuation tables, captions or the PDF
- 2
- no severity flip without the wave-3 frozen-book re-measure; holds need a census declaration first
- 3
- conviction technical_trend weight does not move here (AM-060)
- 4
- MANIFEST_REQUIRED unchanged (own checkpoint)
- 5
- budgets ratchet only via a superseding amendment together with tests/test_editorial_qa.py
- 6
- never a local page re-bake commit — the hub wave ships source-only, AWS bakes
- numbering note
- AM-058 reserved (sim options/overlays per AM-057); conviction weight change pre-assigned AM-060.
AM-057 · 2026-08-19 — MCH Alpha Score (MAS) — a diagnostic composite over published diagnostics, registered before any forward read; plus two Simulation Portfolios that rank on it (Top 10 MAS-1Y / MAS-5Y)
- owner instruction
- Marinus 2026-08-19 (chat): 'AM-057 approved, append it and run the republish' — on the draft at docs/AM-057-DRAFT.md, which states in full what is registered (the block definitions, lambda 0.30, gates, penalties, the harness tests and horizons, the two simulation books at their hashes) and what is NOT (no sizing, no rating, no stance, no optimizer, no sim selection reads MAS; nothing above 'unvalidated' until the record matures). Earlier owner decisions folded in (2026-08-18): ship unvalidated now per spec §8; Block E staleness_multiplier = 1.0 (staleness enters once, as the final penalties); no colour ramp; add the Top-10 MAS-1Y / MAS-5Y simulation portfolios.
- what this authorises
- 0
- compute ros.mas (ros_reemit --percentiles, LAST) and publish it on the hub as MAS-1Y / MAS-5Y marked UNVALIDATED
- 1
- run mas_validate.py daily and publish its record on /research-os/accuracy#mas
- 2
- register SIM_ROS_O11 / SIM_ROS_O12 at the hashes below (executable once the sim job runs; research-row-1.1)
- 3
- nothing else: no sizing, no rating, no stance, no optimizer, no sim selection reads MAS
- registered definitions
- config version
- ros-1.19.0
- config hash
- 61bd1d1d11e378600d952fd3db4516881ff4623ba72b1523fdd0a3bc9938489d
- lambda signal
- 0.3
- weights
- [object Object]
- block components
- [object Object]
- asymmetry
- [object Object]
- catalyst tilt
- [object Object]
- staleness multiplier
- 1
- gates
- [object Object]
- penalties
- [object Object]
- pctile method
- midrank
- collinearity warn
- 0.85
- corr ea aps warn
- 0.85
- validation status
- unvalidated
- spec
- MCH Alpha Score (MAS) — Build Specification v1.0 (owner, 2026-08-18)
- sim registered definitions
- SIM ROS O11
- 8e33fb12b347684e4129aded064a2785ef5d0de9579f6cef87ff077295ffc1a5
- SIM ROS O12
- bdeb84240ed4f0a27678e5ef2d4d17f04c5bc16e2484c78f71e1a1e42f0c59ab
- sim addendum
- addendum to
- AM-056
- n
- 2
- distinct hypotheses
- 2
- conventions version
- sim-conv-1.0
- universe version
- universe-1.0
- row schema
- research-row-1.1
- field absent rule
- not_opened, never a cash tranche
- multiplicity
- 55→57 ids, 48→50 distinct hypotheses; the 55 AM-056 hashes byte-stable (frozen golden)
- harness
- tests
- [object Object] · [object Object] · [object Object] · [object Object] · [object Object] · [object Object]
- horizons trading days
- 21 · 63 · 126
- primary horizon
- [object Object]
- lambda grid
- 0 · 0.25 · 0.5 · 0.75 · 1 · 0.3
- decision rules
- Verdicts are written only when status is `evaluable`: ≥12 non-overlapping monthly cross-sections at the primary horizon. · Block F weight → 0 only via a superseding amendment if, at ≥12 non-overlapping monthly cross-sections, the mean within-tercile slope ≤ 0 with sign consistency < 0.5. · λ, weights, gates and penalties never move on these results without a new amendment and a new multiplicity count. · Dates are never pooled; every evaluated variant is logged to trials.jsonl. · Any implied IR > 1.0 is treated as an implementation bug until proven otherwise.
- bootstrap
- [object Object]
- bug threshold ir
- 1
- sources
- engine mas_scores_<date>.json from ship date; hub ros-snapshots from 2026-07-30 (07-22..07-29 excluded: incomplete)
- prices
- ADJUSTED closes, cache-only (zero AV calls); per date, never pooled
- statement of record
- The §8 pass conditions cannot be evaluated today: no forward window has closed on any date carrying the full input set (earliest 2026-07-30; first 21-day close ~2026-08-31, 63d ~late Oct 2026, 126d ~Jan 2027). MAS ships as an unvalidated diagnostic column that drives no rating, stance or size; its accuracy record is published here under the same pre-registered anchor methodology as the rating tiers when windows close.
- position vs AM038 039
- the retired behaviour composite stood IN FRONT of unproven claims; MAS is a ranked digest BEHIND published, individually-QA'd numbers, drives nothing (QA-437/439), ships unvalidated, is never fitted
- isolation
- 0
- QA-437 per name (stamps + no rating/target/pwev key at any depth)
- 1
- QA-439 no-consumer source scan (qa_audit.py positive control)
- 2
- QA-438 complete · dated · full recompute
- 3
- tests/test_mas.py structural no-op
- 4
- mutations mas_firewall_dropped · mas_check_unregistered · mas_lambda_on_fundamental · mas_pctile_not_midrank · mas_gate_open · sim_field_absent_cash_tranche
- known input defects carried
- 0
- AM-055 SBC dilution charge OFF (disclosed)
- 1
- D-45 performance_store benchmark on raw closes (log only; the harness uses adjusted)
- 2
- corr(expected alpha, alpha/σ) ≈ 0.90 by construction — recorded, never a re-weight
- 3
- conviction term is the published ≤-CDF percentile, not midrank
- 4
- MAS is unvalidated; ~35 names eligible on 2026-08-18
- never this
- 0
- MAS never feeds a rating, a stance, a size, the optimizer, a book variant, the options intelligence or a sim selection
- 1
- no weight, lambda, penalty, gate, clip or horizon moves without a new amendment and a new multiplicity count — the config hash above is what the pass and the harness check
- 2
- MAS is never quoted above 'unvalidated' until a later amendment, on the harness's evaluable status, says so
- 3
- dates are never pooled; lambda is never selected on the grid results
- 4
- MAS is never rendered beside a decision without naming that it did not drive it
- 5
- the two sim portfolios never take a cash tranche standing in for an absent field; no portfolio is added after seeing results without a new amendment
- numbering note
- AM-056's text reserved 'AM-057' for the sim options/overlays amendment; amendments are appended, never edited — that amendment becomes AM-058
AM-056 · 2026-08-17 — Simulation Portfolio Estate — the 55 v1 portfolio definitions registered BEFORE inception (register-then-run)
- owner instruction
- Marinus 2026-08-17 (chat): 'AM-056 approved' — on the draft at docs/AM-056-DRAFT.md, which states in full what is registered, what is NOT (options/overlays = AM-057; no publication; no retrospective history) and the input defects carried. Earlier decisions folded in (2026-08-16): inception A (signal = last August snapshot, first tranche = first NYSE session of September at that session's close); execution next-session CLOSE; SPY the primary null with the equal-weight universe shown alongside; the tab additive and preview-gated.
- what this authorises
- A portfolio may execute ONLY if it appears in registered_definitions at exactly that definition_hash. sim_estate.registry.executable_ids() enforces it; the daily job accrues frozen snapshots and executes nothing while an id is absent or its hash differs. A definition is never edited in place: a change re-hashes it, it stops being executable, and it needs a superseding amendment.
- registered definitions
- SIM BASE B01
- d350130c72b2044ce09c2a2ac06d48d3afff61445277c37e1ea2117654185b3d
- SIM BASE B02
- ef1c7fd1dc5fcec5e068fe54d7164afefd4b743d08a5903776b4edcd9aed170a
- SIM BASE B03
- 8e54c1672102bbdf6cd194c28b6500738ed11ea26bdb84d21a671ad3ac839410
- SIM BASE B04 S01
- 597fda73134c529fa7ac21570d7cc688cf610fba5198d8f96dba052faf73c25b
- SIM BASE B04 S02
- 1890c195ae070238aa289e4a30c1555b0ce74fb238bcf133d286d53701a09d7e
- SIM BASE B04 S03
- df9a2cfba6f9bd0048e30d394d098d88c66c9e0f28de2095cd3f9ff95dd3c1df
- SIM BASE B04 S04
- 62006eb009c2dcbe659638a87fbfcb9a7df48f944ea7095c5254261a1b79d4cd
- SIM BASE B04 S05
- 06dd0006ea0b502d11a68a8482de0efd3ef4e85acdd05ba5c4b1c8d3057f7334
- SIM BASE B04 S06
- 6ee0264cd6b2296bdad39f35a43b645a67c89ea2bded22832e5b55803c6a1acd
- SIM BASE B04 S07
- 2363bf250575954971a78c9dbc18f12b813e991b026297d5b3d7641772bab100
- SIM BASE B04 S08
- 6c1298c8f4a8f86b524129b3b38585cb06f8161b3ea9bc28b57b5c2e91c9bdfc
- SIM BASE B04 S09
- 8a7285669fca2336e1a34eb653ba2b4914e2df639c295457780d80f749074f7b
- SIM BASE B04 S10
- e6dd8da80f142506bf846fa224bd688a97e8ea6f6809b89058c2a560a10aec8b
- SIM BASE B05
- bd34875f7c1217ea0a629ba8fa4415e348be98980828a04e5107282608f7f05d
- SIM COMBO C01
- b65805c5274c357c08d2db4652865121326e029077d724955ab6f5991a9171ec
- SIM COMBO C02
- 103ed100faaea1eebb01dc742e881d507155662d2d0d3a3c2f01363e0efd657e
- SIM COMBO C03
- 88e6ff322ed3f2806074cf81d7b7ea54b155a91c5f50bfcbf6dec74f39b6c42f
- SIM COMBO C04
- 537e43f3c25c558cd06dd083913efb777a9ad468a6f54e1c5721f1596632807b
- SIM COMBO C05
- b6e120b9b3f33b068049e6806c167c162c69e73e9c01104ba1358a162588936d
- SIM COMBO C06
- d7e5550b337fd7f43fbff8d026a391c2898cfeb7ea89fbc732754a1f39a09ab2
- SIM COMBO C07
- 8d0f262b551c4d9d42c65d637511fbb810c00c1c6b284b75d7c8a5abaf8cd4dc
- SIM COMBO C08
- 48601d16d3105e2f1726a23cc1a50cabe1fdcb0672f68a353bf2a22fbdc8525a
- SIM COMBO C09
- 4c01ee8e80ba5a18478616245b8829f57efe0c2925617abd2c08a4debe8c5630
- SIM COMBO C10
- 539510b54a61f0d59c46a1d5370e8f5ccd331259688cfd601da8fd1708087214
- SIM COMBO C11
- 6efbe8b3b289fddab0869b2d94741e48b4ab70d05ad23aabec55b59c88d9abfd
- SIM COMBO C12
- 9215278dd344e2a9defdfb771874b53a32a966fada5e06b714a18060123a1c37
- SIM RANK R01
- 8a2a0478ee5a779f07f4dcbc8520a3cb7e9de946d0e335280a1e01c38f102398
- SIM RANK R02
- 8bc4508885f49099ad01202ef935c04d370647a4d4a1c7a049805887cdfb2fdc
- SIM RANK R03
- 0e74c48b8330c78e6d6b221674be210e2c85403181bd1f75dc3eb716bf04fed6
- SIM RANK R04
- a5f8657ff6ddb76a8eeecdb6b1739c5004df7da1e9bfbe938e3cb1dc4a6e497f
- SIM RANK R05
- eb199c64223fb740fb19a00798daaec73aa40670b14fbccf1786aabd8e41f3c7
- SIM RANK R06
- 269fd5f3d40468d2e97045ddc57a884710d70febb0229643f1d36a57a966d811
- SIM RANK R07
- 07a537615b2ae33caabe32a18cb6a5574d456e29a55ba2ac250bc47bfb78acbb
- SIM RANK R08
- 3c3752a3471c6a60f16fba99953e6ef631c8e43df0a9fff3647502f0b0c0eb58
- SIM RANK R09
- a61a30cf8ad66461e6eda92c571e9d976b00739934f12c86568f4a6b62794718
- SIM RANK R10
- f96bd229d1598dbe8d6c7ae67206b59a122af71899a53d75f687b9e48f5a2089
- SIM RANK R11
- f2d5e2b28840ed3bef6f011cfae2055218d0667c3c600d1dfebc520f02cbe6f3
- SIM RANK R12
- c73f9b23411514a4ba2abf788738509f626fd7ff9554a839093e47309c8a6a4c
- SIM RANK R13A
- 941148bf6565bce2a0aabc44eb111581723ce0d90cc8e516b82af66779ddceff
- SIM RANK R13B
- c6672c287141f84cdeab5fca0159f94f82483886272b3c9b507a32282ab5d644
- SIM ROS O01
- 43ac766083ba00316ed3c53c9aeee9db1e0c8dffe532c3a189006d0510c9586f
- SIM ROS O02
- 48eb849240e80cace904bcf0c8fbf5a72f51da56b1e78d3c6c2ce274641d6438
- SIM ROS O03
- fd70c535dcd5fb649310ea9768ca069a71676a68f2cccd8d77930dc77bbccf4d
- SIM ROS O04
- 6adec7af0e92d5f51eaa42377f24055ddcb4670e5b7ab9319172d7ad586af004
- SIM ROS O05
- fcfb39542f7b56e1f12735d3dd10d4c60de8f9adc617f6ef6cff5cf39bad83a1
- SIM ROS O06
- 71cebc8e5348a15fa0b01adcbaba53808e7b16dbffbf4f189e53775d9e8208f8
- SIM ROS O07
- e99a518f3994eca493a8c380705e4e5e21ee70b098388652f7628e8463c7e8d3
- SIM ROS O08
- ab34a97a818802a71cec3dbf66683228c69d36cc74e5eebf18c82c091f4a24ad
- SIM ROS O09
- a56ff01ce690f12c575c25c66de53910b13dbb135ca838c084b45b92c8af2056
- SIM ROS O10
- 4c4bb76ee60374ef18ab5a43705783fa78893fa5b40096b99abc95341489d0a4
- SIM STATE S01
- 91be109bbe01979a80815c38350f74c3aa85ca0fc119fd863d62ca17a0839c5b
- SIM STATE S02
- fd7a95d0aeb335e209d93b3226244ec66af3760e05944f1b0afcca1601e7d58b
- SIM STATE S03
- eb76b3640f71095ba27cb5b6a7574b97d671e0a6271282ec90d7b8f4d3d074ad
- SIM STATE S04
- e839db3a5d4e7b0ba734d6b47c9e4b7fefa42ffe45d6470ea1bfab04f9438c7f
- SIM STATE S05
- 96b4ec6dfebf4bc5a0f8fd38ce311b6558be06b799cc1264117677df010b5546
- n
- 55
- distinct hypotheses
- 48
- duplicates collapsed for multiplicity
- 0
- SIM_ROS_O05→SIM_ROS_O04
- 1
- SIM_ROS_O07→SIM_RANK_R11
- 2
- SIM_ROS_O08→SIM_RANK_R12
- 3
- SIM_ROS_O09→SIM_RANK_R08
- 4
- SIM_COMBO_C06→SIM_COMBO_C02
- 5
- SIM_COMBO_C09→SIM_STATE_S02
- 6
- SIM_COMBO_C12→SIM_STATE_S01
- hash source
- definitions (registry not yet born — the first run writes these same hashes)
- conventions
- version
- sim-conv-1.0
- capital
- [object Object]
- timing
- [object Object]
- prices
- [object Object]
- costs
- [object Object]
- dividends
- credited as cash to the tranche on the ex-date (amount × shares held), never reinvested
- corporate actions
- cav-1.0 — hold, never drop: splits by declared coefficient; cash acquisitions close the lot at the deal price on the close date; ticker changes re-key the same lot; delisting requires two of {LISTING_STATUS date after open, price stall ≥5 sessions vs SPY, absent from the estate}; spin-offs, stock mergers and bankruptcies are MANUAL events registered before they are applied
- benchmarks
- [object Object]
- evaluation
- [object Object]
- terminology
- Simulation Portfolio · hypothetical capital · not personalised investment advice
- universe
- version
- universe-1.0
- base
- every name in the declared canonical estate snapshot
- excluded
- deal_quarantined · review_held · price_quarantined · split_quarantined · delisted · price_data_anomaly · asset_type != Stock · no price in the snapshot
- per portfolio
- a name whose selector metric is missing/None is INELIGIBLE — never imputed
- fwd pe
- the 19 EV/EBITDA-lens names carry an equity/EBITDA multiple in fwd_pe and are ineligible for R5/R6 (fwd_pe_lens == 'ev_ebitda'); non-positive fwd_pe ineligible
- band
- a stale 52-week band (spot outside the recorded band) makes R2/R3/R4 ineligible
- every exclusion named
- true
- b4 seeds
- 0
- 1009
- 1
- 2003
- 2
- 3001
- 3
- 4007
- 4
- 5003
- 5
- 6007
- 6
- 7001
- 7
- 8009
- 8
- 9001
- 9
- 10007
- spec
- MCH_Research_OS_Simulation_Portfolio_Estate_Full_Specification.md (38 §, 2026-08-16)
- draft reviewed
- mch_starter_kit/docs/AM-056-DRAFT.md (engine repo, commit c778753)
- legal wording version
- sim-legal-1.1
- evaluation
- primary metric
- net annualised time-weighted-return alpha vs SPY (SIM_BASE_B01) with identical cash flows
- shown alongside
- alpha vs the equal-weight eligible universe (SIM_BASE_B03) on every surface, so the size/equal-weight effect stays visible; XIRR is investor-experience only and never ranks
- minimum before any read
- 12 months and 9 tranches
- maturity tiers
- [object Object]
- retained or rejected only at
- Decision-grade
- multiplicity
- Benjamini-Hochberg q-values within family and across the estate + deflated Sharpe; hypotheses counted net of duplicate_of (48, not 55)
- risk free pct
- 4
- known input defects carried
- 0
- AM-055: the SBC dilution charge is OFF on 839 of 858 names — the PWEV portfolios select on a gross per-share figure (recorded per definition in known_input_defects).
- 1
- Forward P/E is withheld on the EV/EBITDA-lens names; they are ineligible for the P/E portfolios.
- 2
- 30 names carried a stale 52-week band on 2026-08-16; ineligible for the band portfolios and named per tranche.
- never this
- 0
- Never add a portfolio after seeing results without a new amendment and a new multiplicity count.
- 1
- Never re-select a tranche because the signal was later revised — a correction is a superseding row, never a re-selection.
- 2
- Never quote a portfolio above its maturity tier, and never sort a published surface by return.
- 3
- Never splice a backtest into a prospective series, and never back-date a missed tranche.
- 4
- Never treat the Research OS 'Max size %NAV' field as an allocation.
AM-055 · 2026-08-16 — SBC dilution charge left OFF (owner decision) — the zero charge is DISCLOSED on every page rather than left silent; owner second-source price verifications (LITE/WDC/TTD/SN) recorded; the AM-041 stage-6 choke-point escape that failed the estate publish; peer-anchor prose written from what is computed
- owner instruction
- Marinus 2026-08-16 (chat): 'Switch off SBC' → ambiguity flagged (turn OFF the 19 configured rates, or do NOT roll the charge OUT to the 839 names at 0.0?) → 'leave SBC off'. Read as: the status quo stands — the dilution charge is NOT rolled out estate-wide; nothing that is currently charged is changed without a further instruction.
- state registered
- annual dilution pct zero
- 839
- annual dilution pct nonzero
- 19
- nonzero names
- [object Object]
- note
- The 19 hand-configured rates keep charging pwev = gross/(1+d) exactly as QA-512 checks (BRK.B −0.5% = net buyback accretion). Zeroing them would move 19 published PWEVs/ratings — not done under this amendment; it needs its own instruction.
- render disclosure
- mch_report_template: when annual_dilution_pct is zero/absent and a PWEV exists, the Probability-Weighted Return Profile prints '*Share-count charge: none applied. The probability-weighted value above is the gross per-share figure — no annual dilution is deducted — stock-based compensation runs at {measured sbc_pct} of revenue; free cash flow net of SBC is {fcf_minus_sbc_b}. SBC is therefore disclosed, not charged.*' The two facts come from extended.accounting_quality / extended.fcf_quality and are dropped (not invented) when unmeasured. Under the house standard SBC is a real cost, so a ZERO charge is a stated position, never a silent absence.
- price verifications
- file
- data/price_verifications.json
- names
- [object Object]
- finding
- Owner second-source spot CONFIRMED on all four; each name's recorded 52-week band is the wrong number (LITE/WDC ranges span ~10×; TTD/SN spot sits outside the recorded band). Recorded and never used to overwrite a fetched price; the engine's AV close stays the priced input. Disposes of the 'verify a PRICE-DATA-ANOMALY name against a second source before publishing' rule for these four.
- operational defects found
- 0
- weekly_snapshots stage 6 (AM-041 behaviour scoring) shipped with its own subprocess.run beside _run_module — the single choke point tests stub. tests/test_vol_xs_expiry runs weekly_snapshots.main() 'with no subprocesses'; the stub never reached stage 6, so the CI gate ran the REAL behaviour_score pass over the estate: 849 cold daily_adj refetches on EFS in 687 s, and the 2026-08-16 estate publish (run-once.sh daily, 16:17 SAST) failed at stage 0 on the hygiene line. Fixed by routing stage 6 through _run_module(ok_codes=(0,3)); the suite now BLOCKS subprocess.run/Popen inside main() and fails by name; mutation snapshots_stage_escapes_choke_point certified. NOT declared as a fetcher-TTL write — it was never that class.
- 1
- Peer-anchor prose printed a dangling dash where the engine carries no P/E-implied price ('12.945x → —' in 741 of 858 figure captions; 'implies **—**' in prose; unrounded floats like 25.634999999999998x in 168). QA-515 (BLOCKING at publish since T3) would have HELD ~741 names on the next publish. The T3 sweep missed it because it rendered from a summary-only tempdir and never produced a <figure>. Fixed: prose and captions are written from what IS computed (EV/Rev re-rate when the P/E anchor is absent; the reason named — EV/EBITDA lens or no forward-EPS basis); multiples formatted to one decimal; a real-render sweep (charts present) is now the pre-publish check.
- never this
- 0
- Do not read 'leave SBC off' as permission to zero the 19 configured rates, and do not roll the charge out to the 839 without an instruction that names it.
- 1
- Do not declare a gate write in ALLOWED_WRITES to make a red hygiene line go away before the mechanism is known — the 849 refetches were a wiring escape, not a fetcher TTL.
AM-054 · 2026-08-16 — Mid-cap narrative pass (367 names off the Python template) + the vendor-data defects it exposed: wrong company descriptions, placeholder SBC, mislabelled EV/EBITDA multiples
- owner instruction
- Marinus 2026-08-16 (chat): 'the reauthor is great. approved. what about the rest of the 855?' — batch 1 approved, all 479 stamped reviewed, and the remaining mid-cap tier authored bespoke.
- scope
- 367 sp900_enrich/2 names (tokenised Python templates, one archetype skeleton per cluster) re-authored bespoke by 16 parallel Claude Opus 5 sessions, drafted STRICTLY from each name's own payload (description, segment drivers, named exposures, scenario tree) with an explicit prohibition on inventing products, customers, contracts or dates. 12 names whose pre-existing authored prose already passes the guard were left alone.
- render corrections
- 0
- profile.description is WITHHELD when it does not name its own company (mch_report_template.description_names_company, tested on the description's leading entity so a generic word cannot satisfy it, with a strict full-name containment fallback for 'Founded in 1902, Lamar Advertising…' cases). 16 of 858 withheld — 7 describe a DIFFERENT company (CART a bank holding co, COR CoreSite, P Pandora, Q IQVIA, RBC Regal Beloit, SN Sanchez Energy, UHS UnitedHealth), 5 are predecessor names, 4 are unusable fragments. The Company Overview had been printing these verbatim under the wrong company's name.
- 1
- {{sbc_pct}} now resolves from the MEASURED extended.accounting_quality.sbc_pct_of_revenue, and is WITHHELD when only the engine default is available: summary.sbc_pct_revenue is exactly 0.01 on 840 of 858 names while the measured figure exists for 716 of them (ABNB 12.5%, TWLO 11.3%). Publishing 'SBC ≈ 1% of revenue' estate-wide is a false disclosure under the house SBC standard.
- 2
- {{fwd_pe}} is WITHHELD for the 19 EV/EBITDA-lens names (reconciliation.equity_ebitda_multiple present): that figure is an equity/EBITDA multiple, so 'WHR trades on 3 times forward earnings' is simply wrong.
- state flagged for owner decision
- 0
- ARCHETYPE MISMATCH is systematic in the mid-cap tier: ~60 names carry a driver set, segment label and scenario vocabulary from a different industry (LOPE education→waste/uniforms services; M Macy's→online marketplace GMV/take-rate; SCI funeral→marketplace; RYN timber REIT→office/hotel obsolescence; THO RVs→toys/screen substitution; VNOM royalties→fee-based midstream; NYT→Dow Jones/REA drivers; PINS→interactive entertainment; KBR, SAIC, ARMK, BCO, HRB, HQY, HLI, GLPI, MUSA, SMG, SGI, SHC, SIRI, OLLI, PII, PFGC, PAG, OLED and more). The valuation math is unaffected (EPS x PE / scenario tree) but the rendered drivers and scenario NAMES are wrong. Narratives were written to the business and kept clear of the wrong vocabulary; the configs need re-archetyping.
- 1
- annual_dilution_pct is 0.0 on 839 of 858 names, so the PWEV carries NO SBC dilution charge at all — against the house standard of ~3%/yr for growth SaaS. This is a VALUATION-level input; changing it moves every target, so nothing was changed. Owner decision.
- 2
- 30 names sit outside their recorded 52-week range with NO sanity warning (the refresh fires selectively); 3 names are priced behind the estate (EA 08-04, TMHC 07-24, NSA 07-22); ticker P carries a three-way identity conflict (config name Everpure, description Pandora, archetype hardware/storage) and is not decision-grade; WHR/LAD/ARW/CHDN guided-EPS inputs look wrong beyond the lens issue.
- 3
- 10 names still carry the engine's PRICE-DATA ANOMALY flag (LITE, WDC, TTD, SN and others).
- never this
- As AM-052/053. No valuation input changed by this pass; every defect above is reported, not silently corrected.
AM-053 · 2026-08-16 — Estate narrative re-authoring pass (Option 1 at scale) — 476 names drafted by Claude Opus 5 under supervision; render-token and QA-511 corrections found by the pass
- owner instruction
- Marinus 2026-08-16 (chat): 'approve batch 1 and continue with the full estate and publish to the hub'. Batch 1 (AAPL/MSFT/NVDA) stamped reviewed_by Marinus.
- method
- 20 parallel Claude Opus 5 sessions, each given a fixed slice and the same contract: read mch_author --show (existing prose, refusals, tokens_today, scenarios, segments, exposures, sanity warnings, 52-week range), PRESERVE the existing analysis, convert every figure to a {{token}} or a qualitative phrase, delete stale DIRECTIONAL claims the guard cannot see, write only through mch_author.py (guard at write, all-or-nothing, provenance stamp, tier llm-drafted + RESEARCH cap). No session ran git, the gate, or any pipeline stage.
- what the pass found
- 0
- STALE VERDICTS, not just stale numbers: ~160 names carried a rating word contradicting today's class, and on a material subset the published prose argued the OPPOSITE conclusion to the live rating (APP, APTV, FLEX, CHRW, KLAC, LII, DVA flipped to a buy class; DASH, DGX, DPZ, CBOE, CME, CLX, LDOS, IT, IVZ, JKHY, KKR, J, JPM to a sell class). Those theses were re-oriented, not merely de-numbered.
- 1
- SIGN ERRORS in authored prose: DVN, EOG and EQT described a net-CASH cushion where canonical state carries net DEBT; EXE likewise. The balance-sheet claim was inverted in the published reports.
- 2
- FROZEN 52-WEEK EXTREMES and 'target sits above spot' claims that had gone false with the tape (BNY, BR, BKNG, BBY, BAC, BRK-B, CCI, CME, CFG, CPRT, CSGP, CTSH, CRM, CSX, COIN, CPT, LULU, MCD, LNT, LYB, MCK, MDLZ, NI, NXPI, NEM, NOC, NOW, NTAP, NWSA, NUE, OKE, ORCL, OXY, PAYX and more).
- 3
- UNSUPPORTED specifics: BRK-B's 'two-thirds of sum-of-parts in securities and cash' (payload says 47%); AKAM's capex share contested by its own exposure block; FSLR's '15.5% blended margin' against a 37% mix-weighted margin; CAG's mis-stated modal scenario.
- corrections made during the pass
- 0
- render_tokens {{op_margin}}: one decimal below 10% — BG's 0.5% rendered as '0%' (round-half-even) and read as 'no operating profit at all'; 11 estate names sit under 2%.
- 1
- render_tokens {{net_cash_b}}: UNAVAILABLE (sentence drops) for bank-like names (Banks/Insurance/Reinsurance/Mortgage/Consumer Finance sub-industries) — a deposit-funded balance sheet's 'net debt of ~$725.6B' (C) or 'net cash of ~$350.6B' (JPM) is an artifact, not leverage a reader should weigh; and UNAVAILABLE below $0.05B, where '~$0.0B' asserts a balance sheet that nets to nothing (CCL, CFG, FDXF, 11 names).
- 2
- QA-511 false positive corrected: 'target' alone matched a falsification metric named '…growth guidance (multi-year target)' (EIX), so a CAGR threshold was read as a price gap. The spot-relative context now requires price-relative phrasing, and the falsification section is excluded (its thresholds are per-metric by construction). Before: 1 blocking in 22 sampled; after: 0 in 30.
- state flagged not fixed
- 0
- 10 names carry the engine's own PRICE-DATA ANOMALY (DO NOT PUBLISH) warning — ABNB, CRL, GRMN, HONA, LITE, MPC, SN, TTD, VLO, WDC. Mostly stale recorded 52-week extremes that the engine refreshes from close history at run time; LITE (9.8x range), WDC (11.0x) and TTD (spot far below the recorded low) need a second-source check before their numbers are trusted. Reported, not silently published as sound.
- 1
- 193 names sit below the 30% P(>current) sanity band — a bear-weighting/opex calibration question, disclosed in-prose by the drafters rather than written around.
- 2
- CDE (Coeur Mining) is modelled by its payload as a copper producer; CRS (specialty alloys) inherits the diversified-industrial-machinery archetype. Both narratives were drafted to the payload and kept general; the mappings need an owner decision.
- 3
- Hand-authored falsification rationales carrying frozen figures remain WITHHELD at render (the T3 residue) — a later authoring batch.
- never this
- As AM-052: no draft written unvalidated; no run-time generation; no DECISION-LEVEL without reviewed_by; no rating, target or probability changed by prose.
AM-052 · 2026-08-16 — Option 1 — supervised LLM narrative authoring: the writer (mch_author.py), the RESEARCH review cap, provenance on the page; pilot batch AAPL/MSFT/NVDA
- owner instruction
- Marinus 2026-08-16 (chat): 'start option 1' — Claude Code sessions (Opus 5 on the Max subscription, no API) draft narratives; publication stays the existing AWS nightly.
- contract
- 0
- mch_author.py is the ONLY path from a draft to data/company_context/<T>.json; it never composes prose. Every field is validated with the SAME mch_claims.narrative_guard the renderer uses (predicates 1–5), against a registry built from the name's LIVE canonical summary; any refused field refuses the whole write (nothing written, reasons named).
- 1
- Tokens only: every figure/verdict is a {{token}}; facts without a token are stated qualitatively; nothing invented; ratings/targets/probabilities untouched.
- 2
- Provenance stamped and rendered: authored_by (model · mode), authored_model, authored_date, author_version author/1, authored_against, tier llm-drafted, reviewed_by/date. The report prints 'Narrative drafted <date> by <model> under supervision; awaiting analyst review (RESEARCH tier)' until reviewed.
- 3
- REVIEW CAP: mch_publication_tier caps an unreviewed llm-drafted narrative at RESEARCH; mch_author --review <T> --by <owner> lifts it at the next engine pass. Nothing an LLM drafted reaches DECISION-LEVEL unreviewed.
- 4
- No LLM call anywhere in the pipeline; drafting is off-cycle (laptop/cloud sessions), committed as data through the pre-push gate; the AWS nightly (Tue–Sat 03:00 SAST) publishes whatever main holds by ~02:55.
- 5
- Worklist at start: 479 names (467 Claude-authored S&P 500 contexts refused today for frozen literals/rating words, 12 with no narrative); the 374 sp900_enrich/2 names are not on it (they land at the next engine pass).
- pilot
- AAPL, MSFT, NVDA re-authored 2026-08-16 by claude-fable-5 (this session; the analysis kept, frozen figures → tokens/qualitative; stale directional claims removed) — 3/3 pass the guard, 0 document_qa blocking on render, tier RESEARCH pending owner review (digest artifact). Suite tests/test_author.py 10/10 registered in the gate; mutations author_write_unvalidated + review_cap_dropped certified.
- never this
- No draft written unvalidated; no run-time generation; no DECISION-LEVEL without reviewed_by; no rating/target/probability changed by prose.
AM-051 · 2026-08-16 — Scenario maps re-authored under owner sign-off — 22 maps / 11 clusters; ratchet baseline 150 → 0
- owner instruction
- Marinus 2026-08-16 (chat): 'follow the recommendations' — A (adopt the --suggest vocabulary-rank proposal) for the 20 three-state members incl. RGEN/IPGP (their 20% downside read as the growth lens's view); B (hand-authored) for LRCX and META.
- what changed
- 0
- data/industry/{ai_compute,disc_autos,disc_retail,disc_travel,health_pharma,ind_machinery,ind_transport,it_hardware,materials_commodity,materials_metals,staples_food_bev}.json scenario_map re-keyed to each member's LIVE config scenario names; per-member sign-off note under _map_notes (the original general note kept as _general).
- 1
- A (20): Structural+Cyclical → downside state, Base → mid-cycle, Upcycle+Peak → upside; implied ≈ 44/32/24 vs house ≈ 37/35/28 for the ev_ebitda names; RGEN/IPGP (ev_ebitda_growth) 20/56/24 vs 37/35/28 — a member-vs-house divergence the Industry Context table exists to show.
- 2
- B (LRCX): AI Capex Bust ← Structural — WFE Reset / China Restriction; Digestion ← Cyclical Downturn — Capex Cut; Sustained Build ← Base — Normalised WFE (modal ↔ modal, the shape the previous author used); Supercycle ← Upcycle — Leading-Edge / HBM Capex + Bull — Supercycle Re-Rate → implied 20/17/35/28 vs house 22/20/38/20.
- 3
- B (META): the live config's 'Recession / TikTok' (the brief mapped the stale 'Recession / TikTok Squeeze'); otherwise unchanged → 20/15/30/35 vs house 22/20/38/20.
- 4
- data/industry_consistency_baseline.json rewritten 150 → 0 (may only shrink; the gate now blocks on ANY mapping problem). industry_consistency --check-all: 0 problems.
- 5
- Effect: no published number changes; the Industry Context 'withheld — map does not match' block returns as a real implied-vs-house table for the 22 members at the next engine pass (Tue 2026-08-18 nightly).
- never this
- No scenario_map edit without owner sign-off; no probability invented; a state with no mapped scenario is not created (a 0% row would be the failed-mapping signature).
AM-050 · 2026-08-16 — T3 build addendum — deviations from AM-049 as registered, recorded before the push (append-only; AM-049 stands as written)
- owner instruction
- Standing: amendments appended never edited; every deviation from a registration is itself registered.
- deviations
- 0
- {{roe}} WITHDRAWN. AM-049 registered {{roe}} in the token vocabulary; canonical state (summary.json) carries no return on equity, so the token could only resolve to an unavailable value. Authored prose has no ROE clause; the token key is not offered — an authored {{roe}} stays literal and visible to QA-5xx.
- 1
- {{op_margin}} resolves from the segments the report's own segment table displays (mix-weighted), not from an AV overview figure — same source as the table, so prose and table cannot disagree. {{net_cash_b}} resolves from reconciliation.net_cash_billions (signed).
- 2
- UNAVAILABLE ⇒ SENTENCE DROPPED: resolve_tokens removes the sentence containing a token whose canonical value is None (never prints a dash where a figure was promised). Missing input ⇒ the narrower statement.
- 3
- CLUSTER = BRIEF MEMBERSHIP (root cause found in build): the ev_ebitda / ev_ebitda_growth valuation adapters carried cluster 'ind_transport' (set for the first airline/rental batch) and sp900_enrich resolved a name's cluster from the archetype — 15 of 20 adapter names live in a different brief (OLN → materials_commodity). Adapter cluster fields nulled; sp900_enrich resolves cluster from data/industry/*.json membership, the same source mch_weekly_run reads.
- 4
- EDITORIAL GATES PROMOTED: AM-049 registered them WARNING-tier this tranche. Measured on an 88-name rendered sample after the render fixes (role labels, structure prose, load-bearing sentence, overlay reasoning de-dupe, live disclaimer text, identifier labels): QA-514 identifiers/config strings, QA-515 empty parentheses/dangling dashes and QA-517 paths/e-mails are BLOCKING at 0 sample hits; QA-513 duplicate sentences and QA-516 mixed spelling stay WARNING (2 sample hits). QA-511 spot-relative % BLOCKING as registered (0 sample hits after the sign-carried-by-word and known-percentage refinements).
- 5
- GUARD PREDICATES 4 and 5 added to mch_claims.narrative_guard: an internal identifier / config string in authored prose; a frozen anchor COUNT contradicting anchors.n. Both apply at write time and render time through the one implementation.
- 6
- --suggest emits proportional worst→best RANK proposals (config scenarios onto cluster states), not edit-distance: config names and brief states share almost no vocabulary (LRCX: 5 mapped names, 0 in config), so edit distance had nothing to match. Proposals written to output/industry_map_suggestions_<date>.json — 22 maps / 11 clusters — for owner sign-off per map; nothing applied.
- 7
- sp900_enrich/2 no longer authors catalyst_events (the /1 frozen next-earnings date is the OLN 'past Next catalyst' class); an /1-authored entry is removed on re-authoring; hand-authored events are kept.
- 8
- 374 engine-enriched contexts re-authored on 2026-08-16 (all stale by author_version; 374/374 passed write-time validation; second pass selects 0). Their tokenised prose reaches the estate at the next weekly merge (08-22); until then the guard keeps them on machine fallback as today.
- 9
- profile.description NOT routed through the guard (registered as a T1-residual surface): the vendor description is descriptive text; QA-502's aggregate exclusion covers its dollar figures. sp900_enrich refuses to embed a description sentence carrying a literal $/% (WPC's 2020 '$ 18 billion' EV) and uses the archetype label instead.
- 10
- NO-BRIEF NAMES (17 of 374; bullet added before AM-050's first commit): a name in no industry brief has no industry_context in its report, so cluster_resolver returns None for it (NO archetype fallback — that field was the contamination source) and the thesis authors 'and its own scenario tree' instead of 'and the {{cluster_label}} house view'; otherwise the whole rating sentence would have been dropped at render as unavailable. --stale-only detected the 17 as moved (archetype cluster → None) and re-authored them; second pass 0.
- measured
- 88-name rendered sample: blocking findings 72 → 1 (BRK-B PWEV, a state defect under separate trace); QA-505 warnings 116 → 48 (own-cluster label whitelisted); mutations certified caught: claim_helper_bypassed, enrich_writes_unvalidated, enrich_literal_restored, editorial_gates_dropped (4/4).
- never this
- As AM-049.
AM-049 · 2026-08-16 — Document-level reconciliation, tranche 3 (registered before code): re-authoring the authored layer as tokenised, verdict-free prose; the estate scenario re-map; the rating CLAIM() helper; editorial gates
- owner instruction
- Marinus 2026-08-16: 'start T3'.
- the defect
- sp900_enrich.py interpolates LITERALS at author time — the stance word ('trading rich to'), the upside (up:+.0%), operating margin, ROE, net cash, the engine's rating, the archetype id and the CLUSTER'S state names — into thesis_narrative, anti_thesis_narrative, scenario_macro and the falsification trigger. Every one freezes: OLN's 'the engine's SELL … (-67%)' was true on 2026-07-22 and false by August. The AM-047 guard now withholds 833 theses / 810 anti-theses / 650 scenario-macro rows / 2,286 rationales to machine fallbacks. The guard is CONTAINMENT; this tranche is the FIX.
- what t3 builds
- 0
- TOKENISED TEMPLATES in sp900_enrich: every figure and verdict becomes a {{token}} resolved at render from canonical state — the existing vocabulary ({{spot}}, {{pwev}}, {{tri}}, {{fwd_pe}}, {{base_target}}, {{sbc_pct}}, {{date}}) grows by {{rating}}, {{stance_word}} (cheap-to / rich-to / fairly-valued, computed at render from tri vs spot), {{tri_vs_spot}}, {{op_margin}}, {{roe}}, {{net_cash_b}}, {{cluster_label}}. Prose carries NO literal %, $ or rating token. resolve_tokens gains these keys; unresolved tokens stay visible (the existing QA discipline).
- 1
- SCENARIO_MACRO from the LIVE config's scenario names, keyed by NAME with the cluster state referenced by {{cluster_state:<i>}} tokens rather than a frozen label — so a cluster re-map never leaves a stale taxonomy in prose. The falsification cluster_link is dropped as an authored field (the report resolves the name's live cluster).
- 2
- sp900_enrich --stale-only: selects names where the AM-047 guard refuses any field, or scenario_macro keys are not the live config's scenario names, or the frozen authored_against.cluster != the live cluster; --stale-only --dry-run lists them.
- 3
- WRITE-TIME VALIDATION: after building ctx, sp900_enrich runs the SAME mch_claims.narrative_guard (imported, never duplicated) against a registry built from the name's current summary; a context the renderer would refuse is NOT written (named, counted). Plus an authored_against stamp {cluster, config_scenario_names_sha256, date} so drift is detectable without string forensics.
- 4
- SCENARIO-MAP RE-AUTHORING for the 11 red clusters (150 baseline problems): industry_consistency --check-all --suggest emits an edit-distance suggested mapping per member; each cluster's scenario_map is re-authored ONLY under owner sign-off (authored data, human-signed per map — the ratchet baseline shrinks as each lands).
- 5
- CLAIM() helper: the 8 rating render sites in mch_report_template read the registry value through one helper; an AST guard suite asserts no `rec` f-string interpolation survives outside it (the migration path toward claim-referencing rendering, rating sites first).
- 6
- EDITORIAL GATES in document_qa (spec §8.5, WARNING-tier this tranche, ratchet later): duplicate adjacent sentences; raw snake_case tokens and config-version strings (ros-1.x) in reader-facing sections; empty parentheses; dangling em-dashes; mixed UK/US spelling in one document (individualised vs individualized); QA-511 percentage residue promoted to BLOCKING for spot-relative percentages only.
- 7
- The T1 residual frozen-figure surfaces (profile.description, moat prose, recommendation prose — 184 blocking) routed through the guard as the thesis is.
- effect and cadence
- No published number changes. As names are re-authored and pass write-time validation, their theses return from machine fallback to authored prose — per re-emit, per name, tier moving RESEARCH toward DECISION-LEVEL only when every other gate is clean too. Scenario maps land per owner signature; the ratchet baseline shrinks accordingly.
- fixtures and mutations
- A tokenised template rendered against two different spots must produce two different upside figures and ONE stance word each (no frozen literal). Write-time validation must REFUSE a context carrying a literal rating/percentage. Mutations certified caught: enrich_literal_restored (a template interpolating up:+.0% again), enrich_writes_unvalidated (write-time guard bypassed), claim_helper_bypassed (a rating site interpolating rec directly).
- never this
- No authored literal number or verdict; no scenario_map edit without owner sign-off; no re-authoring that changes a rating, target or probability; write-time validation and render-time guard share ONE implementation.
AM-048 · 2026-08-15 — Document-level reconciliation, tranche 2 (registered before code): chart + PDF manifests with vintage blocking, and computed publication tiers
- owner instruction
- Marinus 2026-08-15: 'begin with T2' — confirmed estate-wide (all 858 names), OLN is the fixture not the scope.
- the defect
- Charts and the PDF render at ENGINE time from engine state; the body renders later from summary.json. Nothing stamps or compares them. OLN's published dashboard chart says 'HOLD | Target $24 | Spot $22 | 2026-07-30' and its PDF carries $22.17 x6 against a BUY / $18.82 body — three spots, two ratings, two dates in one bundle, all three artifacts passing every existing check.
- what t2 builds
- 0
- charts_manifest.json per ticker directory, written at ENGINE time by ONE _save_chart() wrapper over the 10 savefig sites in mch_stock_engine.py: per chart {file, sha256, rendered_at, price_date, spot, rating, target, engine_version} read off the engine object (the same values generate_dashboard's suptitle already interpolates). mch_stock_pdf appends a 'pdf' entry with the same fields. ros_reemit already copies sidecars; publish-ticker vendors the manifest beside the PNGs.
- 1
- document_qa QA-503: manifest.price_date != summary.price_date, or manifest.rating class != registry rating class, or |manifest.spot - registry spot| > tol => BLOCKING. Missing manifest => WARNING until the next engine pass has minted them estate-wide, then BLOCKING (flip is a one-line constant, recorded in the checkpoint). No retro-minting from file mtimes: a vintage nobody stamped is unknown, not inferred.
- 2
- document_qa QA-504: pdftotext extraction of the PDF; every rating token and every spot-magnitude dollar in the PDF text reconciles to the registry exactly as the md does (QA-501/502 applied to the PDF); the pdf manifest entry vintage-checked as 503. Absent pdftotext in estate mode is a HARD error (a gate that silently skips is the estate-wide summaryPath skip again); under --ci the estate sweep is already skipped.
- 3
- mch_publication_tier.py: tier computed, never authored — DECISION-LEVEL requires no RED, 0 qa_audit blocking, 0 document_qa blocking incl. 503/504 clean, manifests present, no quarantined narrative fallback, industry cross-check present (not withheld); RESEARCH = no blocking but degradations (machine-fallback narrative, withheld industry block, missing manifest pre-flip, synthetic single-100%-segment provenance); DRAFT = any blocking finding / ANALYST holes / --force. Persisted per name (<wk>/<T>/publication_tier.json), stamped at publish into the page badge and the data JSON — NOT into the md (post-emit checks feed it; md stamping would churn the incremental-publish sha). publish-ticker: DRAFT held via gateFail() and reported; --force publishes at RESEARCH MAXIMUM — force can never mint the top label.
- 4
- First slice of spec §8.3 arithmetic recomputation on the header strip: market cap = spot x diluted shares; upside = FV/spot - 1; PWEV = sum(p x target) — recomputed independently by document_qa from the registry, tolerance stated per quantity.
- fixtures and mutations
- Stale-chart pair (manifest price_date != summary) must fire QA-503; PDF-vintage pair must fire QA-504; a planted DRAFT name must be HELD and a --force publish must stamp RESEARCH not DECISION-LEVEL. Mutations certified caught: manifest_missing_tolerated (post-flip, absent manifest passes), vintage_check_dropped, force_mints_decision_level.
- checkpoint
- Rides the next mch_full_update ENGINE pass (manifests are minted at engine time — a template re-emit cannot mint them): the pass after this lands writes every manifest at zero incremental cost -> document_qa --estate with 503/504 -> flip missing-manifest to BLOCKING -> tier stamps live -> held-name report.
- never this
- No manifest vintage inferred from mtimes; --force never lifts a blocking finding or mints DECISION-LEVEL; the tier is never written into the md; the estate re-emit contract (summary.json untouched by re-emit) stands — manifests are ENGINE artifacts.
AM-047 · 2026-08-14 — Document-level reconciliation, tranche 1 (registered before code): claim registry, whole-document scanner QA-501..510, narrative quarantine, loud scenario joins, and the two owner-decided layout changes
- owner instruction
- Marinus 2026-08-14: the OLN-derived QA remediation spec, FULL programme approved via plan mode; past catalysts removed entirely; Company Overview after the decision card.
- what t1 builds
- 0
- mch_claims.py build_registry(summary) -> claims.json per name at re-emit. Schema claims.v1: id -> {value, basis, as_of, tol}. Contradiction matrix v1: rating.coarse vs rating.stance_machine is a DESIGNED divergence requiring the strip's reconciliation sentence; all other rating-class pairs must agree. Values derive ONLY by importing the existing single homes (coarse_rating, mch_triangulation) — nothing re-derived.
- 1
- mch_vocab.py: runtime cluster vocabulary (never persisted) from data/industry/*.json + matrix archetype labels; a name's forbidden set = other clusters' full em-dash state names and labels.
- 2
- narrative_guard(text, registry, vocab): refuses authored prose containing a rating token outside the registry's rating class, literal %/$ figures not produced by {{tokens}}, or another cluster's vocabulary. Refused field renders the machine fallback + provenance note + a row in output/narrative_quarantine_<date>.json. Field-level fail-closed; every withheld item reported.
- 3
- document_qa.py v1 (md-only): reads the FINISHED md + summary + claims.json; extraction by TOKEN CLASS (every $ amount, %, rating token, ISO date, scenario-shaped string); each token resolves to a claim (within tol), a declared non-claim section context, or a finding. Checks: QA-501 rating-token reconciliation (BLOCKING; stance divergence requires the disclosure sentence), QA-502 spot-magnitude dollars + spot-relative % recompute (BLOCKING), QA-505 scenario-name identity incl. the 0%-from-failed-mapping pattern (BLOCKING), QA-506 cluster-vocabulary containment (BLOCKING), QA-507 anchors-count truthfulness (prose count == len(anchors.active)), QA-508 no past-dated event anywhere + next catalyst future-or-absent (BLOCKING), QA-509 registry byte-identity recompute, QA-510 whole-document rating scan replacing mch-html-qa's unreachable check. Generic unresolved percentages WARNING in T1, ratchet at T3. --estate resolves the week via CANONICAL_ESTATE.json; exit = blocking count; frozen vintages reported-not-restated (the QA-3xx convention).
- 4
- Loud joins: mch_weekly_run probs.get(n,0) -> implied=None + join_errors + an honest 'withheld — scenario map does not match' render; industry_consistency --check --all + exit codes + canonical-dir resolution, registered BLOCKING in ci_gate; validate_context_v2 scenario guard fixed to read configs/<T>.json scenarios (config missing => error, not skip); qa_audit --week default via CANONICAL_ESTATE.json.
- 5
- Layout: Company Overview relocated after the decision card (before Rating Bridge); past catalysts REMOVED ENTIRELY (mch_v3_modules catalysts_block days_out>=0 filter; IC Summary next-catalyst re-pointed at ros.catalyst_timeline.next_catalyst — one source with the strip; ros_catalysts window-bypass closed; qa413 gains its own docstring's date rule and becomes BLOCKING).
- 6
- Hardcoded 'five independent anchors' blurbs (template + glossary) become computed from the active anchor count; the strip gains the stance-reconciliation sentence.
- 7
- publish-ticker.mjs: summaryPath resolved from the engine's CANONICAL_ESTATE.json — the estate-wide silent QA-3xx skip becomes a loud refusal when a summary is absent — plus a per-name document_qa call; mch-html-qa.mjs retires the unreachable rating scan.
- intended side effects
- QG-14/Q7 flip GREEN->AMBER for ~243 all-past-catalyst names (AMBER does not block publish; disclosed in the checkpoint impact table); mch_market_view thesis driver recomputes from the filtered catalyst list (names whose only events were past stop being labelled event-driven — a semantic correction, counted); source-log catalog rows vanish where calendars empty.
- fixtures and mutations
- Frozen oln_broken pair (today's published OLN md + ROS5 summary) must FIRE QA-501/505/506/507/508/510; oln_fixed pair post-re-emit must pass; synthetic wrong-cluster + zero-join pairs. Mutations certified caught: claims_write_skipped, probs_default_restored, narrative_guard_neutered, anchor_blurb_rehardcoded.
- checkpoint
- ros_reemit --all -> document_qa --estate first-run triage (the finding list IS the remediation worklist) -> owner sign-off on the visible narrative-fallback changes -> publish --skip-gated -> held-name report + impact table.
- never this
- No silent defaults in any new join; no retro-minted vintages; --force never lifts a document_qa BLOCKING finding; frozen vintages reported-not-restated; the tranche ships NOTHING driving — every new check gates publication quality only.
AM-046 · 2026-08-14 — T5-D5 sizing liquidity truthing: days-to-exit computed from real 21-day average dollar volume, not one quote-day or an mcap guess
- owner instruction
- Marinus 2026-08-14: 'start the D1 selector bridge and then D5'.
- what
- ros_sizing._liquidity's adv_usd_m basis becomes, in order: (1) adv_usd_21 — the 21-day average daily dollar volume from the T1 OHLCV capture, split-adjusted, BY IMPORT from behaviour_features (single source with the behaviour layer); (2) the single-day quote volume x spot (the current first choice, demoted: one day is noise against a 21-day average); (3) the market-cap proxy (0.4%/day). Every basis is STAMPED in adv_basis exactly as today; a missing input falls DOWN the chain, never up. The mcap-keyed liquidity rating and every other sizing field are unchanged.
- published effect
- days_to_exit (and adv_usd_m) change where the 21-day average differs from the quote-day figure or the proxy. Impact on the current book measured at build and recorded in the commit; the estate republishes on the next scheduled run.
- mechanical gate
- The ADV21 basis is used ONLY while AM-046 is in this registry; absent it, the current quote-then-proxy chain runs byte-identical. Certified mutation sizing_adv_gate_bypassed.
- conservatism
- No OHLCV vintage (pre-capture cache) => honest fallback with the basis stamped — the same self-healing the T1 capture registered.
AM-045 · 2026-08-14 — AM-044's precedence rule contradicts its own purpose for v1 None entries (corrects AM-044)
- why appended not edited
- AM-044 is append-only and stands as written; this lands BEFORE any run in which the bridge resolves a match, caught during implementation review (the v1 map deliberately carries 'Iron Condor': None from AM-035's honest-null era — under blanket v1 precedence the one family the bridge names in its title could never bridge).
- what am 044 said
- 'The v1 map takes precedence where both define a key.'
- what it means
- NON-None v1 mappings take precedence (the rename aliases such as 'Covered Call' -> 'Covered Call (if held)' must not be shadowed by identity). A v1 entry of None FALLS THROUGH to the gated v2 tier, which resolves only keys in SETTLEABLE_STRUCTURES — so 'Iron Condor' bridges when AM-044 is registered, while 'Long Stock' and 'Cash-Secured Put' (not settleable options families) stay honestly None under every gate state.
- gate off invariant unchanged
- With AM-044 absent from the registry the v1 map alone resolves, byte-identical, including every None.
AM-044 · 2026-08-14 — T5-D1 selector bridge: the matcher covers all twelve v2-priced families (capability only — the nine decision rows stay byte-identical)
- owner instruction
- Marinus 2026-08-14: 'start the D1 selector bridge'.
- what
- ros_options_intel's exact-alias matcher gains a GATED v2 tier: identity aliases for every options_econ.SETTLEABLE_STRUCTURES family (Iron Condor, Long Straddle, Long Strangle, Put Spread (income), Call Ratio Backspread, Long Butterfly, plus the six the v1 map already reaches). The v1 map takes precedence where both define a key; matching stays EXACT (the AM-035 discipline — no fuzzy revival).
- what does not change
- 0
- The nine decision_table rows: directions, buckets, preferred, alternatives — byte-identical, asserted by test. No row emits a v2-only family today, so published matched-structure economics are unchanged by construction.
- 1
- Behaviour-conditioned rows remain blocked on live OOS evidence (the two-tier gate); IV-conditioned NEW rows remain blocked on two scored live quarters of the options v2 cohort (first scoreable 2026-10-31 window). Each future row change is its own owner-signed amendment.
- 2
- Ledger/cohort semantics: none. The options ledger keeps its own cohort discipline; the selector only matches and copies economics verbatim (QA-412 contract).
- 3
- QA-459's vocabulary (ALLOWED_STRATEGIES) is unchanged — Iron Condor was already in it.
- mechanical gate
- The v2 tier resolves ONLY while AM-044 is in this registry (amendment_registered clone); absent it, the v1 alias map alone resolves, byte-identical. Certified mutation selector_v2_unregistered: the v2 tier resolving without the registration must turn the suite red.
- effect on published output
- NONE today, by construction (no row emits a v2-only family). The bridge is structural readiness for evidence-gated table amendments.
AM-043 · 2026-08-14 — AM-015 completion: the optimal-expiry scoring function (registered before code)
- owner instruction
- Marinus 2026-08-14: 'start D2a' — the T5-D2a tranche of the MBE programme (AM-037 owner decisions), behaviour-independent by design.
- implements
- Exactly the mode AM-015 (2026-08-04) registered and never received: 'options_generate gains an optimal-expiry mode that scores across all liquid expiries in the chain rather than snapping to the nearest target, on theta profile, term structure, capital efficiency and catalyst coverage — explicitly NOT pick-the-longest.' The four dimensions are AM-015's verbatim; NO fifth dimension is added (expected-CAGR / vol-exposure sketch ideas are NOT registered and NOT built).
- candidate set
- 0
- Every listed expiry >= 7 DTE whose at-the-money contract of the required right passes the existing per-contract quality gate _q (bid floor, two-sided quote, OI floor, spread cap). The structure's own legs still face the full gate downstream.
- 1
- HARD exclusion preserved from the legacy picker: an expiry within 10 days BEFORE a known earnings print is excluded — scoring must not resurrect expiring-into-the-print.
- dimensions
- 0
- theta_profile = 1 - min(1, |theta_ATM| * min(H_hold, dte) / premium_ATM), H_hold = min(valuation-horizon days, dte): the premium fraction NOT burned by decay over the expected holding window. ATM proxy stamped (net structure theta when legs known).
- 1
- term_structure: cheapness of the candidate's ATM IV against the name's own interpolated ATM-IV curve from the SAME chain, rank-normalized, cheaper higher.
- 2
- capital_efficiency: rank of |delta_ATM| / premium_ATM — exposure per premium dollar.
- 3
- catalyst_coverage: (catalysts within the valuation horizon resolving BEFORE expiry) / (all such catalysts); 1.0 when none exist. Earnings keeps its hard rule regardless.
- composite
- score = 0.30*theta_profile + 0.25*term_structure + 0.25*capital_efficiency + 0.20*catalyst_coverage; each dimension min-max normalized 0-1 within the candidate set; TIE-BREAK: SHORTER dte (the registered anti-pick-the-longest backstop); top BOUNDS['expiries'] distinct expiries selected; detail output carries every candidate's per-dimension scores. Missing dimension input for a candidate scores 0 for that dimension (conservative, never renormalized away), stamped.
- bounds and cohorts
- options_distribution._scale_to_tenor r<=1.5 stands (beyond ~550 DTE the house MC is withheld; model-econ prices from term-structure IV alone per AM-015's known_limitation). LEAPS_MIN_DTE=365 blocking label check and the leaps tenor cohort (never pooled) stand.
- mechanical gate
- Scored mode ACTIVE iff AM-043 is in this registry (amendment_registered() clone, the AM-027 pattern); absent it the legacy nearest-target snap runs byte-identical. Certified mutations: expiry_gate_bypassed; expiry_longest_wins (a composite degenerating to DTE-monotone must be caught by a planted chain where the longest expiry is strictly worse on three of four dimensions).
- effect on published output
- Per AM-015: the options surface may shift expiries for liquid names at the next estate run; structures inherit all existing inception/cohort machinery. No equity rating, target or probability changes.
AM-042 · 2026-08-14 — AM-027/028 closure evidence: the correlated sleeve CVaR5 resolved the breach; the 2026-08-06 override is retired EARLY, before its 2026-09-06 expiry
- owner instruction
- Marinus 2026-08-14: 'retire the override and close it out'.
- evidence
- 0
- Grant condition (2026-08-06): absolute-sum sleeve CVaR5 1.514% vs the 1.5% NAV limit after five consecutive tightenings (07-29 1.425 → 08-06 1.514), driven by defined-risk put spreads whose full width counted −100% each; override ceiling 1.6%, expiry 09-06.
- 1
- AM-027 (registered 2026-08-07) made the correlation-aware aggregate the DRIVING measure; AM-028 (same day) corrected the copula loading's sigma source BEFORE the measure ever drove a published verdict.
- 2
- Production under the correlated driver: 2026-08-11 artifact — aggregate 0.076% NAV, driver 'correlated (AM-027)', 650/670 simulated, 20 comonotonic fallbacks NAMED, breach false. 2026-08-14 nightly (first unattended AWS run) — 0.198% NAV across 638 structures, headroom 1.302%. The comonotonic bound overstated sleeve tail risk ~7-8x, exactly as the override file's own 'why_the_number_is_probably_overstated' paragraph suspected.
- 3
- The 12-assertion sleeve-CVaR suite (register-then-run gate, conservative fallback, ordering, determinism, settlement contract at the extremes) runs in ci_gate.
- action
- data/options_budget_override.json DELETED. Rationale: (1) it is inert — there is no breach to excuse; (2) it was written for the ABSOLUTE-SUM artifact, which no longer drives — a future breach OF THE CORRELATED MEASURE would be a genuinely new event that must BLOCK, not be excused by a stale file written about a different measure; (3) its own review_at_expiry clause required choosing a resolution rather than renewal — resolution (b), the correlation-aware aggregation by amendment, was chosen 08-07 and demonstrably worked. Early retirement is the conservative direction: the limit fully re-closes today.
- unchanged
- monthly_cvar5_limit_pct_nav stays 1.5% NAV; the sleeve is not trimmed; the absolute sum remains computed and displayed as the conservative bound (AM-027's own terms); the override MACHINERY (reason+expiry+ceiling, breach-still-reported) stays in ros_options_budget for any future use.
AM-041 · 2026-08-14 — T6 — live OOS accrual + scoring for the behaviour shadow (registered before code)
- owner instruction
- Marinus 2026-08-14: 'start T6. Set a reminder for two quarters to include T4' — the T4 reminder is set (claude.ai routine trig_01V5UZuWHmUNUhHjEzYzoZp1, fires once 2027-02-14); calendar time alone never triggers a cutover, the reminder prompts the decision only.
- scope
- 0
- Score ONLY the two calibrated claims that passed their AM-039 gates: breakout_follow p_follow_63 and tail_risk p_mae_63. Retired models (containment, vol_persistence) emit no claims and are NEVER scored; their states are recorded facts, not forecasts.
- 1
- Cohort: behaviour_v2_live, inception 2026-08-14 (the first production estate attach). NEVER pooled with the anchor or options scoring cohorts, and never with the T3 backtest numbers — a live row and a backtest row must not meet in one statistic.
- outcome definitions
- 0
- IDENTICAL to the T3 harness, by import not by copy: direction join = behaviour_backtest.trend_outcome (the single-source function the suite asserts); sector benchmark = behaviour_backtest.SECTOR_ETF; breakout outcome = trend_outcome(63td forward sector-relative alpha, direction from the row's state).
- 1
- tail_risk outcome = (63td forward minimum cumulative return <= -mae_sigma_mult x RV63 x sqrt(63/252)) with mae_sigma_mult from the registered config (1.5). RV63 and the forward window are recomputed from the adjusted-close cache at score time using the row's as_of — retrospective computation of historical values, stamped in the scorecard.
- 2
- Maturity: both claims score at as_of + 63 trading days. Before maturity a row is scoreable:false with first_scoreable_date. A partial forward window (delisting, series end) is a REFUSAL counted with its reason — never a silent drop, never a partial return.
- metrics
- 0
- Per claim and per state: n, realized frequency, claimed p, Brier score.
- 1
- Skill vs climatology, where climatology = the accrued LIVE base rate of that state cell (never the backtest rate). A skill NUMBER is reported only when the cell has n >= 200; below that the scorecard says 'insufficient n' and shows counts only — no number invented from a thin cell.
- 2
- A reliability summary is added once total scored rows >= 2,000.
- 3
- No IR-like figure anywhere in the scorecard (house rule: any IR > 1.0 is a bug).
- machinery
- 0
- behaviour_score.py (new): reads output/performance/behaviour_ledger.jsonl (verifying the hash chain first — a scorecard from an unverified chain is refused), scores matured rows, writes output/behaviour_lab/live_scorecard.json. The ledger is NEVER mutated: scores live in the scorecard artifact only. Labeled 'live OOS shadow accrual — internal diagnostics — not performance evidence'.
- 1
- weekly_snapshots gains stage 9 'behaviour': chain verify + accrual telemetry (dated cross-sections present, freshest as_of) + the scoring pass + the dated scorecard write.
- 2
- Heartbeat: scheduler_heartbeat JOBS['snapshots'] min_work 4 -> 5. The 5th work unit is the DATED scorecard written by THIS run — deliberately NOT the ledger rows, because stage 2a0 runs before the attach at 2k in the same pipeline, so ledger rows visible at 2a0 predate the run and counting them would let a dead attach clear the gate (the VD-SCHED-001 class). The one pre-inception run (2026-08-14 nightly, ledger not yet born at 2a0) beats 4/5 short and is excused by the standing scoped snapshots override (expires 2026-08-18); every later run has the ledger and lands 5/5.
- 3
- QA/tests: tests/test_behaviour_score.py planted-world suite in ci_gate; certified mutation behaviour_cohort_pooled (the scorer pooling live rows with another cohort must turn a suite red); the scheduler wiring-audit updated to the new registry truth.
- cadence
- 0
- Quarterly evidence amendments: first due ~2026-11-14, second ~2027-02-14 — each records the scorecard verbatim at that date, pass/fail against nothing (accrual evidence, not a gate): the gates that matter are the cutover decisions.
- 1
- Every cutover (T4 display, any T5 driver) remains a separate owner-signed amendment in the AM-016 wording. Calendar time alone never triggers.
- 2
- T4 decision reminder: 2027-02-14 (two quarters), routine trig_01V5UZuWHmUNUhHjEzYzoZp1.
AM-040 · 2026-08-14 — Behaviour V2 verdicts — 2 PASS / 2 FAIL against the AM-039 gates; calibration adoption + claim retirements
- registered by
- session 2026-08-14, evidence from output/behaviour_lab/backtest_2026-08-14.json (841 estate + 120 delisted, 245 monthly dates 2006-01→2026-05, fitted on pre-2016, judged out-of-fold on the 10y tier, placebos N=200 per model)
- verdicts
- containment
- [object Object]
- breakout follow
- [object Object]
- tail risk
- [object Object]
- vol persistence
- [object Object]
- post verdict actions
- 0
- data/behaviour_calibration.json (new, version-controlled, value_ref-registered): train-fitted tables for the two passing models only, with provenance (train window pre-2016, artifact backtest_2026-08-14.json, this amendment). Shipped values are the train fits BECAUSE they replicated OOS; the 10y observed rates are recorded alongside as validation evidence, never as the fit.
- 1
- market_behaviour.py: containment and vol_persistence emit state-only descriptor blocks with claim_retired reasons naming this amendment; breakout_follow and tail_risk claims read the calibration table; the block gains calibration_version and drops uncalibrated:true (every emitted claim is now calibrated; retired models claim nothing).
- 2
- ros_config behaviour family: uncalibrated_probs_v2 marked superseded_by behaviour_calibration.json (kept for the record, never read for claims).
- 3
- QA-463 (v2 branch): claim probabilities must equal the calibration-table values.
- 4
- The 'behaviour flags' composite threshold (>= 2 models passing, AM-039) IS met (breakout_follow + tail_risk). Its introduction remains a SEPARATE owner-signed amendment — nothing composite ships under this one.
- 5
- Live shadow (T6) remains the real out-of-sample test for all carried numbers; in_sample:true and the no-backtest-marketing label stay on every artifact.
- epistemic note
- This round was adversarial against its own carried effects: one of the three effects carried from AM-038 (containment) died under a control that AM-039 added. Passes here are calibration-window results, not out-of-sample proof; the shadow ledger accrues the real test.
AM-039 · 2026-08-14 — Behaviour V2 — four narrowed models re-registered around the effects that survived T3, with the epistemics stated plainly: three are derived from the same window and their real test is live accrual; only the repaired volatility outcome faces genuinely fresh evidence
- owner decision
- Marinus Heymann, 2026-08-14: 'start the refinement round' (AM-038 reserved the refine-or-stop call to the owner).
- retired
- Directional trend (inverted in every regime), momentum-as-selection (lives in the factor library at portfolio-tilt strength; no second home), participation direction (wrong-signed). Retirement is permanent for these DEFINITIONS; any future revival is a new registration, not a revision.
- v2 models
- containment
- The v1 trend classifier's range state, promoted to a standalone claim: P(|fwd 63td return| <= RV63*sqrt(63/252) | state=range). v1 measured 68.0% (n 14,955). Gate: fitted OOF skill >= +2% AND reliability slope [0.8,1.2] AND range-state containment freq >= 62% with Wilson CI excluding the unconditional containment rate (computed in-run), n >= 10k.
- breakout follow
- The v1 structure model's passing half, alone: breakout states only, claim = 63td direction follow-through. The expansion claim is retired. Gate UNCHANGED from v1 (freq-0.5 >= +2pts with Wilson lower bound > 0.5, n >= 5k) plus fitted OOF skill >= 0.
- tail risk
- The v1 risk ranks unchanged; the CLAIM downgrades from strength to reliable ordering: strictly monotone MAE-event frequency across states (Kendall tau = 1.0) including the bear slice (unchanged), AND high-vs-low Wilson CIs non-overlapping — DISCLOSED: this replaces v1's ratio >= 1.8 bar, is weaker, and v1's measured 1.39 with its ns would already satisfy it. The model may therefore only ever be published as an ordering descriptor, never as a strength claim.
- vol persistence
- NEW outcome repairing the v1 degeneracy: per state, P(forward RV21 band == current band) — proper persistence, non-degenerate at every band including the boundaries. Gate: fitted OOF skill >= +5% AND reliability slope [0.8,1.2], n >= 50k. This is the only v2 model facing evidence not already inspected.
- epistemics
- SECOND LOOK, SAME WINDOW, SAID OUT LOUD: containment, breakout_follow and tail_risk were identified BY the 2006-2026 window they will be re-measured on — their v2 backtest passes are expected BY CONSTRUCTION and establish calibration tables plus implementation correctness, NOT independent evidence. Their genuine out-of-sample test is the live shadow accrual (T6) and its quarterly evidence amendments. Both looks are logged in trials.jsonl. vol_persistence alone carries a real ex-ante question into the run.
- stated expectation before running
- containment ~68%, breakout ~53%, tail ordering monotone — expected to pass (derived from this window; see epistemics). vol_persistence: genuinely open — band persistence should exist (vol clusters) but the repaired outcome has never been measured; a fail retires the model.
- composite
- NO composite in v2. If >= 2 models pass, a 'behaviour flags' block (named binary/tiered flags with their measured frequencies) may be proposed for T4 by a further amendment; a 0-100 score over four narrow claims would manufacture precision the parts do not carry.
- mechanics
- behaviour family reworked in place -> ros-1.17.0 (NO production block was ever emitted under 1.15/1.16 — the 2k stage ships with the unpushed local chain, and the production behaviour_ledger does not exist yet, so v2 replaces v1 before any row accrues). market_behaviour.py emits the four v2 models under model_set: 'v2'; behaviour_backtest.py gains the v2 outcomes; the same harness, universe, PIT rules, placebos and OOF discipline as AM-037/038. Verdicts by follow-up amendment.
- effect on published output
- NONE — shadow-only, structural no-op unchanged (behaviour absent from _ros_public).
- approved by
- Marinus Heymann, 2026-08-14 ('start the refinement round')
AM-038 · 2026-08-13 — Behaviour backtest verdicts: ALL SIX models FAIL their pre-registered gates — no composite is calibratable at the registered bars, publication does not proceed, and the three real effects inside the failures are named
- methodology resolution
- AM-037 left heuristic-vs-fitted scoring unstated. Resolved conservatively before judging: gates were judged on FITTED state->frequency tables scored OUT-OF-FOLD (fitted on pre-2016 rows, scored on the 10y tier). Scoring the heuristic placeholders would judge the placeholder; scoring a same-sample fit would be tautological. Heuristic-scored figures are recorded alongside in the artifact.
- run
- 2026-08-13 · 841 estate + 120 delisted (longest-span, delisted >= 2008) · 245 monthly cross-sections 2006-01-31 -> 2026-05-29 · outcomes per AM-037 verbatim · placebo N=200 p95 <= 0 for every model (the machinery does not hallucinate skill) · artifact output/behaviour_lab/backtest_2026-08-13.json
- verdicts
- trend
- FAIL — the DIRECTIONAL claim is inverted at scale: AUC 0.4475 on 101,719 rows, and below 0.46 in EVERY regime slice (bull 0.456, bear 0.415, gfc 0.410, covid 0.419, rates 0.422). Directional states carry no edge (state frequencies 0.49-0.50). The one real signal: RANGE names stay contained 68.0% (n 14,955) — containment, not direction.
- momentum
- FAIL by both clauses, narrowly and honestly: median monthly IC 0.0198 vs the >= 0.02 gate (the platform's own live IC record reproduced almost exactly), 58.1% months positive (passes that clause), pooled AUC 0.5024 vs >= 0.53. A monthly IC that does not survive pooling is not a per-name selection signal at this horizon.
- participation
- FAIL — the distribution leg is WRONG-SIGNED on raw returns: P(up|accumulation) 0.587 vs P(up|distribution) 0.618 — distribution names rose MORE. Measured lift -3.1pts vs the registered +3pts. The AUC 0.591 is partly an artifact of pooling oppositely-coded outcomes and is not credited.
- volatility
- FAIL — reliability slope 1.298 vs [0.8,1.2] (skill +0.055 passed its clause) — AND the registered outcome is DEGENERATE AT THE BOUNDARY: 'same or lower band' is tautologically true for the top (stressed) band — 12,700/12,700 hits by construction. The monotone frequency ladder (0.57/0.75/0.82/1.00) is partly baked in by the definition, not evidence. The outcome definition itself is the defect; any retry requires a NEW registered definition.
- structure
- FAIL on the compound gate — but the follow-through half PASSED alone: breakout follow-through 53.21% at n 24,426 (lift +3.2pts, Wilson lower bound above 0.5). The expansion half failed (AUC 0.5172 vs 0.53). Breakout follow-through is a real, small effect.
- risk
- FAIL — MAE-event frequency IS strictly monotone across states (0.0883/0.0935/0.1009/0.1230), monotone within the bear slice too, but the high/low ratio is 1.39 vs the registered 1.8. Real separation, below the registered bar.
- consequences
- Per AM-037 discipline: a failing model is excluded from the composite; with ALL SIX failing there is NO calibratable composite and NO behaviour_calibration.json ships. The T4 publication gate (owner decision 2026-08-13: historical validation gates display) is NOT met — the Market Behaviour section does not publish. The T2 shadow continues accruing, stamped uncalibrated:true and authoritative:false, harming nothing. Every T5 driving integration that depends on behaviour states stays unbuilt.
- real effects for any refinement round
- Three narrow effects survived honest measurement and one reproduction: (1) range-state containment 68% — a range-detection claim, not a direction claim; (2) breakout follow-through +3.2pts at 24k events; (3) monotone MAE separation across risk states (ratio 1.39); plus momentum's known weak monthly IC reproduced at 0.0198. A refinement round would re-register redefined models around these (and a non-degenerate volatility outcome) via a new amendment and a fresh T3 — an owner decision, not a default.
- stated expectation check
- AM-037 predicted trend/momentum AUC 0.52-0.56 (WRONG for trend — inverted; momentum at the bottom edge pooled), volatility strong on calibration (partly an artifact of the degenerate outcome), participation/structure genuinely uncertain (participation failed wrong-signed; structure half-passed), risk monotone (CORRECT, below ratio bar). Recorded against the predictions to keep the register honest about its own foresight.
- published numbers changed
- NONE. Nothing was published from this tranche; that was the design.
- approved by
- recorded by the build session 2026-08-13; the refinement-or-stop decision is explicitly reserved to Marinus Heymann
AM-037 · 2026-08-13 — Market Behaviour Engine — six behaviour models, a composite and an Entry Quality score, computed, recorded and scored in shadow; nothing published, nothing driven (programme registration: tranches T1-T3)
- config
- ros-1.15.0 on first emit · new `behaviour` family: model weights exactly the owner's spec (trend .25 / momentum .20 / participation .20 / volatility .15 / structure .10 / risk .10, registered PROVISIONAL), EQ weights .45 valuation / .35 MBS / .20 liquidity, BEHAVIOUR_MISSING_SCORE 40, all state thresholds — every constant carries an assumption-register entry with a live value_ref (behaviour.* namespace).
- what
- behaviour_features.py (pure, PIT-refusing feature functions over OHLCV; the volume/OHLC capture is an ADDITIVE key on the existing daily_adj cached parse — the splits precedent, zero extra API calls) and market_behaviour.py emitting per-name states with probabilities and confidence for trend, momentum, participation, volatility, structure and risk, a 0-100 composite (MBS) and an Entry Quality score, to a block stamped authoritative:false + advisory:true and an append-only hash-chained output/performance/behaviour_ledger.jsonl (MAX_AGE_DAYS=10 no-backfill, --verify). Probabilities stamped uncalibrated:true until the calibration harness (T3) replaces the registered heuristic mapping with measured tables (behaviour_calibration.json, version-controlled).
- never this
- No raw retail indicator is ever published as advice — no 'RSI says BUY'. States, probabilities, confidence. Behaviour NEVER drives rating, target or scenario probabilities (refused permanently; QA-461 enforces).
- backtest registration
- scope
- Walk-forward monthly cross-sections 2006-01 to present (5y/10y/20y tiers) on PRICE/VOLUME signals only. IV-touching components accrue live only — 25 months of IV history exist and a 20-year IV backtest is refused as impossible. Fundamentals-conditioned features refused beyond the 2026-07-28 PIT boundary (none exist by construction).
- universe
- estate 858 + delisted extension capped at 120 (options_backtest_pit.listing_map pattern); survivor-only AND extended reported side by side; gates judged on the EXTENDED set; delisting truncations counted as per-state refusals and the bias quantified.
- outcomes
- trend p_persist_63/p_contained_63 · momentum p_pos_alpha_63 · participation p_confirm_63 · volatility p_calm_21 · structure p_follow_63/p_expand_63 · risk p_mae_63 — all machine-checkable, all >= the 21-trading-day house minimum; outcomes joined via alpha_common.fwd_return/fwd_alpha_sec (refuse partial windows).
- keep drop gates 63td 10y extended
- [object Object]
- discipline
- A failing model is EXCLUDED from the composite and the exclusion recorded; MBS may ship with fewer than six models and says so. Composite weights are NOT fitted to the backtest — one pre-registered 20%-holdout re-fit clause only, exercised by its own amendment, else refused. Placebos (N=200, date-scramble + sign-flip) must show ~0 skill; every variant logged to trials.jsonl; all artifacts stamped 'calibration backtest — in-sample — no backtest marketing'; any IR-like figure > 1.0 anywhere is treated as an implementation bug.
- evidence gate
- Per-model keep/drop verdicts against the gates above, recorded by follow-up amendment with the numbers. PUBLICATION (display-only, advisory-stamped) is gated on those verdicts plus owner sign-off; every DRIVING integration (selector, expiry scoring with behaviour inputs, optimizer preference, conviction component) is a separate owner-signed amendment with its own mechanical amendment_registered() gate and certified mutation.
- stated expectation before running
- Trend and momentum: weak-positive discrimination (AUC 0.52-0.56), consistent with the platform's own momentum IC record. Volatility: strong calibration on persistence (vol clusters) — that is the bar it must clear, not discrimination. Participation and structure: genuinely uncertain — either may fail its gate and be dropped. Risk: monotone MAE separation expected. Anything dramatic is treated as a leakage bug first. Recorded before any backtest runs to prevent post-hoc rationalisation.
- fallback
- Missing inputs are conservative by construction: a missing model contributes 40/100 (below neutral) and is logged with a reason; more than two missing withholds the composite; a missing valuation component withholds Entry Quality entirely — behaviour never stands without valuation. Names without volume (pre-capture cache vintages) run with participation missing, which is the honest state of the estate until the capture completes.
- effect on published output
- NONE in T1-T3. The estate run that introduces the shadow block must produce byte-identical public JSON and reports — that no-op is the acceptance test (AM-016 precedent). `behaviour` stays out of the _ros_public allow-list until the T4 amendment.
- approved by
- Marinus Heymann, 2026-08-13 (plan approval: Market Behaviour Engine programme, ExitPlanMode; 'start T1')
AM-036 · 2026-08-13 — AM-035 evidence: the frozen-book impact table, measured and signed — and the R5 rule was structurally DEAD, not merely biased (corrects AM-035)
- evidence
- measured on
- frozen weekly_ROS5, 858 summaries, read-only (ops/am035_impact.py, 2026-08-13)
- conviction
- 846/858 moved · mean +1.73 pts · range [-3.80, +5.20] · 581 with |delta| >= 2 pts — inside the pre-stated 10-point component bound
- r5
- 49 names newly eligible (real spot/SMA200 as low as 0.663), 0 reverse; the synthetic constant 1.415 > 0.85 means R5 had never fired for any name since the rule shipped
- stance
- 5 changes after the full rules engine with hysteresis, ALL Increase->Hold: ACI, APP, BLDR, CPRI, TSLA — names genuinely below their real 200d average that the dead rule never flagged
- selector
- 91 matched-structure changes: 58 Bear Put Spread -> Protective Put (if held) family corrections, 19 Put Spread (income) -> None and 14 Long Call (LEAPS) -> None honest nulls
- withheld
- 3 names with no usable cached series (EXPD, HONA, SPCX) now say so instead of fabricating
- verdict
- AM-035's evidence gate CLEARED — the correction is accepted on truthfulness; conviction stayed inside its pre-stated bound, the stance blast radius is 5 names in the conservative direction, and every selector change is either the correct family or an honest null.
- transition rule
- QA-457/459 carry a vintage gate: summaries with no technicals.source field are pre-AM-035 history — reported, never blocking (the frozen-vintage precedent). The gate self-retires: after the corrected estate publishes, no canonical summary lacks the field and the checks bite in full.
- published numbers changed
- On the next estate run: conviction scores (846 names, <=5.2 pts), rules stance on the 5 named, options_intel matches on 91. Ratings, targets and authored scenario probabilities untouched (QA-405/473 re-verified).
- approved by
- Marinus Heymann, 2026-08-13 ('sign-off, publish the correction' — impact table reviewed)
AM-035 · 2026-08-13 — Real technicals, canonical sectors, and an exact-match selector — three measurement defects corrected before the Market Behaviour programme stands on them (AM-034 remains reserved for the Phase-8 event queue, drafted and unregistered)
- what 1 technicals
- The estate pipeline never passed price_history, so every published technical came from a fixed-seed random walk anchored at 0.3×spot (mch_stock_engine.run_technical_analysis) — 858 names printed RSI 39.5685 and SMA-200 off real values by 20-30%. The block is now computed from the cached 27-year adjusted-close series (858/858 coverage) and stamps source + as_of. The synthetic branch is DELETED, not gated: a fallback that fabricates is the defect class itself.
- what 2 sectors
- ros_factors sector-demeaning and alpha_common.sector_map read canonical GICS from data/taxonomy.json (890/892 names) instead of a fallback chain that resolved 404/858; the residual fallback share is logged per run.
- what 3 selector
- ros_options_intel._match_structure becomes exact-match against a registered alias map — the token matcher was mapping 'Protective Put' to 'Bear Put Spread' on 58 live names and 'Long Stock' to 'Long Call (LEAPS)' on 14; unmatched names keep the honest matched_structure null with a reason. The withdrawn PMCC (AM-018) is removed from qa_audit.ALLOWED_STRATEGIES and from both decision_table alternatives lists. The stale G15 documentation claim that Iron Condor and true LEAPS are unpriced is corrected (AM-017 prices both on the v2 path). The selector is NOT re-pointed at the v2 families here — that is a driving change and waits for its own amendment.
- why
- The Market Behaviour Engine (next amendment) must stand on real measurements. Publishing corrected numbers first keeps its introducing no-op acceptance test meaningful: a shadow layer proven against fabricated inputs would prove nothing.
- evidence gate
- A frozen-book impact table BEFORE the publishing run: the conviction delta distribution, every R5-technical-breakdown eligibility flip, every stance change, every matched-structure change, and factor-percentile movement counts. The correction is accepted on truthfulness, not on outcome direction; the measured table is appended by same-run follow-up amendment (AM-032 pattern) and the publishing run is gated on the owner's sign-off of that table.
- stated expectation before running
- Conviction moves bounded by the 0.10 component weight (≤10 points absolute). R5 eligibility will change materially in BOTH directions: the synthetic SMA-200 was biased low by construction (series anchored at 0.3×spot), so spot/SMA200 ratios were overstated and R5 under-fired. Matched structures change on exactly the 72 mis-mapped names; the 188 nulls gain no economics. Sector-demeaned factor percentiles move for roughly the 450 newly-labelled names. Recorded before the run to prevent post-hoc rationalisation.
- effect on published output
- Conviction scores; rules stance where R5 flips; options_intel matched structures and their copied economics; factor percentiles. Ratings, targets and authored scenario probabilities are untouched (QA-405/473 firewalls unchanged and re-verified).
- config
- ros-1.14.0 — decision_table alternatives (PMCC removal) + the selector alias map as a registered structure.
- approved by
- Marinus Heymann, 2026-08-13 (plan approval: Market Behaviour Engine programme, Tranche 0 owner decision 'Fix first, in this programme'; the publishing run itself additionally gated on impact-table sign-off)
AM-033 · 2026-08-11 — Earnings-revision activation attempt REJECTED by its own evidence; estate pairs gated behind a METHOD boundary; the fy[0] bug's fourth surface fixed (corrects AM-032)
- activation attempt
- First estate-to-estate pairs (2026-08-07 -> 2026-08-11, 823 computable rows) produced Sigma carve +4.14 and the signed estate residual fell +5.23 -> +1.11 — but the row-level evidence AM-032 required said NO: 600/848 rows nonzero over TWO trading days (implausible as revision breadth), the carve reduced |residual| on 294 rows vs increased on 306, Sigma|residual| on carved rows was 14.1% WORSE, and shrugged rose 13.6% -> 17.4%. The signed improvement was cancellation, not explanation.
- verified cause
- The pair spanned a METHOD change inside the source. mch_v3_modules built consensus_fy_eps by first-'fiscal year'-row selection — the fy[0] index bug's FOURTH consumer surface — so the fiscal year the field quoted moved with cache generations (verified: SYNA 08-07 vintage quoted FY2027 5.27, 08-11 vintage FY2028 6.56; the top 'revisions' replicated the 08-10 cross-source basis-gap table almost exactly).
- changes
- 1) mch_v3_modules now selects FY+1 BY LABEL (earliest fy_end >= today; latest when all past; rows without fy_end never selected) — all four surfaces now share the rule. 2) attribution EST_METHOD_BOUNDARY = 2026-08-12: estate rows dated before the boundary never serve as either pair leg (mutation phase9_est_boundary_dropped; suite proves the filter is the live mechanism by moving the boundary and watching the same pair compute).
- status
- Category returns to REGISTERED-INERT (0 computable rows). It re-activates when two post-boundary estate vintages exist (first candidate pair: the next two estate builds), measured against the same row-level standard before it may stand.
- ledger note
- The 2026-08-11 attribution ledger row appended by the activation attempt carries the method-artifact revision component (Sigma +4.14). Append-only: it stands, flagged here, superseded by the same-day corrected row (revision 0.0000).
- published numbers changed
- none — attribution artifacts have no rating/target consumer; consensus_fy_eps display fields refresh estate-wide at the next engine run
AM-032 · 2026-08-10 — Tranche-2 evidence recorded; earnings-revision pairs restricted to SAME-SOURCE after a pre-activation measurement caught a basis gap (corrects AM-031)
- evidence noise band
- before
- 2026-08-10 pre-AM-031 run: 848 attributed, shrugged 72%, REFUSED
- after
- 2026-08-10 same book, AM-031 in force: 848 attributed, shrugged 13.6% (115 rows), within_noise_band 496 rows, EMITS; median annualized sigma_idio 0.3252 (sanity [5%,200%] clear), vol-uncomputable 0.0%
- split
- 100% of the un-shrugging came from the noise band; the revision carve contributed 0 rows (see correction below) — matching AM-031's stated_expectation_before_running
- reading
- 13.6% sits at the ~13-18% tail a 1.5-sigma band leaves by chance under fat tails — the criterion now flags genuine outliers (PLTR +29.5% idio vs a ±9.1% band in a reporting week) instead of drowning in quantified noise
- verdict
- AM-029's evidence gate CLEARED for the noise-band classification — it stays
- correction same source
- what am 031 said
- srcs estate/av pooled (same measurand: street FY1 consensus)
- what measurement showed
- Cross-source E0/E1 pairs (estate consensus field at 2026-08-07 vs AV EARNINGS_ESTIMATES at 2026-08-10) measured BEFORE activation: 71% of 810 pairs nonzero, |delta| p50 0.23% / p75 1.40% / p95 6.75%, 19% beyond 2% — over ONE trading day, in which true consensus-revision incidence is a few percent of names. That is a basis gap (fiscal-year row alignment, adjusted-vs-GAAP estimate bases), not revision activity; carving it would have injected fiction into 500+ residuals
- rule
- E0 and E1 must come from the SAME src; a name with an in-tolerance E0 but no same-source later snapshot is not-computable with the reason named. Certified by mutation phase9_cross_source_pooled + a planted cross-source world in the suite
- consequence today
- 0 rows computable on the current book (the only in-tolerance E0 is the 2026-08-07 estate snapshot; no later estate snapshot exists yet) — the category is REGISTERED-INERT and contributes an honest 0 estate-wide
- activation path
- estate→estate pairs begin at the next estate build; the category's own evidence (does the carve reduce shrugged share on rows where it computes?) is measured then. Its failure mode is contribute-0, never fiction: unmeasured implies 0 by construction
- rail keepalive fix
- The consensus rail's 24-day dormancy (2026-07-17→08-10) was a wiring defect: weekly_snapshots stage 1 invoked alpha_consensus_snapshot with NO arguments — the report-only branch, appending nothing, while recording ok=True. The stage now runs --from-estate against the canonical estate (free, dedup-safe, rows dated by analysis_date); certified by mutation phase9_rail_starved + a structural suite check. Wiring-vs-function, occurrence #7 on this platform
- first learning pass
- First non-refused attribution artifact → learning_engine appended 4 proposals (expectation-bias: Energy n=37 mean err -0.029, Information Technology n=106 +0.062, Materials n=52 +0.041; kill-switch telemetry 9/33 observable triggers fired). Propose-only; each states that adoption requires its own amendment
- published numbers changed
- none — no published surface consumes attribution artifacts
AM-031 · 2026-08-10 — Attribution tranche 2: noise-band classification + earnings-revision as the first independently MEASURED category (AM-029 expansion clause exercised)
- motivating measurement
- First live attribution run (2026-08-10) REFUSED: 848 attributed, 72% of rows residual-dominant against the 25% cap. Diagnosis: at ~1-week elapsed windows the dominance criterion asks the impossible — single-name short-horizon error is mostly idiosyncratic noise, and market+sector explain a minority of weekly single-name variance as a fact of equity markets, not a modeling failure. No category list can attribute noise; the criterion must distinguish QUANTIFIED noise from genuine attribution gaps.
- change 1 noise band
- definition
- sigma_idio from trailing 120 trading days of daily log returns ending at the last trading day <= d0 (point-in-time at registration; no look-ahead); >=60 observations required; sigma_idio^2 = max(sigma_d^2 - beta^2*sigma_spy_d^2, (0.25*sigma_d)^2) — the floor keeps the band from collapsing when beta^2*sigma_spy^2 >= sigma_d^2, and a SMALLER band is the conservative direction (more rows shrug); band = 1.5 * sigma_idio_d * sqrt(N_trading_days(d0 -> as_of)) counted from the name's own close series
- classification
- a row that shrugs under the dominance rule but whose |residual after measured carves| <= band is classified within_noise_band — explained-as-noise, with the band, sigma, and day count emitted per row; it is NOT counted toward the shrug cap
- missing input rule
- insufficient closes for sigma (or no beta) => the row CANNOT claim within-noise and stays shrugged when dominance fires — missing input, more conservative
- change 2 earnings revision
- measure
- earnings_revision_surprise = E1/E0 - 1 (signed, unclamped), entering the component sum; E = street consensus FY1 EPS from the alpha-P2 consensus rail (output/alpha_lab/consensus_history.jsonl, srcs estate/av)
- E0 join
- snapshot nearest d0 within the ASYMMETRIC tolerance [d0-3d, d0+5d]: a later-than-d0 E0 UNDERSTATES the carve (conservative); an earlier snapshot beyond 3 days would attribute pre-window revisions into the window (anti-conservative), hence the tight early side
- E1 join
- latest snapshot dated <= as_of and strictly after the E0 snapshot date
- not computable
- either snapshot missing, |E1/E0 - 1| > 0.30 (fiscal-year-rollover signature, reusing the rail's own guard), or |E0| < 0.05 => the component is EXACTLY 0 and the residual keeps the full burden — when unmeasured, the category claims nothing
- change 3 run guards
- existing 25% shrug cap UNCHANGED; NEW BLOCKING refusals: (a) vol-uncomputable share > 20% of attributable rows (the measurement layer is broken, the band would silently degrade to the old criterion estate-wide), (b) median annualized sigma_idio outside [5%, 200%] (degenerate vol inputs make the band a carpet or a wall). Both to be mutation-certified.
- change 4 rails and capture
- consensus rail restarted (dormant since 2026-07-17): estate snapshot from the canonical week appended (dated by each summary's analysis_date, the rail's existing convention) + fresh AV snapshot dated 2026-08-10; forecast-ledger rows gain guided_eps + pe_mean (config-sourced, None when absent) at schema_version 2 — forward-only, NO backfill of existing rows; hash-chain semantics unchanged
- evidence gate
- Per AM-029's expansion clause, the categories STAY only if the same-day live re-run measurably reduces the shrugged share. Before: 72% (run of 2026-08-10). The after number, split into rows un-shrugged by the noise band vs by the revision carve, is recorded by follow-up amendment (AM-032) the same day. If the re-run still refuses, that is recorded too and activation is not forced.
- stated expectation before running
- at ~1 week elapsed the revision carve is expected near-zero for most names (consensus moves at earnings reports); the noise band is expected to do the bulk of the un-shrugging. Recorded before the run to prevent post-hoc rationalization.
- published numbers changed
- none — no published surface consumes attribution artifacts yet
AM-030 · 2026-08-07 — AM-029's in-module identity check was tautological; the binding refusal is residual dominance (corrects AM-029)
- what am 029 said
- the identity error == sum(components) + unattributed must hold to 1e-9 on EVERY row ... BLOCKING in the emitting module
- what is true
- With the initial category set, unattributed is DEFINED as the remainder, so the in-module assert could never fire (caught by the mutation harness within the hour: removing the assert survived every suite — certified by nothing). The identity is real but is enforced by construction plus the test suite's independent recomputation from planted series, not by a runtime assert. The category that closes the identity is renamed idio_residual and disclosed as a RESIDUAL; unattributed activates only when independently MEASURED categories (multiple-change, earnings-revision, ...) land per AM-029's expansion clause.
- binding check restated
- The module's BLOCKING refusal is RESIDUAL DOMINANCE: a row shrugs when |idio_residual| > 60% of |error|, and the run refuses to append when more than a quarter of rows shrug — i.e. when the MEASURED categories (market, sector) explain too little of the book's error for the decomposition to mean anything. This is computable and falsifiable on real data today.
- published numbers changed
- none — no consumer has read an attribution artifact yet; the first ledger append postdates this correction.
AM-029 · 2026-08-07 — Phase 9 (A8): attribution engine, thesis/kill-switch coverage, and the propose-only learning engine — methodology registered before first emission
- attribution
- INTERIM attribution of every standing forecast (earliest forecast-ledger row per ticker, the scoreboard's dedup convention) at each run date: realized return from adjusted closes (dividends folded, stamped) vs the forecast's linearly PRORATED expected return (upside_pct x elapsed/horizon — an approximation, stamped, replaced by the quantile-grid time-slice when the maturation machinery earns it). The error decomposes into named categories: market_surprise (shrunk beta x [SPY realized - ERP-prorated expectation, ERP 4.5% locked]), sector_excess_surprise (GICS ETF vs SPY), idiosyncratic_surprise, and unattributed. FIVE categories now; the register's ~9 (multiple-change, earnings-revision, scenario-shift, fx, dividend-explicit) attach as their inputs accrue, and any expansion toward the spec's 24 requires evidence that the new category reduces unattributed share, recorded here by amendment.
- reconciliation blocking
- The identity error == sum(components) + unattributed must hold to 1e-9 on EVERY row (catches drift between independently computed components), and the run REFUSES to emit when unattributed exceeds 60% of |error| on more than a quarter of rows — a decomposition that mostly shrugs is not an attribution. Both checks are BLOCKING in the emitting module, not advisory.
- kill switch coverage
- Per name: authored falsification triggers counted as observable-now (price-expressible) vs not-yet-observable, and observable ones evaluated against adjusted closes. This is COVERAGE and firing, not yet accuracy — accuracy claims wait for matured outcomes and are never backfilled.
- learning engine
- learning_engine.py reads the attribution + calibration artifacts and APPENDS dated proposals to output/performance/model_change_proposals.jsonl. PROPOSE-ONLY IS STRUCTURAL: the module's only write target is its proposals file; it imports no config writer; the constraint is asserted on the AST in tests and re-planted as a harness mutation. Auto-deploy of proposals was refused permanently in the Phase 0 decisions memo and that refusal stands.
- interim semantics
- Every emitted figure is stamped interim: true (mark-to-date, horizon open). Matured attribution lands only through scoring_core's maturation path; interim and matured are never pooled.
- deferred in phase
- Governance portal (hub surface) and calendars/diagonals schema v3 ship as their own tranche; the WS2 shadow-scenario cutover evaluation is NOT due before ~2027-02 (two scored quarters from the 2026-08-04 shadow inception) and its gate is unchanged.
AM-028 · 2026-08-07 — AM-027's copula loading used the MC quantile grid's sigma — valuation uncertainty, not return volatility — which understates co-movement and flatters the tail (corrects AM-027)
- why appended not edited
- AM-027 is append-only and stands as written; this correction lands BEFORE any run in which the correlated measure drives a published breach verdict, caught by AM-027's own comonotonic check during build verification (median lambda 0.181 was implausibly low for large-cap equities; the one-factor implied pairwise correlation would have been ~3%).
- what am 027 said
- per-name loading lambda_i = clamp(beta_i * sigma_SPY / sigma_i, 0.05, 0.95) with sigma_i from the name's own 1001-point MC quantile grid
- what is true
- The grid's dispersion is 1-yr VALUATION uncertainty (multiple/margin/scenario risk, typically 40-70% wide) — not the return volatility that the one-factor identity corr(i,m) = beta_i * sigma_m / sigma_i requires. Both sigmas must live in the same return space: sigma_i is now the name's 1-yr realized volatility from its cached adjusted closes (the same series the shrunk beta was estimated on, so the estimator is internally consistent). The MC grid remains the MARGINAL distribution — fat tails preserved; only the coupling strength changes.
- conservatism unchanged
- Names with no cached closes (or under 120 observations, the MIN_BETA_OBS floor) fall back comonotonic and are named, as AM-027 specifies. The comonotonic check and the absolute sum stay published beside the driver.
- published numbers changed
- none before this correction — no published run has yet used the correlated measure as its driver.
AM-027 · 2026-08-07 — Options-sleeve monthly CVaR5: correlation-aware aggregation becomes the DRIVING measure; the absolute sum remains the displayed conservative bound
- owner decision
- Marinus Heymann, 2026-08-07 (recorded choice among the three resolutions the 2026-08-06 override file's review_at_expiry clause names). NOT a sleeve trim; NOT a limit change — monthly_cvar5_limit_pct_nav stays 1.5% NAV.
- what changes
- ros_options_budget's breach test binds on a JOINT sleeve CVaR5: one-factor Student-t copula (df=5, the house fat-tail standard; a common chi-square shock supplies tail dependence), per-name loading lambda_i = clamp(beta_shrunk_i * sigma_SPY / sigma_i, 0.05, 0.95) with sigma_i from the name's own 1001-point MC quantile grid; terminal prices by inverse transform on that grid, re-horizoned to each structure's own tenor by the canonical options_distribution._scale_to_tenor; settlement P&L by the canonical options_econ contract (parse_legs + entry_premium + intrinsic) as % of model_econ.capital_base; per-structure sqrt-time monthly scaling unchanged from the absolute-sum convention; 20,000 seeded draws (seed = sha256('AM-027:'+as_of)).
- what does not change
- The absolute-sum measure keeps being computed and PUBLISHED in the same document as measures.absolute_sum_pct_nav — it is the comonotonic upper bound and the continuity series (1.425/1.450/1.495/1.514/1.526 across 2026-07-29..08-07). The sleeve universe definition, the sleeve_nav_pct reference arithmetic, the override mechanics (reason+expiry+ceiling) and QA-418's role are unchanged.
- conservatism clauses
- A name whose legs, premium, quantile grid, beta, or SPY sigma cannot be read is NOT simulated: its ABSOLUTE-SUM contribution is added on top of the simulated CVaR (its comonotonic worst case) and it is named in fallback_names — a missing input makes the output more conservative, never less. The measure also emits its own comonotonic_check (all lambdas=1, shared shock), which must sit at or above the correlated figure.
- register then run
- The driver switch is gated IN CODE on this amendment's presence in this registry (ros_options_budget.amendment_registered). Without this entry the absolute sum drives and the correlated figure is diagnostic-only. This entry is appended 2026-08-07 BEFORE any run in which the correlated measure drives.
- why
- The absolute sum assumes all ~656 structures realize their 5% tails simultaneously. Five consecutive tightenings against the 1.5% limit were driven by defined-risk put spreads each counting its full width (-100% CVaR5): a mix artifact of the aggregation, not a change in the sleeve's risk posture. Fixing the measure by pre-registered amendment — rather than trimming positions into a measurement artifact or raising a limit while it binds — is the resolution the override file itself prescribes.
AM-026 · 2026-08-06 — AM-023's enumeration of the options_hist store accounts for 25 of the 26 dates it names (corrects AM-023)
- why appended not edited
- AM-023 is append-only and stands as written. This is the same treatment AM-018 gave AM-017, AM-022 gave AM-021, AM-024 gave AM-023 and AM-025 gave AM-024: an amendment is corrected by a later one, never rewritten, so the record shows what was believed when the decision was taken as well as what is true now.
- what am 023 said
- "26 distinct chain dates in that store, month-end from 2024-07-31 to 2026-06-30, plus a single 2026-08-03 date"
- what is true
- Measured 2026-08-06 over output/cache/options_hist (20,556 files, which AM-023 states correctly): 26 distinct dates = 24 month-end dates from 2024-07-31 to 2026-06-30, PLUS 2026-04-17 (a single MSFT file, not a month-end, left over from the chain-date verification on the premium key), PLUS 2026-08-03 (60 files, the registration-snapshot rescue). 24 + 1 + 1 = 26. AM-023's sentence names 24 + 1 = 25 and so omits 2026-04-17.
- load bearing claim unaffected
- AM-023's REFUSAL does not move. The claim it rests on is that exactly ONE options_hist date falls on or after the v1 book's first entry (overlay_as_of 2026-07-23), and that holds: re-measured, the only such date is 2026-08-03. 2026-04-17 is three months BEFORE the v1 inception, so it cannot narrow the overlap. Historical calibration of exit rules stays REFUSED, on sampling — one usable date is not a sample — and not on field availability, which AM-023 already corrected.
- evidence
- glob over output/cache/options_hist/*@*.json, dates parsed from the filename: 20,556 files, 26 distinct dates, 24 month-end-ish, non-month-end = [('2026-04-17', 1 file), ('2026-08-03', 60 files)], dates >= 2026-07-23 = ['2026-08-03'].
- published numbers changed
- none — this corrects a count inside an amendment, not an output.
AM-025 · 2026-08-06 — The FIRED arm reported zero unobserved days when it could not look, and two shipped summaries reported post-filter denominators (corrects AM-024)
- discipline
- AM-023 and AM-024 are NOT edited. This appends beside them, exactly as AM-024 appended beside AM-023 and AM-018 beside AM-017. The sentences corrected below must still be readable in AM-024 or the record is not a record.
- correction 1 the fired arm disclosure
- what AM 024 registered
- 'The exit day of a FIRED arm now carries unobserved_days_before_this_exit, because unobserved days before an exit make the recorded date an upper bound.'
- what the code did
- options_exit_policy.paired_outcomes() computed `before = [] if expected is None else [...]` and emitted len(before), so when expected_observation_days() REFUSED the window the disclosure read unobserved_days_before_this_exit: 0. Zero is the strongest possible coverage claim — every day before the exit was watched, so the exit date is exact — and it was being made by code that could not check a single day. Every other reader of `expected is None` in that function withholds and names the missing denominator; this branch alone converted the refusal into a number. Reproduced by re-chaining the live policy with entry_as_of 'not-a-date' on one subject plus one fired mark: the arm returned {'outcome': 'rule_exit', 'unobserved_days_before_this_exit': 0}.
- how it stayed green
- tests/test_exit_policy.py exercised only the well-formed window (asserting == 4). Neutering len(before) to 0 went red there while the unreadable-window path stayed green, and verify() does not check date parseability, so such a row passes the chain gate.
- the fix
- When the window cannot be established the count is None and a sibling field, unobserved_days_before_this_exit_unknown, states that the number is UNKNOWN and the recorded exit date is an upper bound of unknown tightness. 0 now means zero. Proven by planting a fired mark on a policy row with an unreadable date and asserting the field is not 0 and that the text names the unestablished window; the well-formed case additionally asserts the sibling is ABSENT, so the two cases cannot certify each other.
- correction 2 two post filter denominators in shipped output
- options mark v2.status
- Printed '{len(last)} trades · N priced · N rules evaluated' — a post-filter row count presented as the population, with the live book mentioned nowhere. Three mark rows written for one day against the 79-trade book printed '3 trades · 3 priced · 3 rules evaluated', which reads as a complete day and is a 4% one. This is the identical defect --limit's summary was fixed for in AM-024 defect 3, left standing in the sibling reporting function, and --status is a shipped CLI entry point that nothing in the suite called. It now prints '{covered} of {len(book)} live trade(s)', the count of live trades carrying NO row that day, and a PARTIAL warning naming the append-only consequence.
- options exit policy.verify
- The headline 'covered at {version}: n/m live' could be narrowed to a post-filter denominator with the suite fully green, because both assertions on that line were satisfied by numerator == denominator and the 1-of-4 population injection asserted only the failure sentence. The ratio is now asserted in the FAILING cases (1/4 and 3/4), which is where a narrowed denominator shows.
- gates certified that previously were not
- last chain as of priced filter
- Dropping the `priced` term left the suite green while permanently suspending every rule on any trade that had hit the look-ahead refusal — that path writes priced=False WITH a future chain_as_of, so the staleness reference jumps ahead of every chain that will ever arrive. Both halves of the distinction are now certified separately: a not-priced row is refused as the reference, and a stale-but-priced row IS the reference.
- paired outcomes no policy branch
- `elif not p:` could be replaced with `elif False:` with the suite green, because no live trade lacked a policy row in any fixture. Any trade appended to the ledger after the policy chain was written lands there and would have been misdiagnosed as a date problem. Now planted by removing one row from the chain.
- gates that went red only by CRASHING
- Four gates aborted the run with a traceback before the collected failures were printed, so the suite exited 1 having said nothing about which branch was disabled. The (ref or {}) discipline is now applied to every refusal accessor, the domain table's rule keys are read with .get(), and both paired_outcomes() call sites turn an exception into a named failure carrying the exception text.
- verification debt
- The uncertified checks NOT closed in this pass are written into defects_register/mch_defect_register.json under 'verification_debt', one entry each with the neuter that would prove it and why it was not closed. A named list of uncertified checks is worth more than a scramble of fixtures that certify each other.
- what stands from AM 024
- Everything else. The coverage-gap denominator, the calendar-day over-count and its withholding direction, verify()'s refusal on an absent or empty chain, the --limit refusal, the 6,816 structure count and the entry_premium_sign correction are unchanged and re-asserted on every run.
- status
- REGISTERED. Nothing has been marked and nothing has settled in the v2 cohort — options_trade_marks.jsonl and options_trade_resolved.jsonl do not exist, checked 2026-08-06 — so no rule-arm number exists to have been fitted to any of this.
AM-024 · 2026-08-05 — Three false sentences in AM-023, and three gates in the exit-policy code that could be deleted with the test suite green (corrects AM-023)
- discipline
- AM-023 is NOT edited. This entry is appended and names what it corrects, which is the same treatment AM-018 gave AM-017 and AM-022 gave AM-021. Every figure below was recomputed from the shipping files while writing this, and the ones that can drift are now recomputed by tests/test_exit_policy.py on every run rather than restated from memory.
- correction 1 the tenor ratio
- what AM 023 said
- dte_floor_21.what_it_does_not_assume: 'so 21 days is 47% of the shortest trade's remaining life and 6.6% of the longest's.'
- why that is false
- The shortest DTE at registration is 18, so 21 days is 117% of that trade's ENTIRE remaining life — the floor exceeds the life. No trade in the book yields 47%: the distinct DTEs are exactly 18 / 46 / 74 / 137 / 165 / 228 / 318 and 47% would require a 44.7-DTE trade. The figure understated the disparity about 2.5x precisely where the disparity matters.
- the corrected sentence
- 21 days EXCEEDS the shortest trade's remaining life (18 DTE — 117% of it, which is exactly why those 7 trades are out of domain), is 45.7% of the shortest IN-DOMAIN trade's life (46 DTE) and 6.6% of the longest's (318 DTE). The 6.6% and the median of 165 in AM-023 were both correct and stand.
- why it mattered
- The sentence exists to DISCLOSE the tenor disparity. Understating it concealed the one fact that makes the disclosure worth making — that for the shortest tenor the rule would fire before the trade was recorded — and so contradicted the domain table printed beside it, which withholds those same 7 trades for that same reason.
- where fixed
- options_exit_policy.py module docstring; asserted in tests/test_exit_policy.py section 6b, which recomputes 21/18, 21/46 and 21/318 from options_trade_ledger.jsonl and fails if the false ratio returns.
- correction 2 convention derived for
- what AM 023 said
- profit_target_50.assumes: '...the weaker justification is recorded on those rows as `convention_derived_for` rather than papered over with a second invented constant.'
- why that is false
- `convention_derived_for` is the identical constant 'short premium, ~45 DTE' on ALL 79 policy rows and on BOTH rules — read back off output/performance/options_exit_policy.jsonl 2026-08-05: 79/79 for profit_target_50 and 79/79 for dte_floor_21. It describes the CONVENTION the parameter came from, which is the same sentence whatever trade it is attached to, so it identifies none of the debit verticals and discriminates nothing. Comparing an in-domain debit row's profit_target_50 block against an in-domain credit row's, the only field that differs is max_profit.
- the corrected sentence
- The three debit verticals in the profit-target domain are WDAY Bear Put Spread, CRWD Bear Put Spread and TSCO Bull Call Spread. What identifies them on the registered rows is `entry_premium_sign`, which is 'debit' on exactly 3 of the 17 in-domain rows and 'credit' on the other 14; the AM-017 family_cohort is what keeps the two from being pooled when the outcomes are read. `convention_derived_for` records the provenance of the parameter and nothing about the trade.
- what was NOT done and why
- No flag was backfilled onto the 79 registered rows. The chain is append-only and hash-chained; the rows were written before this was found, and editing them to make an old sentence true is the failure the chain exists to prevent. The false sentence is corrected instead — here, and in the options_exit_policy.py docstring — and tests/test_exit_policy.py now asserts against the REGISTERED chain both that convention_derived_for is a constant (79/79 on each rule) and that entry_premium_sign picks out exactly WDAY, CRWD and TSCO. A further check asserts no registered row carries a flag invented after the fact.
- counts that stand
- 17 in domain = 14 net credit (13 Put Spread (income) + 1 Iron Condor) + 3 net debit (2 Bear Put Spread + 1 Bull Call Spread). Re-measured 2026-08-05; unchanged from AM-023.
- correction 3 the 6816 challenge examined and REJECTED
- the challenge
- An adversarial pass reported that AM-023's refusal_v1_marks_are_not_a_daily_series is wrong — that '6,816 structures marked on BOTH 2026-07-30 and 2026-07-31, 6,160 (90.4%) identical' counts ROWS as structures, that the distinct structures marked on both days is 2,485, and that the structure-wise identical share is 84.5%.
- what the recount found
- The challenge is WRONG and its replacement figures are not adopted. It keys a 'structure' on entry_hash alone. entry_hash identifies an overlay EMISSION, not a structure: on 2026-07-31 its 2,485 distinct values cover 6,816 rows because one emission carries up to three structures for the same ticker — A on that date carries Covered Call (if held), Put Spread (income) and Protective Collar (if held), all under entry_hash 0c2e3ba9. Reading 2,485 as a structure count merges three different bets into one, inside an argument about a ledger that cannot count bets.
- the key that is correct
- options_rec_mark.py's own de-duplication key, (entry_hash, structure) — used at lines 111, 116 and 152 to decide what has already been marked and what has already settled. Under it, measured 2026-08-05: 2026-07-31 carries 6,816 rows and 6,816 distinct structures (one row each); all 6,816 were also marked on 2026-07-30 (which carried 8,236 rows / 8,236 structures); 6,160 of the 6,816 (90.38%) carry an exit_value identical to the previous day's; and 6,754 of the 6,816 were priced off the same 2026-07-28 chain. AM-023's four counts are therefore CORRECT AS WRITTEN and stand unchanged.
- what changed anyway
- options_mark_v2.py now names the structure key it counted with, and records that keying on entry_hash alone gives 2,485 'structures' at 78.6% identical — so the recount cannot be re-run wrongly and re-litigated a third time. tests/test_exit_policy.py section 6c recomputes all five figures from options_rec_marks.jsonl (47,435 rows) on every run, so none of them can drift.
- why this is in the record
- A challenge to a registered figure is examined and its outcome recorded whichever way it goes. Accepting a 'correction' that is itself wrong would have put a false number into the amendment chain under the appearance of rigour, and the wrong number would then have been the one everything downstream cited.
- defect 1 the coverage gap branch counted only rows that EXIST
- what AM 023 claimed
- not_evaluated_is_not_not_fired: 'paired_outcomes() WITHHOLDS the rule arm entirely for any trade with such a gap — the rule did not fire is only an outcome if the rule was evaluated on every day it could have fired.'
- what the code did
- It counted mark ROWS whose `evaluated` was false. A day the writer never ran leaves no row at all, so it was invisible. Run against the live book with one evaluated mark planted 43 days after entry, the rule arm returned {'outcome': 'no_rule_fired', 'days_evaluated': 1} — 'did not fire' meaning 'did not look', which is the exact conflation the branch was built to prevent, surviving inside the branch built to prevent it. `blind_days_before_first_mark` was computed and gated nothing.
- how it stayed green
- The one test that named the branch sampled next(iter(pairs.values())), which is NOC — a trade withheld by the DIFFERENT, unmanaged-by-construction branch, already asserted three lines earlier. Replacing the whole branch with `elif False:` left the suite green.
- the fix
- options_exit_policy.expected_observation_days() computes the days on which the rule COULD have fired — the day after max(entry_as_of, policy_as_of) through min(as_of, expiry) — and the rule arm now reports an outcome only if every one of those days carries an EVALUATED mark row. Missing rows and present-but-unevaluated rows are counted and named separately, with distinct sentences, so no fixture can certify the wrong half. The exit day of a FIRED arm now carries unobserved_days_before_this_exit, because unobserved days before an exit make the recorded date an upper bound.
- what this costs and why it is kept
- The window is CALENDAR days. This module has no session calendar and will not invent one: 'skip Saturday and Sunday' is silently wrong for every market holiday, and a hand-written holiday list would be a constant fitted here with nothing behind it. While the chain cache advances only on session days, every trade therefore carries unobserved days and the rule arm stays WITHHELD. That is the truthful reading of the standard AM-023 registered, and the direction of the error is the safe one — the diagnostic is withheld and says how many days it could not see, rather than published on days nobody looked at. A trading calendar may be introduced later as its own registered change.
- proof
- tests/test_exit_policy.py section 5a plants four states on a trade that HAS a rule in domain (so no sibling branch can catch it): 0 of 41 days observed, 1 of 41, 41 of 41 with one stale, and 41 of 41 all evaluated. The first three are WITHHELD with the branch named in the failure text; the fourth returns no_rule_fired, which is what proves the gate is a check and not a constant.
- defect 2 verify returned 0 on an ABSENT chain
- what AM 023 claimed
- status: 'Policy REGISTERED (79 rows, hash-chained).'
- what the code did
- options_exit_policy.verify() printed 'exit-policy chain is empty (no policy registered yet)' and returned 0 when the file was missing, and printed 'chain: INTACT' on a chain holding 1 of the 79 rows — it never compared the policy population against live_trades(). Both new tests/ci_gate.py lines passed with output/performance/options_exit_policy.jsonl moved away: the gate that exists to prove the registration artifact exists passed with the artifact deleted.
- the fix
- verify() REFUSES (non-zero, with the reason) on an absent or empty chain, and asserts that every live trade carries a policy row at the current POLICY_VERSION, printing 'covered at exit-policy-1.0.0: 79/79 live'. tests/test_exit_policy.py proves the refusal by deleting the file and asserting on the text, and also verifies the REAL registered chain from disk, so deleting it now fails both gate lines instead of neither.
- also certified
- Three of verify()'s four integrity checks had no test that isolated them. prev_hash linkage — the only check that detects a deleted or reordered row — had none at all; the orphan test passed off the row_hash check because its injected row also broke row_hash and prev_hash; the duplicate (trade_id, policy_version) check had none. Each now has its own injection with the other hashes recomputed valid, and asserts the ABSENCE of the sibling failure texts as well as the presence of its own.
- defect 3 limit could narrow an append only inception day
- what the code did
- options_mark_v2.py passed argparse's --limit straight into mark(), three lines below a docstring claiming 'a caller cannot narrow the marked population by accident', and the summary reported the POST-filter count as the denominator: '--mark --limit 3' printed '3 trade(s) ... 3 priced' with no mention of the other 76.
- why it matters here specifically
- The v2 marks ledger is append-only and nothing has been marked yet, so a limited FIRST mark would write a permanently partial inception day — 3 of 79 trades observed on a day that can never be re-marked — and under the coverage fix above those 76 trades would be withheld from the rule arm for the life of the book.
- the fix
- --limit is REFUSED outside --dry-run, with the true book size in the refusal text, and the refusal happens before the ledger is opened. The summary line now prints the UNFILTERED denominator and accounts for every trade in the book ('N of 79 live trade(s) ... M already marked today ... K skipped (expired) ... L not reached (--limit)'), with a loud line if those counters do not sum to the book. The default population is now tested directly by calling mark() with no trades= against a temp ledger and asserting 79 rows.
- seven further gates certified
- Each of these could be neutered with the suite green and now has an injection asserting on its failure text: the no-entry-cash refusal (neutering it to cash = 0.0 computes P&L against a price nobody paid); the unreadable-date DTE refusal; the one-mark-per-trade-per-day dedupe (without it a second --mark appends 79 duplicate rows — the same defect AM-023 quantifies at 9.2% for v1); the partial-leg-parse refusal (a two-leg spread would have been priced as a naked short); index_chain's first-listing-wins; read_chain's no-mtime-TTL contract; the register-then-run refusal in the mark writer (no test had ever marked an unregistered trade); unevaluated()'s out-of-domain preservation (its test read a stale variable bound before the injection); the no-evaluator rule status; the withheld profit_fraction on rows with no denominator; and policy_index()'s latest-wins, which is the only thing that makes a superseding policy take effect.
- supersedes field added
- build_rows() now writes `supersedes` — {policy_version, policy_as_of, row_hash} of the row it replaces — onto any policy row for a trade that already carries one, so the succession AM-023's shape argument depends on is checkable against the chain rather than inferred from file order. None of the 79 registered rows carries it and that is correct: none of them supersedes anything. They are not edited.
- what stands from AM 023
- Everything not listed above. The two rules and their parameters, the domain and every named exclusion (17/62 and 72/7, 36 unbounded and 26 held-stock, the -3 days of room), the shape argument, the paired design, the two-day blind window, the settlement-unchanged statement, the earliest-evidence dates, the asymmetry disclosure, the historical-calibration refusal on SAMPLING, the no-API-cost statement and the four v1-mark counts are all unchanged and were re-verified against the shipping files while writing this.
- status
- Policy still REGISTERED at exit-policy-1.0.0, 79 rows, hash-chained, unedited (md5 4ec1ffcac908ed943e7dc82bb4fc5063 before and after this amendment). The first mark has STILL NOT been taken: options_trade_marks.jsonl and options_trade_resolved.jsonl do not exist, checked on disk 2026-08-05 after this work. Nothing here observes an outcome, so the window AM-023 describes is still open and these corrections are still pre-observation.
AM-023 · 2026-08-05 — v2 managed-exit policy — two rules, a named domain, and a rule-exit number that is a diagnostic and not the published outcome
- registered before any observation
- Nothing in the v2 cohort has been MARKED (options_trade_marks.jsonl did not exist) or SETTLED (options_trade_resolved.jsonl did not exist) when this was written — both checked on disk 2026-08-05. This is therefore the last moment an exit policy can attach to the existing 79 live trades with no possibility of having been fitted to an observed outcome, and after the first mark it never can be again. options_exit_policy.py, options_mark_v2.py and tests/test_exit_policy.py were written before this entry; NO mark has been taken, by deliberate instruction, so the owner sees these parameters before the window closes.
- shape
- The policy is a SEPARATE hash-chained row keyed to trade_id (output/performance/options_exit_policy.jsonl), never a field inside the id and never a field on the ledger row. Inside the id it would mint a second registration of the same legs at the same price and destroy the de-duplication content addressing exists for — v1 re-appended 1,420 duplicate structures out of 15,463 (9.2%, measured 2026-08-05). On the ledger row it would be worse because it fails silently: append() skips an id already in the chain, so re-registering under a different policy is a no-op, the second policy is dropped and the caller is told 'already recorded'. Both failures are proven by injection in tests/test_exit_policy.py rather than argued.
- rules
- profit target 50
- [object Object]
- dte floor 21
- [object Object]
- conventions not imported wholesale
- Both parameters come from a 45-DTE short-premium regime, and this book is neither: median DTE at registration is 165 and 56 of the 79 live trades are long-premium debits (23 net credit), all measured 2026-08-05 from options_trade_ledger.jsonl. The numbers are kept because inventing bespoke ones on a book with no observed outcomes would be fitting a rule to nothing; what is NOT kept is the pretence that the derivation transfers. Each rule states its assumption above and each states where the assumption does not hold.
- domain and named exclusions
- live trades
- 79
- profit target 50 in domain
- 17
- profit target 50 withheld
- 62
- withheld unbounded max profit
- 36 — 19 Long Strangle, 17 Long Call (LEAPS). The payoff has positive slope at infinity, so there is no maximum to take a fraction of. Withheld and named; never a capped, sampled or substituted denominator.
- withheld held stock position not in the record
- 26 — 16 Protective Put (if held), 7 Covered Call (if held), 3 Protective Collar (if held). This is AMBIGUITY, not absence, which is why it cannot be repaired by fetching anything. The ledger row carries the OPTION LEGS ONLY (options_trade.Trade.identity() excludes the share leg on purpose), so two different maxima can be computed from the same row and they disagree. Measured 2026-08-05 with the registration-date closes from options_registration_snapshot_2026-08-03.json: NOC Protective Put 530.15 option-legs-only vs UNBOUNDED as a position; STT Protective Collar 169.95 vs 11.28 (15.1x); DHR Covered Call 8.45 vs 10.95 (1.30x). For 16 of the 26 the two readings do not agree on whether a maximum exists at all. options_payoff.structure_legs refuses to default `include_held_stock` for exactly this reason. Picking either reading silently would make the rule measure a position nobody was shown — the option-only reading of a protective put maxes out when the underlying goes to zero, which is a hedge paying off, not a trade working.
- dte floor 21 in domain
- 72
- withheld would have fired at registration
- 7 — the 2026-08-21 expiries were 18 DTE at as_of 2026-08-03 (5 Long Call (LEAPS), 1 Covered Call (if held), 1 Protective Put (if held)). A rule that exits a trade before the ledger records it is not a policy, so these are named OUT OF DOMAIN rather than exited at t=0, which would have manufactured 7 same-day round trips. The room each had is recorded as -3 days, not clamped to zero.
- trades with no rule in domain
- 7 (the same 2026-08-21 group). They are unmanaged by construction and their only outcome is settlement at expiry. paired_outcomes() says so in those words rather than reporting 'no rule fired'.
- settlement is unchanged
- AM-019 registered settlement at expiry for the v2 cohort and this amendment does not restate it. hold_to_expiry, from options_trade_resolve, remains the PUBLISHED outcome and is labelled `published: true` on every paired row. The rule-exit number is labelled `published: false, diagnostic: true` and carries no evidential standing until it has a sample.
- paired design
- Every trade settles TWICE — once at the rule's exit price, once held to expiry — on the SAME trade. Name, strikes, expiry, entry price and market path are identical by construction, so the only difference between the arms is the rule. That is why this needs no cohort split, no control group and no second registration batch: a 'next batch carries policies' design would compare trades struck on different dates into a different tape, and the date/market confound would be inseparable from the rule's effect.
- blind window
- TWO DAYS, disclosed exactly as AM-019 disclosed its five unregistered days. The 79 trades were struck at as_of 2026-08-03; today is 2026-08-05 and no mark has been taken. Any exit that would have fired on 2026-08-04 or 2026-08-05 before the first mark is UNOBSERVABLE and is not reconstructed. In practice the window is longer than two days and the reason is recorded here: on 2026-08-05 all 60 cached chains for the live names are still as_of 2026-08-03 — the registration date — so a first mark taken today would be priced off the very chain the trades were selected from and would carry measures_nothing = true on every row. The rules cannot begin to be evaluated until the chain cache advances.
- earliest evidence
- first hold arm settlement
- 2026-08-21 — 7 trades, all of which have NO rule in domain, so they produce a published settlement and no rule arm.
- first possible rule trigger
- 2026-08-28 — the dte_floor_21 trigger for the 2026-09-18 expiries.
- first paired outcome
- 2026-09-18 — 21 trades whose earliest expiry carries a rule in domain.
- per family AM 017
- Cohorts are never pooled across families, so evidence arrives per family and mostly late: Bull Call Spread and Iron Condor 2026-09-18 (n=1 each); Long Strangle and Put Spread (income) 2027-01-15 (n=19 and n=13); Bear Put Spread 2027-03-19 (n=2). The profit-target rule therefore reaches n=2 in September 2026 and its only family with a usable n — Put Spread (income), 13 trades — not before 2027-01-15. No number from either arm may be quoted before then.
- asymmetry disclosed
- NO loss-side rule is registered. Both rules truncate the right tail and leave the left tail intact, which mechanically raises the win rate of the rule arm without raising its expectancy — the classic way an exit policy flatters itself. A stop-loss was NOT added because its threshold would be a constant invented here with no derivation for this book (a debit structure's max loss is already the debit, and '2x credit received' has no meaning for the 56 long-premium trades). The bias is therefore disclosed and made analysable instead: every mark records WHICH rule fired, so any rule-arm outperformance can be decomposed by firing rule rather than taken at face value.
- refusal historical calibration
- These parameters are NOT calibrated on history, and the reason is SAMPLING, not data quality. The earlier statement that the frozen chains lack quotes is wrong and is corrected here: output/cache/options_hist holds 20,556 chain files and they DO carry bid, ask and greeks (verified 2026-08-05 on A@2026-06-30 — 522 contracts, 522 with delta, 69 with no bid). What they lack is DATES. There are 26 distinct chain dates in that store, month-end from 2024-07-31 to 2026-06-30, plus a single 2026-08-03 date covering only the 60 rescued v2 names — so exactly ONE options_hist date falls on or after the v1 book's first entry (overlay_as_of 2026-07-23). A rule that fires on a specific day cannot be calibrated on a monthly grid: the exit price at the moment the rule triggers is simply not in the store, and a month-end price is a different, path-dependent number. Stating this as a data-quality problem would have been refuted by the first person to open a chain file, and the refusal would have been discarded with it.
- refusal v1 marks are not a daily series
- A NEW refusal, not in the plan. v1's 47,435 marks (options_rec_marks.jsonl) look like a daily series and are not one. Of the 6,816 structures marked on BOTH 2026-07-30 and 2026-07-31, 6,160 (90.4%) carry an IDENTICAL exit_value, because 6,754 of the 6,816 rows on 07-31 were priced off the SAME 2026-07-28 chain as the day before (all counts measured 2026-08-05). Nothing in those rows distinguishes 'the book did not move' from 'the book was not repriced'. A rule back-fitted or evaluated against that series would trigger on days the chain never advanced. This is AM-022's defect — a zero mark meaning 'no new bar' rather than 'no move' — in a ledger AM-022 does not cover, so it is named here. options_mark_v2 fixes it forward: every mark records chain_as_of, the chain_as_of it is compared against (the previous mark's, or on the first mark the chain the trade was STRUCK on), and an explicit measures_nothing flag; when that flag is set EVERY rule goes to not_evaluated, including the pure-calendar DTE rule, because firing it would mean recording a fill at a price that was never refreshed.
- not evaluated is not not fired
- A chain that cannot be read is 'not evaluated today', never 'did not fire'. v1's marker skipped the ticker and wrote no row, which downstream is indistinguishable from a rule that looked and found nothing. options_mark_v2 writes the row anyway with evaluated=false and the reason, and paired_outcomes() WITHHOLDS the rule arm entirely for any trade with such a gap — 'the rule did not fire' is only an outcome if the rule was evaluated on every day it could have fired.
- no api cost
- The daily mark reads output/cache/options/<SYM>.json by path and never fetches. All 60 live v2 names are present at parse_version 2 with bid and ask (verified 2026-08-05, 60/60). Across the 61 registered names' chains, 10,171 of 76,282 contracts (13.3%) have no usable bid; a leg with no bid — or a short leg with no ask — REFUSES the mark for the whole structure, because an exit is a package and a package with one unquotable leg has no exit price. No substitution of mark, last, mid or zero.
- status
- Policy REGISTERED (79 rows, hash-chained). The first mark has NOT been taken and is held for owner review of these parameters. Taking it is what closes the window described above.
AM-022 · 2026-08-05 — AM-021's description of the live paper marks was imprecise — a zero mark can mean 'no new bar', not 'no move' (corrects AM-021)
- what AM 021 said
- part 3: 'the other three marks are 0.0 because price_ref equalled the mark-date close.'
- why that is imprecise
- The causal claim is right — those rows are zero because the two prices are the same number — but 'the mark-date close' is wrong for at least one of them, and wrong in a way that hides a real defect. Checked against the shipping ledgers: the 2026-07-31 emission priced MO at 74.92, which is the 2026-07-29 close, not the 2026-07-31 close. The mark taken on 2026-07-31 returned 0.0 because the newest bar the fetcher held was STILL 2026-07-29 — the cache had not advanced. The market had moved (MO closed 67.94 on 07-30 and 68.33 on 07-31); the ledger simply could not see it.
- the defect this exposes
- A paper mark run against a cache that has not advanced emits 0.000%, which reads as 'the book was flat' when it means 'there is no new data'. A reader counting marks as observations counts that row. On the live optimizer's ledger three of four rows are 0.0 and nothing in any of them distinguishes the two cases.
- what changed in the code
- ros_book_variants.mark() now records `marked_close_window` beside `price_ref_window`, counts `positions_on_the_same_bar_as_price_ref`, sets `measures_nothing` when every priced position is on the same bar it was priced at, and emits a `staleness_note` generated from those counts. The CLI prints a warning on any non-zero count. The live optimizer's own mark is NOT changed by this amendment — its ledger is append-only and already written, and altering how it marks is a separate decision.
- what stands from AM 021
- Everything else. The count of marks (four), the elapsed span (six calendar days), the single non-zero observation (+0.025% on the 2026-07-31 book, spanning the 2026-07-29 to 2026-07-31 closes — two trading days) and the conclusion that the accrual carries no evidential weight are all unchanged and were verified against the ledgers again while writing this.
- discipline
- AM-021 is NOT edited. This entry is appended and names what it corrects, which is the same treatment AM-018 gave AM-017.
AM-021 · 2026-08-05 — Covariance minimum-history rule (LIVE) + HRP twin paper book + CVaR as a diagnostic only
- registered before
- ros_book_variants.py, ros_cvar_diag.py and the ros_cov_model minimum-history rule were written to disk AFTER this amendment was appended. No result from any of them had been read when it was written.
- part 1 live change
- what
- ros_cov_model now requires each name to carry at least `min_history_td` of its OWN adjusted-close history before it may enter the factor covariance. Names below the floor are EXCLUDED and NAMED in `dropped`, with their observed bar count as the reason. The intersection calendar is then taken over survivors only.
- constant
- [object Object]
- why 504
- Two years of trading days. The covariance is the ENTIRE risk story of the P4 optimizer — there is deliberately no risk-aversion lambda, so the 10% volatility ceiling is the only thing standing between the alpha ranking and a concentrated book. A name with less than two years of history cannot support a risk estimate that a ceiling can lean on. 504 was chosen as the shortest window spanning two full annual reporting cycles, before any book was constructed on it, and is not tuned: the alternatives examined were 252, 378, 504, 630 and 756, and 252 through 630 all produce the SAME survivor set on the 2026-08-05 pool, so the choice inside that range changes nothing.
- defect it fixes
- ros_cov_model._returns intersected dates across ALL pool names, so the newest listing set the estimation window for everybody. Measured on the 2026-08-05 ROS5 pool (103 names, expected alpha > 0): the common calendar was 188 dates / 187 observations spanning 2025-11-03 to 2026-08-04, against a requested lookback of 756. Two names caused it — Q (188 bars) and PSKY (250 bars) — while the median pool name carries 6,729 bars. The shipped optimized_book.json records the consequence as covariance.obs = 185. The 10% ceiling was therefore estimated on roughly six months of a single recent regime.
- measured effect
- On the 2026-08-05 pool the rule drops 2 of 103 names (Q, 188 bars; PSKY, 250 bars) and lifts the estimation window from 187 to 725 observations, 2023-09-12 to 2026-08-04. The calendar is then capped by TKO (726 bars), which is reported as `binding` with caps_window true — 725 is short of the requested 756 and the file says so rather than implying a full window. Median annualised name volatility moves from 40.8% (187 obs) to 35.7% (725 obs).
- material consequence disclosed
- Q is HELD in the shipping optimized_book.json of 2026-08-03 at 0.83% of NAV. Under this rule Q has no risk estimate, is therefore not in the pool, and cannot be held. WITHHOLD, NEVER CLAMP: a name whose risk cannot be estimated is excluded from the book, not admitted at a default. This is a reduction in investable names, i.e. the conservative direction.
- not applied retroactively
- The shipping optimized_book.json (as_of 2026-08-03) was NOT re-run and NOT restated. Its as_of already carries a row in optimized_book_history.jsonl, and re-solving it would leave the live file disagreeing with its own recorded emission — a restatement in all but name. The live book picks the rule up at its next emission, and until then optimized_book.json.covariance.obs = 185 is the honest record of how it was actually built.
- withholding rule
- If, after the drops, the achieved intersection is still shorter than min_history_td, the covariance is WITHHELD (returns None with a stated reason) and the optimizer emits nothing. It does not fall back to a shorter window.
- part 2 twin cohort
- cohort id
- variant-hrp
- what
- Hierarchical Risk Parity over the same pool and the same covariance as the live optimizer, emitted to its OWN ledgers under output/ros/variants/ and never written into optimized_book*.json.
- never pooled
- variant-hrp is a separate cohort from the live P4 book on every axis: separate emission ledger (hrp_book_history.jsonl), separate NAV ledger (hrp_paper_nav.jsonl), separate registry chain (variant_registry.jsonl). No figure from either may be averaged, maxed, blended or compared as performance with the other. The twin does not replace the live optimizer and is not a candidate to.
- construction
- Correlation distance sqrt((1-rho)/2) -> single-linkage agglomerative tree -> quasi-diagonal leaf order -> recursive bisection with inverse-variance allocation between the two halves. Deterministic; no seed enters the weights.
- why hrp and not the other two
- It is the only variant whose inputs exist today. It introduces NO new free constant: no risk-aversion lambda, no delta, no tau, no view-confidence matrix. Everything it needs is a covariance, which the live book already computes.
- registered constants it reuses
- vol_target_pct 10.0, each name's own published ros.sizing.max_pct_nav, adv_participation_pct 10.0, reference_nav_musd 10.0, min_initial_pct_nav 0.5, max_names 62 — all read from ros_config, none redefined here.
- scaling rule
- HRP returns relative weights. They are scaled by ONE scalar s, found by bisection, to the largest value at which BOTH the registered 10% volatility ceiling and sum(w) <= 1 still hold, with each name additionally capped at its own registered limit: w_i = min(s * r_i, cap_i). One derived scalar, no free constant. Cash is the residual and is an OUTPUT, exactly as in the live book.
- constraints it does NOT enforce
- Sector caps and style bands. HRP has no mechanism for them and inventing one would be an unregistered constant that sets the answer. The twin therefore MEASURES its realised sector concentration and style tilts against the registered limits and reports every breach explicitly in `constraint_breaches`. A twin that breaches a registered limit is a finding about the construction, not a licence to trade it: variant-hrp is not investable as emitted and its artefacts say so.
- trade rules off
- The no-trade band and the cost hurdle are DISABLED on both sides of the construction comparison. They are an executability layer whose effect depends on the prior book, so leaving them on would attribute path dependence to construction.
- part 3 what the sample can and cannot show
- inception
- 2026-08-05. Zero trading days of variant-hrp accrual exist at registration. The live optimizer's own paper ledger holds four marks over six calendar days and exactly ONE non-zero return observation (+0.025% over two trading days on the 2026-07-31 book); the other three marks are 0.0 because price_ref equalled the mark-date close.
- detection arithmetic
- The smallest annualised return difference distinguishable from zero at t=2 over T trading days is 2 * TE * sqrt(252/T), where TE is the annualised tracking error between the two constructions. This is emitted on the face of every divergence report, computed from the shipping TE rather than quoted.
- cannot show
- Which construction is better. At a tracking error of roughly 4-5%/yr, detecting a difference of size IR*TE at t=2 requires about 4/IR^2 years — 16 years at the top of this platform's own registered plausible band (expectations_of_record: net IR 0.15-0.5) and roughly 178 years at the bottom. Six weeks, six months, or a single pair of books will not settle it, and any framing that implies otherwise is backtest marketing.
- can show
- Everything ex-ante, on day one, with zero accrual: volatility at a matched risk ceiling, concentration and effective N, the share of the live book's weight set by the objective versus by the caps and the covariance, one-way turnover to switch constructions, maximum single-name weight difference, which constraints bind and what each costs in alpha, and realised sector/style exposure against the registered limits.
- what accrual adds
- Exactly one thing: proof that the ledger discipline holds — dates recorded, nothing restated, no backfill, a mark landing on every scheduled run. That is a PROCESS proof and it is worth having. It is not a performance proof and may not be presented as one.
- precedent being avoided
- output/alpha_lab/paper_nav.jsonl holds four alpha-lab twin books at two marks each, last marked 2026-07-21, because alpha_paper_book.py was never wired into any orchestrator. The twin's record and mark are wired into mch_full_update.py in the same stage as the optimizer's, before the QA step, so a run that does not mark is a run that fails its gate.
- part 4 cvar diagnostic
- status
- DIAGNOSTIC ONLY. CVaR does not choose a single weight anywhere in this platform.
- what
- 1-year expected shortfall of the book ALREADY HELD, from the engine's own per-name Monte Carlo terminal-price quantile grids joined by an imposed copula.
- constants
- [object Object]
- why t5 not gaussian
- A Gaussian copula has ZERO tail dependence, which is precisely the wrong assumption in the tail an expected shortfall is measuring. df=5 is not fitted: it is the same df the engine already registers for the marginals it is joining, so the joint and the marginals do not disagree about tail thickness.
- the joint is IMPOSED not measured
- The engine runs each name's Monte Carlo independently and does not persist raw draws, so no joint structure exists on disk. The correlation fed to the copula is a DAILY-return correlation from cached adjusted closes over the same window and the same minimum-history rule as the factor covariance. It is applied to ONE-YEAR marginals. That horizon mismatch is not repaired by any copula family and is disclosed on every emitted figure.
- marginals are UNSHRUNK
- The quantile grids are the engine's raw Monte Carlo. The optimizer's alpha is shrunk by kappa=0.35 precisely because raw expected alpha is not an unbiased forecast. The diagnostic therefore reports the mean of its own distribution alongside every tail figure, so the optimism is visible rather than buried in a tail number.
- why it may not become an objective
- Three decisions are missing and all three would set the answer: (a) an owner-signed alpha-versus-CVaR trade-off constant, (b) whether the marginals get the kappa shrink before entering the objective, (c) a copula family defended rather than defaulted. Shipping the objective before those exist installs exactly the unregistered free constant the live design's 'no lambda' note exists to prevent. This is a SIGN-OFF block, not a sample block, and could clear quickly.
- withholding rule
- If names covering less than coverage_floor_pct_of_invested of the book's invested weight carry a usable quantile grid, the portfolio figure is WITHHELD with the uncovered names listed. It is not renormalised onto the covered subset: both renormalising and zero-filling understate the tail, and a tail number that is quietly too small is worse than no number.
- part 5 black litterman NOT BUILT
- status
- NOT BUILT. Waiting on SAMPLE, not on sign-off.
- why
- Three of its four inputs do not exist. delta (risk aversion) is not registered and the live design excludes it by name; pinning it to the registered 4.5% ERP gives 1.92, 2.68 or 2.41 depending purely on whether SPY volatility is measured over 3 years, 1 year or 187 days — a 40% spread set by a window choice, i.e. a free constant wearing the costume of a derived one. tau has no basis anywhere on disk. Omega, the view-confidence matrix, has NO empirical basis at all: forecast_ledger_2026-08.jsonl holds 852 open 1-year forecasts with zero resolved outcomes, the pre-registered REG-004 anchor cohort of 480 names shows n_forecasts = 0, and calibration.json disclaims its own 48.5% hit rate as 'NOT one the engine published'.
- additional defect
- The equilibrium prior would be built by cap-weighting the expected-alpha>0 pool — the cap-weighted subset that already passed the alpha filter, so the 'equilibrium' would be contaminated by the very signal it exists to be independent of. There is no market portfolio on disk to use instead.
- unblock condition
- At least 12 monthly cross-sections of RESOLVED forecasts on the REG-004 anchor cohort. The platform's own evidence_gate_2027 judges that no earlier than 2027-07.
- why it is recorded here anyway
- So that building it later cannot be presented as a new idea whose constants were chosen freshly. The constants it needs are named now, before anyone has seen what they would produce.
- amendment discipline
- This entry is APPENDED. Nothing above it in amendments[] was edited. If any constant here is later changed, a further amendment says so; this one stays as written.
AM-020 · 2026-08-05 — Portfolio simulation — forward-only, seeded, registered, and currently worthless as evidence
- what
- portfolio_sim.py simulates a book from the PUBLISHED rating history. Every run is seeded, config-hashed, and appended to a git-tracked hash-chained sim registry before its result is looked at.
- why forward only
- Point-in-time FUNDAMENTALS do not exist — Alpha Vantage serves current values only, and the one PIT snapshot on disk is 2026-07-28. A research sim run over history would therefore be scoring today's fundamentals against yesterday's prices, which is look-ahead of the most flattering kind. The publication ledger IS a genuine point-in-time record of what was PUBLISHED and when, so the sim uses that and only that: at each rebalance it may read ratings dated on or before that day.
- sample
- 12 rebalance dates spanning 2026-06-26 to 2026-08-03 — about six weeks. This carries NO evidential weight and no result from it may be quoted as performance. What is being built is the apparatus and its record, so that when a real sample exists the machinery is already pre-registered rather than fitted to the answer.
- deep history
- Rule-based sims over 5-25 years are permitted on the 26.7y adjusted-close history because they use NO fundamentals — but they must run on a delisted-inclusive universe or they are survivorship-biased. Any deep-history sim that cannot state its delisting treatment does not run.
- registry
- output/performance/sim_registry.jsonl — append-only, hash-chained, git-tracked. A row is written BEFORE the result is read, carrying the seed, the config hash and the universe, so a sim cannot be quietly re-run until it flatters and only the winner reported.
- surfaces
- No simulation output appears on any marketing surface, and none is published as a track record. The forward records that count are the forecast ledger and the two options ledgers, all of which settle against real outcomes.
- variants
- Optimizer variants (Black-Litterman, HRP, CVaR) run as TWIN PAPER BOOKS against the live optimizer — never as a replacement, and never pooled with it.
AM-019 · 2026-08-05 — v2 options ranking enters its forward record — inception 2026-08-05
- scope
- The rows the /options page PUBLISHES: the top N per bucket across all four buckets (income, leverage, hedge, volatility) — 80 structures on this build. NOT the full ranked pool of 1,424: a forward record must cover what was put forward as a recommendation, and registering structures no reader was shown would inflate every denominator with trades that were never offered.
- why all four
- Owner decision, 2026-08-05. The ledger is append-only and never backfilled, so an inception date is permanent. Registering the volatility families alone would leave them with no same-day, same-engine cohort to be compared against — the only alternative being the v1 book, selected on different rules and never poolable. Every family therefore starts on the same day under the same market, or the comparison this record exists to support cannot be made later.
- cohorts
- THREE axes, none collapsible into another. `cohort`=v2 (ledger identity; never pooled with v1). `family_cohort`=the structure family (AM-017; a straddle and a covered call are different bets with different loss shapes). `tenor_cohort` (AM-002; a 45-day call and an 850-day call are not one thing). This build is 100% `standard` tenor — no LEAPS-tenor structure reached the published rows.
- not registered
- /options went live 2026-07-31 and ran for five days with no forward record, because no path existed from a ranked row to a Trade. Those five days are NOT reconstructed. The record begins today and the gap stands in the open — backfilling it would be inventing pre-registrations after the outcomes began.
- families absent
- Long Straddle, Long Butterfly and Call Ratio Backspread were generated and assessed but did not reach the published rows, so they have no cohort yet. Their inception is the first build on which they rank — a later date than their siblings, which is a real limitation on comparing them and is stated rather than smoothed over.
- settlement
- options_trade_resolve at expiry. Nothing is scored at registration.
AM-018 · 2026-08-05 — Poor Man's Covered Call withdrawn — it is a diagonal, which AM-017 itself excluded (corrects AM-017)
- what
- PMCC is removed from SETTLEABLE_STRUCTURES and every co-change point. 13 families -> 12.
- why
- A Poor Man's Covered Call is a LONG-DATED long call financed by repeatedly selling SHORTER-DATED calls against it. The two expiries are the structure — that is what makes the long call a stock substitute and the short call 'covered'. AM-017 excluded calendars and diagonals because the ledger row carries a single `expiry` field, and then admitted a diagonal in the same breath. Constructed at one expiry it is not a PMCC at all, it is a deep-ITM bull call spread, which the set already covers.
- second defect
- The name also collided with the report router: 'poor man's covered call' CONTAINS 'covered call', so mch_json_summary routed it to `income` while options_rank.BUCKETS declared it `leverage` — the same structure in two different buckets on two surfaces (D-28). Withdrawing it removes the instance; the contract test was strengthened to assert bucket AGREEMENT rather than mere presence, which removes the class.
- readmission
- PMCC may be readmitted at schema v3, when a row can carry per-leg expiry. It is a legitimate structure — it is not legitimate at one expiry.
AM-017 · 2026-08-04 — Six single-expiry options families admitted, each as a NEW COHORT at inception
- what
- Long Straddle, Long Strangle, Long Butterfly, Iron Condor, Call Ratio Backspread and Poor Man's Covered Call join the settleable set. A fourth published bucket, `volatility`, carries the structures whose thesis is dispersion rather than direction.
- why
- The engine could form a volatility view and had no way to express one: all seven existing families are directional or covered overlays, so a name whose IV/RV and term structure argued for a straddle got a bull call spread instead.
- cohort rule
- Each family is its OWN cohort from its first pre-registration. Never pooled with the existing seven for any accuracy claim, and never pooled with each other — a straddle and a covered call are different bets with different loss shapes, and an average over both describes neither. Same discipline as v1/v2 and the LEAPS tenor.
- excluded
- Calendars and diagonals are NOT admitted. The ledger row carries a single `expiry` field, so a multi-expiry structure cannot be keyed or settled; that needs schema v3 with per-leg expiry and is deferred.
- ratio constraint
- Ratio backspreads are v2-trade_id-only. The v1 leg-display format carries no quantity, so a 1x2 and a 1x1 at the same strikes collide in the v1 keys. v2 identity hashes leg descriptors including qty and is safe.
- no weight change
- No existing weight, threshold or published number is altered. This admits new structures to generation and ranking; it does not re-rank or re-price the old.
AM-016 · 2026-08-04 — Shadow scenario engine — dynamic probability weights computed, recorded and scored, but NOT published
- config
- ros-1.12.0 · new scenario_shadow family. The authored archetype tuples are UNCHANGED and remain the sole published source of scenario probabilities.
- what
- A deterministic, config-versioned function derives a per-name five-scenario probability vector from inputs that already exist per name: the nine conviction components, shrunk beta, macro betas (rates/oil/USD with t-stats), style and thematic factor percentiles, three volatility measures, the forensics set (Piotroski, Altman, Beneish, Sloan, ROIC, operating leverage), FCF quality, EPS beat rate, insider flow, transcript disconfirmation and the IV term structure. Output goes to a summary block stamped authoritative: false and to an append-only scenario_shadow_ledger.jsonl, so the shadow, the authored view and the realized outcome can all be scored against each other as forecasts mature.
- why
- Scenario probabilities are the load-bearing input to the PWEV, which sets the published target and rating. Measured on the live book, only TWO distinct probability vectors exist across four archetype templates: (0.20, 0.17, 0.35, 0.20, 0.08) covers 684 of 893 configs (77%) and the top two cover 82%. Twenty-five distinct probability values exist across 854 names. Presenting that as a name-level probabilistic judgement overstates what was assessed per name. The inputs to do better already exist and are already computed every week.
- shadow first is the point
- The shadow publishes NOTHING and changes NOTHING. It exists to accumulate a scoreable record so the question 'are dynamic weights actually better calibrated than the template?' is answered by evidence rather than by preference. Building it and switching to it in one step would be exactly the failure this platform's pre-registration discipline exists to prevent.
- fallback
- Where any input is missing, the authored archetype tuple is used and the fallback is logged. A missing input therefore returns the system to the more conservative, already-published behaviour rather than to a partially-informed estimate.
- firewall
- ros_prob_diagnostics keeps its rule that diagnostics never feed back. QA-470 is amended NARROWLY to permit the shadow block to exist while retaining every invariant it enforces. New QA-473 proves non-leakage by machine: the authored probabilities in every published config must be byte-identical to the archetype tuples in sp500_archetypes.py. Sum-to-one, the eps x multiple scenario-target reconciliation, the 20% structural-impairment floor and the PWEV chain check are all unchanged.
- cutover evidence gate
- Cutover requires ALL of: at least two quarters of scored shadow forecasts; shadow Brier AND CRPS better than the authored view by a pre-stated margin at a pre-stated minimum n; no sector or rating cut materially worse; and a separate owner-signed amendment at that time. Calendar time alone never triggers it. If the gate fails, the authored tuples stand and the shadow is retired or redesigned.
- effect on published output
- NONE. The estate run that introduces this must produce zero diff in any authored probability, rating or target — that no-op is the acceptance test.
- approved by
- Marinus Heymann, 2026-08-04
AM-015 · 2026-08-04 — True LEAPS — the options engine may select expiries beyond 365 days, as its own tenor cohort
- config
- ros-1.11.0 · options_generate. HORIZONS gains a 'leaps' bucket [450, 550, 650, 750, 850] and the per-name horizon is passed rather than defaulted. No liquidity threshold, ranking weight or scoring constant is changed.
- what
- options_feed passes a per-name horizon derived from catalyst timing and the valuation horizon (unmapped names keep 'medium', so the change is behaviour-preserving where no mapping exists). options_generate gains an optimal-expiry mode that scores across all liquid expiries in the chain rather than snapping to the nearest target, on theta profile, term structure, capital efficiency and catalyst coverage — explicitly NOT 'pick the longest'. Structures selected beyond 365 days are tagged tenor_cohort: 'leaps'.
- why
- The engine has never published an option beyond 164 days, and the reason turned out to be a missing argument rather than a judgement: options_feed never passed a horizon, so every name in the estate ran the 'medium' default of [60, 120] days, and the HORIZONS table capped at 365 in any case. Alpha Vantage's chains DO carry true LEAPS — MSFT lists 23 expiries out to 865 days — and those long-dated contracts are the most liquid part of the chain by the engine's own gate, passing at 72% against 21% at zero days to expiry. The platform was declining its best-quoted contracts by accident.
- cohorts
- LEAPS structures are a NEW COHORT from first emission and are never pooled with the 45-164 day cohort for any accuracy or return claim. The v1 overlay is untouched; its r=0.04/q=0 freeze and the 8,236 pre-registered v1 structures stand.
- label discipline
- The 'LEAPS' label may appear only where days to expiry exceeds 365, enforced as a BLOCKING check. The 860 existing structures carrying that label at 45-164 days (defect D-19) are corrected by a DISPLAY-NAME MAPPING AT RENDER ONLY — settlement keys on (entry_hash, structure_name) and the summary block buckets by substring, so a rename would orphan settlement keys and silently empty a bucket. The ledger key is immutable.
- known limitation
- The house distribution is re-horizoned from a one-year Monte Carlo, so an 865-day option implies a scale ratio of 2.37 — beyond anything the tenor scaling has been tested at. Bounds and 865-day test cases land BEFORE any LEAPS candidate is scored, and the term-structure IV is preferred at long tenor.
- effect on published output
- The options surface may show longer-dated structures for liquid names. No equity rating, target or probability changes.
- approved by
- Marinus Heymann, 2026-08-04
AM-014 · 2026-08-04 — Forecast-distribution ledger — the predicted probability becomes a stored, scoreable object
- config
- ros-1.11.0 · UNCHANGED. No weight, threshold, rating rule or scenario probability is touched by this amendment. It adds a capture surface and unifies scoring; it does not alter what the engine predicts.
- what
- A hash-chained, append-only forecast ledger (output/performance/forecast_ledger_YYYY-MM.jsonl, monthly segments with the chain spanning segments) records, for every name at every publish: rating, expected return, conviction, horizon, the Monte Carlo percentile vector p5-p95, prob_above_current, prob_above_target, all five (scenario probability, scenario target) pairs, a 41-point decimated quantile grid plus the SHA-256 of the full 1001-point grid, and the engine/config/input hashes. Scoring is unified into a single core (scoring_core.py); ros_calibration and performance_store become adapters over it and track_record.py is retired. Matured rows are scored with PROPER scoring rules — Brier, log-loss and CRPS — computed from the stored probabilities, cut by sector, market cap, rating and volatility regime.
- why
- The platform has published probabilities for four months and has never stored one. The Brier score on the accuracy page today is a PROXY, reconstructed after the fact as p = clip(0.5 + predicted_upside, 0, 1) from a price target — it scores the target, not any probability the engine emitted. Meanwhile the real distribution IS computed and persisted per run in output/weekly_FULL_<date>/*/summary.json, but that tree is untracked by git and is never joined to an outcome. Three separate scorers (track_record, performance_store, ros_calibration) implement three different hit-rate methods on three different inputs and produce three different answers. This amendment fixes the capture side first because capture cannot be backfilled.
- cohorts
- The proxy-era numbers and the distribution-era numbers are SEPARATE COHORTS and are never pooled, on the same discipline that keeps the v1 and v2 options cohorts apart. The distribution cohort begins at the ledger's first emission and carries no history.
- no backfill
- The 26 retained price dates (2026-04-24 onward, ~13,600 summaries) are NOT loaded into the ledger. They remain diagnostics. Backfilling them would manufacture a track record the platform did not commit to at the time, which the standing constraints forbid.
- defect fixed in the making
- performance_store assigns a forecast to the pre-registered anchor cohort only if its earliest rated point falls on or after the anchor date. Every REG-004 name already had a rated point at 2026-06-26/27, so all 494 scored forecasts fall into 'legacy' and the pre-registered REG-004 cohort publishes n_forecasts: 0 — the anchor's forward record has never been scored by the scoreboard. This is a defect in the scoring code, fixed here. No stored record is altered; only the classification logic.
- effect on published output
- NONE on any rating, target, probability or number derived from them. The accuracy surface gains a second, separately labelled cohort and states that the legacy figure is a proxy.
- enforcement
- QA-490 chain integrity across segments (BLOCKING) · QA-491 every published name has a same-day ledger row (BLOCKING) · QA-492 scenario probabilities sum to 1 within tolerance in every row · QA-493 no-mutation, proven by replaying the prefix hash.
- falsifiable
- Once n is sufficient, a reliability diagram built from stored probabilities either lies on the diagonal or does not. If the published distributions are systematically overconfident or miscentred, this ledger is what will show it — including if it shows the engine is worse than the proxy era suggested.
- approved by
- Marinus Heymann, 2026-08-04
AM-013 · 2026-07-31 — IV-measure attribution — the vol measure that selects the structure is now declared and enforced
- config
- ros-1.10.0 · UNCHANGED. No weight, no threshold and no options_selector.decision_table row was touched; data/ros_config.json is byte-identical to the archived ros-1.10.0.json (md5 e25666a44a9c47637acafeb9faf6aebd), iv_low_pctile 33 / iv_high_pctile 66 included.
- what
- ros_options_intel now emits BOTH vol measures and names the load-bearing one. New fields: iv_bucket_driver (iv_rv_percentile_cross_sectional | iv_regime_label) and iv_rank_is_advisory: true. The reasoning states the selecting measure verbatim ("This is the measure that selects the structure above") and flags where the two disagree. New BLOCKING check QA-416 fires when the self-relative rank is displayed alongside a cross-sectional percentile and the block fails to declare the driver, or the prose fails to attribute it. QA-416 is in the enforced ros_broken fixture floor.
- why
- The self-relative monthly IV rank added in options P2 was DISPLAY-ONLY, while the vol bucket that indexes the decision table remained the CROSS-SECTIONAL IV/RV percentile. They are different quantities. Both are present for 467 of 854 names, and on those 467 they disagree on the bucket 311 times (67% of dual-measure names, 36% of the estate). Those 467 were about to publish prose attributing the recommendation to a measure that had not produced it — NTAP: "2th percentile -> low vol -> cheap options, buy defined-risk downside" (IV/RV percentile 2.1, bucket low, direction bearish) beside a self-relative reading of 100.0 (decile 10). The prose was not wrong about a number; it credited the wrong number.
- evidence
- 854 names re-emitted (output/weekly_ROS5). QA-416: 467 names failing before the fix, 0 after. Driver split: 474 cross-sectional, 380 regime-label. iv_rank_method: 838 self_relative_monthly, 16 iv_rv_percentile_interim. All 854 preferred structures verified UNCHANGED against the already-published reports.
- effect on published output
- NONE on decisions — DISCLOSURE-ONLY by construction. 854/854 recommendations unchanged. The only output deltas are reasoning strings and the two new declarative fields. What changed is what the reader is told about which measure produced the recommendation.
- why no config bump
- The registered ritual is explicit that weight/param changes require a version bump. Nothing in ros_config.json changed — all ten prior bumps carried a new or changed numeric constant; none was disclosure-only. Bumping here would archive a file identical to ros-1.10.0 apart from its version string, degrading the archive as a record of parameter states, and would split a mid-flight estate across two config versions for zero parameter change (the hazard flagged in amendment #5). Precedent for a no-bump amendment: amendment #6.
- deliberately not done
- Switching the selector to the self-relative rank. It would change 209 preferred structures, and that measure has collapsed toward its ceiling (median 83.3, 92% of names >= 50, 13% at exactly 100) — it is reading a market-wide vol regime, so it would express one correlated bet as 838 independent name-level calls. A regime-demeaned rank is the real fix and is deferred as its own workstream, consistent with the standing options_sleeve deferral of IV-rank-driven decisions.
AM-012 · 2026-07-31 — Optimism taper and tail-metric redesign for option economics
- config
- ros-1.10.0 · options_distribution (dist-1.0.0)
- what
- Option economics are computed on a MIXTURE of the house MC and a drift-free lognormal under the chain's own ATM IV. The house view is capped at 0.25 and tapered to ZERO at a 20pp authored-vs-MC gap: w = min(w0,0.25)·clip(1−|gap_pp|/20,0,1). Separately, the downside is reported as prob_total_loss_pct / prob_loss_50_pct / expected shortfall with a tail_is_degenerate flag, replacing a CVaR that cannot order.
- why
- Two measurements, both on the live estate. (1) The authored probability layer is one-directionally optimistic: 274 of 854 names diverge >20pp from our own MC and 272 of those are MORE bullish; root cause is 25 distinct authored values across 854 names. (2) cvar5_pct is exactly −100% for 100% of Bull Call, Bear Put and Long Call structures and 99.5% of income Put Spreads — correct, and useless for ranking.
- registered before
- Any ranker exists. The constants are fixed now so they cannot later be tuned against the ranking they produce.
- effect on published output
- NONE. The published economics still come from the house MC alone; the blend is carried in the v2 shadow only. QA-484 blocks it from claiming authority.
- direction is conservative by construction
- The taper shrinks toward the market distribution, which contains no house view, so a name whose probability layer is internally incoherent loses the ability to talk its own options up.
- absolute gap is deliberate
- |gap_pp|, not the signed gap. The diagnostic measures INCOHERENCE between two house-derived quantities; tapering only when the disagreement flatters us would be fitting the rule to the desired answer.
- falsifiable
- Measured impact: Spearman 0.61 house-only vs blended over 2,447 structures, top-20 overlap 10/20, 31.1% tapered to zero. If a later measurement shows the blended ranking is indistinguishable from a pure risk-neutral one, the house weight is doing no work and the cap should be revisited — not raised silently.
- known limitation
- At a 0.25 cap the blend is at least 75% market-driven, so the ranking is closer to a relative-value ranking than to a house-view ranking. That is the honest consequence of the measured bias, and it is coherent with the backtest finding that the options sleeve's historical return was attributable to the volatility risk premium rather than to our selection.
AM-011 · 2026-07-31 — Options liquidity tiers, calibrated and pre-registered before anything gates on them
- config
- ros-1.9.0 · options_liquidity
- what
- Liquidity is classified into four tiers (institutional / retail / illiquid / unquotable) rather than gated by a single pass/fail threshold. `institutional` (bid>=0.05, OI>=100, spread<=20%) is registered NOW as the only tier that will ever be eligible for public RANKING, before any ranking exists to be tuned against it.
- why
- The v2 core inherited thresholds from the specification and nobody had measured what they do. Measured on the contracts the overlay actually selects — not whole chains, which give a meaningless 10.4% — the published book has median open interest 13, 21.7% of legs at ZERO open interest, 7.1% with no bid, and a median spread of 30.1%. The SLAB contract that produced the 1099% failure was not an outlier; it was the middle of the distribution. Registering the ranking threshold before the ranking is built is the point: choosing it afterwards, against results, is how a threshold becomes a fit.
- evidence
- 854 frozen chains @2026-06-30 · 8,510 selected contracts · 7,361 structures across 846 names · docs/OPTIONS-LIQUIDITY-CALIBRATION.md
- effect on published output
- NONE. P2 classifies; nothing is suppressed, reordered or hidden. Every structure publishes exactly as before, now carrying its tier. Whether `illiquid` structures continue to be published at all changes what ~65% of reports show and is explicitly deferred to P4 for Marinus's decision.
- falsifiable
- At `institutional`, 748 structures across 298 names survive — ample for a top-20 in each of three buckets, so ranking loses nothing by being strict. If a later measurement shows the ranked set cannot be filled from this tier, the tier was wrong and this amendment is what it will be checked against.
- known limitations
- 0
- month-end snapshots understate liquidity — re-measure mid-month before P4
- 1
- AV provides no bid/ask size, so 'executable' means quoted, not in size
- 2
- volume deliberately excluded (SLAB 220 P: volume 0, open interest 1,207)
· 2026-07-30 — Alpha P6 — probability diagnostics, stance hysteresis, kill switch
- n
- 10
- approved by
- Marinus (2026-07-30), sequencing P6 before P5
- registered
- probability diagnostics
- [object Object]
- stance hysteresis
- [object Object]
- kill switch
- [object Object]
- first run finding
- The diagnostics' first estate pass flagged 274 of 854 names on internal coherence, and the disagreement is ONE-DIRECTIONAL: all 272 names with a gap wider than 20pp have the AUTHORED view more bullish than our own Monte Carlo, none the reverse (median +11.6pp, 64% authored-higher). Root cause: only 25 distinct authored mass-above-spot values across 854 names, two archetype tuples covering 678 of them. RECORDED, NOT CORRECTED — which side is right is a research judgement this module is deliberately barred from making.
- note
- config ros-1.7.0 -> ros-1.8.0, both archived. Not fitted to MCH outcome data.
· 2026-07-30 — Alpha P4 — the registered trade rules are now ENFORCED, not merely measured
- n
- 9
- approved by
- Marinus (2026-07-30)
- no new constants
- Nothing here is new. book_construction already registered the no-trade band (|target/current - 1| > 0.25), the cost hurdle (marginal alpha > 2x modelled cost) and the cost itself (equity_slippage_bp_per_side = 10). They existed as prose that solve() never applied. Mirrored into ros_config optimizer.trade_rules; ros-1.6.0 -> ros-1.7.0, archived.
- implementation
- cost hurdle
- a PRE-solve bound: a name may be held or sold but never bought above its prior weight unless alpha >= multiple x cost. A new position is charged the ROUND TRIP (0.40%), an increase one side (0.20%).
- no trade band
- post-solve, iterated to a fixed point: a held name whose target is within 25% of its current weight is pinned and the optimiser reallocates around it
- defect fixed in the making
- First implementation applied the hurdle AFTER the solve, in a single pass. The re-solve then refilled freed capacity with names that were never re-checked, so sub-hurdle names still reached the book — JBL held at 0.17% alpha against a 0.40% hurdle. Whether a NEW position clears its cost depends only on the name's alpha, so the test belongs before the solve, where it needs no iteration. The no-trade band genuinely depends on the target and keeps the two-pass shape, now iterated to a fixed point.
- measured
- held
- 53
- invested pct nav
- 43.99
- vol pct
- 8.06
- expected alpha shrunk pct
- 1.77
- cost hurdle blocked
- 15
- lowest alpha held pct
- 0.42
- hurdle pct
- 0.4
- cost of enforcement
- 1.84% -> 1.77% shrunk alpha; the 0.07pp is the price of not buying alpha too small to cover its own spread
· 2026-07-30 — Alpha P4 — book-size cap raised 40 -> 62
- n
- 8
- approved by
- Marinus (2026-07-30), on the amendment-#7 unreachability finding
- change
- max names
- [object Object]
- config
- ros-1.5.0 -> ros-1.6.0 (both archived)
- rationale
- Resolves the mutually unsatisfiable trio by relaxing book size, the constraint Marinus chose to give up. The volatility ceiling now becomes the binding control instead of an inert number, and the name cap continues to inherit the published sizing layer.
- concentration caveat
- A 62-name book drawn from a 76-name qualifying pool holds ~82% of everything that clears its hurdle. That is close to inclusion, not selection: the ~0.7pp alpha gain over a 40-name book comes from holding MORE, not from choosing BETTER. The '500 buys -> own 30' framing of the original review no longer describes this book, and nothing published about it should imply concentration-driven selection.
- note
- kappa, clip, pool rule, name cap, sector cap, style band and vol target are UNCHANGED.
- measured after change
- held
- 62
- invested pct nav
- 51.46
- vol pct
- 9.41
- vol target pct
- 10
- expected alpha shrunk pct
- 1.84
- binding
- volatility ceiling (9.41 vs 10.0) · 62 names at their own cap · style band
- pool coverage
- 62 of 76 = 82%
- defect found applying this
- ros_optimizer's --max-names argparse default was the literal 40 and ocfg['max_names'] was read NOWHERE, so raising the constant updated ros_config, its archive and this amendment while every run still built the book from 40. A pre-registered constant the code never reads is worse than no constant: the record and the artefact disagree silently. Fixed to resolve from ros_config with the CLI as an explicit override. The 0.7pp cost attributed to the 40-name cap in amendment #7 is confirmed REAL, not a solver artefact — 1.11% at 40 vs 1.84% at 62.
- not yet enforced
- C8's 25% turnover band and C9's swap filter are MEASURED and written to the turnover ledger but are NOT constraints in solve(). The first emission reports 25.7% one-way turnover against the prior book, which exceeds the registered band. Meaningless here (there is no real prior rebalance yet) but it must become a constraint before any tracked book runs.
· 2026-07-30 — Alpha P4 — style band widened 0.5 -> 1.0
- n
- 7
- approved by
- Marinus (2026-07-30), on the amendment-#6 constraint-binding analysis
- change
- style band z
- [object Object]
- config
- ros-1.4.0 -> ros-1.5.0 (both archived)
- rationale
- At +/-0.5 the style bands bound so hard that the volatility ceiling — the control the sign-off intended to govern invested weight — never engaged: 24 names, 20% invested, 3.8% vol against a 10% target. At 1.0 the vol ceiling becomes the binding constraint, which is the architecture that was approved.
- honesty caveat
- This constant was changed AFTER observing which value produced a larger book and higher modelled alpha. That is in-sample specification choice and is recorded as such. Two things limit the concern: it is motivated by constraint ARCHITECTURE (making the approved primary control actually bind), not by realised returns; and no out-of-sample clock had started, since the optimizer has never emitted a tracked book. The band is now FIXED and any further change requires the standard evidence gate.
- note
- kappa, the clip, the pool rule, the name cap and the vol target are UNCHANGED.
- measured after change
- held
- 37
- invested pct nav
- 30.46
- portfolio vol pct
- 5.56
- expected alpha shrunk pct
- 1.11
- binding
- book-size cap (37 of a 40 maximum) · 36 of 37 names at their own max_pct_nav · style band 1 still touching at the wider 1.0
- vol ceiling still does not bind
- true
- structural finding
- Widening the band moved the binding constraint from style to BOOK SIZE; the volatility ceiling still does not bind, and under the approved constants it CANNOT. Arithmetic: the name cap inherited from the sizing layer is ~0.83% of NAV, so a 40-name book can be at most ~33% invested, which carries ~7.1% volatility. Reaching the 10% target needs ~62 names (~51% invested). The three approved constraints — inherit the sizing layer, hold 25-40 names, target 10% vol — are therefore mutually unsatisfiable: two of them cap volatility below the third. The 10% target is inert by construction; the real risk control is (name cap x book size). Measured cost of the 40-name cap at band 1.0: 1.12% vs 1.84% modelled alpha at 62 names, i.e. ~0.7pp. Recorded, NOT resolved — which of the three to relax is Marinus' decision.
· 2026-07-30 — Alpha P4 — optimizer built; measured behaviour of the approved constraint set
- n
- 6
- registered
- objective
- maximise w.alpha_input subject to the signed-off constraints
- no risk aversion lambda
- deliberately omitted — a lambda would be an unregistered free constant that silently sets the answer; risk is controlled by the explicit vol ceiling and the caps, all of which were signed off
- covariance
- Sigma = B F B' + D, cross-sectional factor model; Ledoit-Wolf shrunk sample cov as cross-check
- cardinality
- positions below min_initial_pct_nav are dropped from the WEIGHT VECTOR, not just the printout
- measured first emission
- pool
- 76
- held
- 24
- invested pct nav
- 19.92
- portfolio vol pct
- 3.76
- vol target pct
- 10
- expected alpha shrunk pct
- 0.81
- binding
- style_value +0.487 at the 0.5 band · style_lowvol -0.490 at the band · Information Technology at the 25% sector cap · all 24 names at their own max_pct_nav
- finding
- The volatility ceiling does NOT bind. The style bands bind first, at 24 names and ~20% invested, so the approved resolution 'the 10% vol target governs invested weight' is not what happens in practice — the style bands and per-name caps govern. Recorded because the sign-off assumed otherwise; widening the bands is a decision for Marinus, not a silent retune.
- note
- No constant was changed to obtain this result. Constants remain as amendments #4 and #5.
· 2026-07-29 — Alpha P4 — candidate pool and volatility-target resolution (approved)
- n
- 5
- approved by
- Marinus (2026-07-29), completing the P4 constraint sign-off
- registered
- candidate pool
- expected_alpha_pct > 0 ONLY. Alpha Score ranks WITHIN the pool and may never admit a name to it. Replaces the originally planned 'top ~120 by Alpha Score', which measured 76 of 120 names with expected alpha <= 0.
- pool floor rule
- if the pool falls below ~15 names the optimizer emits a SMALLER book and more cash — the filter is never relaxed to reach a target book size
- volatility target
- HOLD 10% annualised (resolution (i)). The target explicitly governs invested weight: because the surviving names are high-beta by construction (top-30 average shrunk beta 1.25 vs 0.81 estate median), the optimizer must solve for invested weight such that modelled portfolio volatility <= 10%, and disclose the resulting cash weight as a deliberate output rather than a residual.
- note
- Completes the P4 sign-off gate (with amendment #4: kappa=0.35, +/-25% clip, name cap inherits ros.sizing.max_pct_nav). All four constants declared before any optimizer code exists and not fitted to MCH outcome data. ros_config carries them once the in-flight phase-4 re-emit completes — a mid-run version bump would split the estate across two config versions (QA-406).
· 2026-07-29 — Alpha P4 — optimizer alpha-input shrinkage and position cap (approved constants)
- n
- 4
- approved by
- Marinus (2026-07-29), on the P4 constraint proposal
- registered
- alpha input to objective
- clip(expected_alpha_pct, -25, +25) * kappa
- kappa
- 0.35
- clip pct
- 25
- single name cap
- inherit the published ros.sizing.max_pct_nav for each name (resolution (a)) — the optimizer may never size a name above what its own report publishes; median 0.83% of NAV today
- consequence disclosed
- a 30-name book is therefore ~25% invested; the model portfolio is a risk-budgeted sleeve, not a fully-invested fund
- rationale
- Raw expected alpha is not an unbiased forecast: the benchmark-relative alpha CI still crosses zero, and an unshrunk optimizer concentrates into the names whose authored scenario targets are most aggressive. kappa=0.35 with a +/-25% clip takes the modelled 30-name book from +23%/yr to ~+7%/yr.
- gate to change kappa
- kappa rises ONLY through the standard evidence gate: >=26 weekly cross-sections, |IC| t>=2, <=5 changes per quarterly review, via a dated amendment + ros_config bump. It is not tunable at runtime.
- note
- Declared BEFORE the optimizer's first emission and before any optimizer code exists. Not fitted to MCH outcome data. STILL OPEN, not registered here: the candidate-pool definition (proposed: expected alpha > 0) and the volatility-target resolution.
· 2026-07-29 — Alpha P2 — options model economics + sleeve risk budget
- n
- 3
- registered
- payoff engine
- model_econ per structure: engine MC terminal distribution sqrt-time-scaled to option tenor, priced through chain premiums; EV/PoP(mc)/PoP(iv)/expected-loss/CVaR5/sortino/kelly (kelly display-only, never a sizing input — QA-417)
- sleeve budget
- [object Object]
- pop divergence threshold pts
- 25
- note
- Constants declared before the first estate emission of model_econ. None of these constants is fitted to MCH outcome data. Chain economics remain verbatim (QA-412 unchanged); model economics live under the separate model_econ key (options_econ split).
2 · 2026-07-29 —
- change
- Alpha Program P1 constants registered BEFORE first emission: required return rr = rf_1y_proxy(3m/2y CMT midpoint, dated macro store) + beta_shrunk(1y OLS vs SPY, Vasicek 0.65/0.35) x ERP 4.5% + size/liquidity premium 0/+100/+200bp (>=10B / 2-10B / <2B; ADV<$10M bumps one bucket, cap 200bp); industry premium slot 0 (deferred). Alpha Score v1 static weights: ea_pctile 0.35 / conviction 0.25 / thesis_health 0.15 / event_tone 0.15 / options_market 0.10 / forensic_quality 0.00 (dormant until the Phase-3 factor registry; any change = config bump + dated amendment). None of these constants is fitted to MCH outcome data. Expected alpha is REPORTED alongside existing metrics; rating/rules mechanics unchanged (QA-434 enforces).
1 · 2026-07-20 —
- change
- Live neutral/banded books now CONCENTRATED to top-30 by post-neutralization weight, with a re-projection within the survivors to restore factor-neutrality. Rationale: monitorability, per-name conviction, lower implementation cost. Cost acknowledged: fewer names → lower Grinold breadth → lower IR ceiling; and 30 long-only names cannot perfectly neutralize 3 factors, leaving small residual style tilts (first estate: val +0.21, mom -0.13; market beta still hedged to ~0). The DIFFUSE full-neutral book is retained in the calibration backtest as a control so the concentration trade-off is measured, not assumed. TOP_N=30 in alpha_book.py.
- affects
- 0
- book_construction.live_variant
Assumption register
56Every load-bearing constant, registered with its value, its code reference and its rationale — QA-495 fails the build when a registered value drifts from the code. Impacts above 5% require an approval reference.
| ID | Value | Why | Reviewed |
|---|---|---|---|
| mc.regime_probability | 0.25 | Mandated conservatism: a tech-selloff state in which the terminal multiple is shocked down. CLAUDE.md requires >=25% weight on such a regime. | 2026-08-04 |
| mc.regime_shock | -0.35 | How far the terminal multiple is cut in the compression state. | 2026-08-04 |
| mc.student_t_df | 5 | Student-t rather than Gaussian draws for growth, margin and multiple. CLAUDE.md mandates fat tails as the realistic case for liquid equities. | 2026-08-04 |
| engine.struct_impair_weight mch_stock_engine:ENGINE_DEFAULTS['struct_impair_weight'] | 0.2 | At least this much probability mass below 0.75x spot, so a bear case is a real bear case. | 2026-08-04 |
| engine.struct_impair_frac mch_stock_engine:ENGINE_DEFAULTS['struct_impair_frac'] | 0.75 | What counts as impairment: a scenario target below this multiple of spot. | 2026-08-04 |
| engine.prob_above_low mch_stock_engine:ENGINE_DEFAULTS['prob_above_low'] | 30 | Below this the engine warns that bear weighting or opex may be too severe. | 2026-08-04 |
| engine.prob_above_high mch_stock_engine:ENGINE_DEFAULTS['prob_above_high'] | 85 | Above this the distribution is likely thesis-confirming. | 2026-08-04 |
| engine.eps_drift_tol mch_stock_engine:ENGINE_DEFAULTS['eps_drift_tol'] | 0.1 | CLAUDE.md: recalibrate if engine-derived EPS deviates >10% from guidance. | 2026-08-04 |
| engine.dcf_mc_gap mch_stock_engine:ENGINE_DEFAULTS['dcf_mc_gap'] | 0.5 | Beyond this the two legs are investigated rather than blended silently. | 2026-08-04 |
| engine.rev_growth_clip mch_stock_engine:ENGINE_DEFAULTS['rev_growth_clip'] | -0.4 · 2 | Truncates the growth draw. Binding mainly for hyper-growth names. | 2026-08-04 |
| engine.gm_clip mch_stock_engine:ENGINE_DEFAULTS['gm_clip'] | 0.1 · 0.9 | Truncates the margin draw. | 2026-08-04 |
| engine.pe_clip mch_stock_engine:ENGINE_DEFAULTS['pe_clip'] | 3 · 200 | Truncates the multiple draw. | 2026-08-04 |
| engine.anomaly_range_x mch_stock_engine:ENGINE_DEFAULTS['anomaly_range_x'] | 6 | A recorded range wider than this is treated as a split/data error rather than volatility. | 2026-08-04 |
| engine.scenario_prob_tol mch_stock_engine:ENGINE_DEFAULTS['scenario_prob_tol'] | 0.01 | How far the five weights may sum from 1.0. QA-470 enforces it at publish. | 2026-08-04 |
| engine.scenario_target_tol mch_stock_engine:SCENARIO_TARGET_TOL | 0.01 | target must equal eps x multiple within this; the engine rewrites and keeps target_authored. | 2026-08-04 |
| triangulation.weights mch_triangulation:TRI_WEIGHTS | [object Object] | How the five valuation anchors combine into the published fair value. Single-sourced; QA-304 blocks a second definition. | 2026-08-04 |
| sbc.annual_dilution | 0.03 | CLAUDE.md mandates SBC as a real economic cost, charged once as dilution against the PWEV. | 2026-08-04 |
| beta.min_obs alpha_common:MIN_BETA_OBS | 120 | Below this the estimator returns NO estimate rather than a beta. It used to return a hard 1.0, which is a fabricated market exposure indistinguishable in any output from a name that genuinely tracks the market. The floor existed as three uncoordinated literals (alpha_common, options_greeks, ros_qa_estate's QA-429), so the gate meant to catch the fallback compared one copy of the constant against another: raising the estimator's floor made every published beta the fallback, moved the options-greeks headline +34.5%, and left QA-429 silent at exit 0. One symbol, imported by both, registered here so QA-495 catches an edit. | 2026-08-05 |
| options.leaps_min_dte options_econ:LEAPS_MIN_DTE | 365 | A LEAP is a Long-term Equity AnticiPation Security — more than a year. D-19: 860 structures were published as LEAPS at a maximum of 164 days. | 2026-08-04 |
| options.gate options_generate:GATE | [object Object] | Applied BEFORE legs are combined, so nothing untradeable can reach a structure. Of 112,968 listed contracts, 11,981 pass (10.6%). | 2026-08-04 |
| options.bounds options_generate:BOUNDS | [object Object] | Keeps the estate-wide search tractable. The published set is the best of a deliberately small pool, not the best available in the chain. | 2026-08-04 |
| options.max_house_weight options_distribution:MAX_HOUSE_WEIGHT | 0.25 | AM-012. Bounds how much the house view may move an option's expected value. | 2026-08-04 |
| options.max_tenor_extrapolation options_distribution:MAX_TENOR_EXTRAPOLATION | 1.5 | The engine MC is a one-year distribution. Past 1.5x the house leg is WITHHELD, not clamped, and the blend falls back to the risk-neutral leg. | 2026-08-04 |
| options.v1_rate | 0.04 | Frozen deliberately: 8,236 pre-registered v1 structures were priced with it and golden tests pin the output byte for byte. The v2 path uses the dated CMT curve. | 2026-08-04 |
| scoring.hold_band scoring_core:HOLD_BAND | 5 | A HOLD 'hits' if it tracked its benchmark within this many points. Unified 2026-08-04 — two definitions existed (+/-10 and +/-5) and the tighter one won. | 2026-08-04 |
| scoring.min_days scoring_core:MIN_DAYS | 21 | Below this a forecast has not matured. | 2026-08-04 |
| scoring.anomaly_floor scoring_core:ANOMALY_FLOOR | -85 | A print this bad is a corporate action the series did not adjust for, not an outcome. | 2026-08-04 |
| ledger.max_age_days forecast_ledger:MAX_AGE_DAYS | 10 | AM-014. A summary older than this is history, not a publication — this is what makes 'never backfill' structural rather than a promise. | 2026-08-04 |
| shadow.structural_floor ros_scenario_shadow:STRUCTURAL_FLOOR | 0.2 | The shadow is held to the same mandate as the published view, so a data-driven vector cannot be less conservative than what is published. | 2026-08-04 |
| return_engine.risk_free | 4 | Sets expected alpha and the Sharpe denominator. In ros_config.return_engine. | 2026-08-04 |
| return_engine.horizon_years | 1 | One global horizon for all 854 names. Makes CAGR equal total return by construction. | 2026-08-04 |
| cov.min_history_td ros_cov_model:MIN_HISTORY_TD | 504 | Two years of a name's OWN adjusted closes before it may enter the factor covariance. The optimizer has no risk-aversion lambda, so the 10% volatility ceiling estimated off this covariance is its entire risk story; a name with less history cannot support an estimate a ceiling can lean on. Below the floor the name is EXCLUDED and named, never admitted at a shortened window. | 2026-08-05 |
| cov.lookback_td ros_cov_model:LOOKBACK_TD | 756 | Three years requested. The ACHIEVED window is emitted in meta with the name that caps it, so a shortfall is visible rather than assumed away. | 2026-08-05 |
| cvar.copula_df ros_cvar_diag:COPULA_DF | 5 | Matches the df the engine already registers for the marginals being joined, so the joint and the marginals agree about tail thickness. Not fitted. A Gaussian copula was rejected: zero tail dependence is the wrong assumption in the tail an expected shortfall measures. | 2026-08-05 |
| cvar.draws ros_cvar_diag:DRAWS | 20000 | 20,000 joint draws puts 1,000 observations in the 5% tail. The bootstrap standard error of every CVaR is published alongside it, so the draw count is disclosed by its effect rather than defended in prose. | 2026-08-05 |
| cvar.coverage_floor_pct ros_cvar_diag:COVERAGE_FLOOR_PCT | 95 | Below this share of the book's invested weight the portfolio tail figure is WITHHELD and the uncovered names listed. It is never renormalised onto the covered subset: both renormalising and zero-filling understate the tail, and a tail number quietly too small is worse than no number. | 2026-08-05 |
| attribution.noise_band_k attribution_engine:NOISE_K | 1.5 | Band half-width in idio sigmas. A residual within K*sigma_idio*sqrt(t) is explained-as-noise by the name's own realized vol; ~13-18% of clean rows land outside by chance under fat tails, absorbed by the 25% run cap. Evidence AM-032: 72% shrug (refused) -> 13.6% (emits) same book. | 2026-08-10 |
| attribution.vol_lookback_days attribution_engine:VOL_LOOKBACK | 120 | Trading days of daily log returns ENDING at the registration date — point-in-time, no look-ahead. The band uses vol known when the forecast stood. | 2026-08-10 |
| attribution.vol_min_obs attribution_engine:VOL_MIN_OBS | 60 | Fewer observations than this and vol is NOT computable: the row cannot claim within-noise and stays shrug-eligible — missing input, more conservative. | 2026-08-10 |
| attribution.idio_floor_frac attribution_engine:IDIO_FLOOR_FRAC | 0.25 | sigma_idio >= this fraction of total sigma. Keeps the band from collapsing when beta^2*sigma_spy^2 exceeds total variance; a SMALLER band is the conservative direction. | 2026-08-10 |
| attribution.vol_uncomputable_cap attribution_engine:VOL_UNCOMPUTABLE_CAP | 0.2 | BLOCKING: band uncomputable on more than this share of rows refuses the run — the vol layer is broken and the band must not silently degrade estate-wide. Mutation-certified. | 2026-08-10 |
| attribution.sigma_sanity_ann attribution_engine:SIGMA_SANITY_ANN | 0.05 · 2 | BLOCKING: median annualized sigma_idio outside this range refuses the run — degenerate vol makes the band a carpet or a wall. Mutation-certified. | 2026-08-10 |
| attribution.rev_tol_early_d attribution_engine:REV_TOL_EARLY_D | 3 | E0 snapshot at most this many days BEFORE d0. Earlier would attribute pre-window revisions into the window — the anti-conservative side, hence tighter than the late side. | 2026-08-10 |
| attribution.rev_tol_late_d attribution_engine:REV_TOL_LATE_D | 5 | E0 snapshot at most this many days AFTER d0. A later E0 understates the carve — the conservative side, hence looser than the early side. | 2026-08-10 |
| attribution.rev_rollover attribution_engine:REV_ROLLOVER | 0.3 | |E1/E0 - 1| beyond this is a fiscal-year-rollover signature: the carve is EXACTLY 0 — an unmeasured category never claims anything. Pairs are SAME-SOURCE only (AM-032). | 2026-08-10 |
| attribution.rev_min_e0 attribution_engine:REV_MIN_E0 | 0.05 | |E0| below this makes the revision ratio meaningless: carve is EXACTLY 0. | 2026-08-10 |
| attribution.est_method_boundary attribution_engine:EST_METHOD_BOUNDARY | 2026-08-12 | Estate consensus vintages before this date were built by index-first FY-row selection (the fy[0] bug's fourth surface, fixed in mch_v3_modules 2026-08-11), so pairs spanning it measure re-selection, not revision — the rejected activation attempt read that as +/-25% two-day revisions. Same-source pairs require a method-stable source. | 2026-08-11 |
| behaviour.donchian_n behaviour_features:DONCHIAN_N | 126 | A 126td (half-year) channel is long enough that a fresh break is information rather than noise, short enough to exist for most names; the T3 backtest judges it, and any re-tuning is an amendment, not an edit. | 2026-08-13 |
| behaviour.gap_threshold behaviour_features:GAP_ABS_THRESHOLD | 0.03 | 3% overnight is beyond normal drift for large caps; below it, gap counting becomes a volatility measure in disguise. | 2026-08-13 |
| behaviour.udvr_window behaviour_features:UDVR_WINDOW | 63 | One quarter of participation evidence — the shortest window the 63td outcome horizon can honestly condition on. | 2026-08-13 |
| behaviour.adv_short behaviour_features:ADV_SHORT | 21 | One month of turnover; the liquidity input Entry Quality ranks on. | 2026-08-13 |
| behaviour.adv_long behaviour_features:ADV_LONG | 252 | One year — the base the expansion/contraction flag compares against. | 2026-08-13 |
| behaviour.atr_short behaviour_features:ATR_SHORT | 21 | Paired with atr_long for the compression ratio; a month of true ranges. | 2026-08-13 |
| behaviour.atr_long behaviour_features:ATR_LONG | 252 | A year of true ranges — the denominator that defines 'compressed'. | 2026-08-13 |
| behaviour.rv_pctile_window behaviour_features:RV_PCTILE_WINDOW | 1260 | Five years places today's realised vol within a name's own regime history; shorter windows re-centre on the current regime and stop discriminating. | 2026-08-13 |
| behaviour.min_window_fraction behaviour_features:MIN_FRACTION | 0.9 | Below 90% coverage a 'window' is a different, shorter window wearing the label — the function refuses instead. | 2026-08-13 |
Scheduler health
QA-494: a job with no fresh heartbeat degrades the publish gate visibly — dead automation is never mistaken for a quiet week
| Job | Last beat | OK | Source |
|---|---|---|---|
| daily | 2026-08-24T03:48:37Z | ✓ | scheduled |
| estate | 2026-08-24T03:20:45Z | ✓ | orchestrated |
| marketdata | 2026-08-22T05:30:14Z | ✓ | scheduled |
| snapshots | 2026-08-15T20:43:00Z | ✓ | orchestrated |
Active override — jobs snapshots until 2026-08-25: snapshots ONLY. Extended ONE WEEK 2026-08-16 (from 08-18 to 08-25) — Marinus's explicit decision, with the condition NAMED: the 2026-08-15 weekly beat was 3/5 because (a) the AM-041 behaviour scorecard read 'pre-inception' — stage 2a0 runs BEFORE stage 2k, which birthed the ledger THAT run (858 v2 rows now on EFS+S3), so the next weekly's 2a0 scores it; and (b) iv_history's weekly_live row can only land after the US close and the run fired pre-close. Both units are structurally able to land on the 2026-08-22 weekly. If that run does NOT produce a clean 5/5 beat, that IS evidence of a persistent condition and the gate must close on it — no second extension without a diagnosed cause. History: granted 08-11 (marketdata dropped same day per its drop-condition), first expiry 08-18.
Self-imposed model-governance discipline in the style of SR 11-7. MCH Advisory Services is an independent research publisher, not a registered investment adviser; nothing on this page is investment advice or a regulatory attestation.