A frontier AI model (Claude’s most advanced) runs two paper-trading accounts in public — one making fast options bets, one investing off a written thesis on the AI economy. This strip is the scoreboard; the plain record is right below.
The experiment: two paper accounts are run by Fable, the most advanced Claude model available over an API, both under a daily-trade mandate (each must act at least once a session — so a trade here is not automatically conviction). Fable-aggressive (FA) makes short-dated options bets (cheap bets that pay off big if the market lurches around scheduled events like a CPI print or a Fed meeting) inside caps enforced by code, not prompt. Fable-economist (FE) keeps a written thesis journal on the AI economy — with web search as its eyes — and buys the companies it argues are the leverage points. Everything else on this page is a reference line to beat: a mechanical control strategy, a market-state overlay, and SPY buy-and-hold.
Where it came from: this plate used to be MERIDIAN — five specialist LLMs deliberating over filings and news sentiment. It produced sophisticated reasoning and, in two months, exactly one trade: an analysis engine with no defined edge. SOLITON is the rebuild (MERIDIAN is kept as an archive). Every strategy here is a mechanical, pre-registered rule set, backtested before it touches even paper money, and run in public.
Success was defined before launch: beat the controls over 100+ logged trades with every cap respected — not “up in week one.” Paper money throughout.
Plain-language key — what the labels on this page mean
paper money
virtual $100k (or a smaller ring-fenced bankroll) traded against real prices — no real money moves.
daily-trade mandate
each Fable account must take at least one action every session, so some trades are obligation, not conviction — those are flagged.
mandate-forced
the model itself flagged a trade as taken only to satisfy that rule.
IV rank (0–100)
how expensive options insurance is right now versus the past year — high means pricey.
put spread / iron condor
defined-risk options bets: you know the most you can lose up front.
evidence label
each track wears its honest status: control = a yardstick, not a bet; negative / unproven = the backtest found no edge; shadow = tracked for signal only, no money.
SPY buy-and-hold
just holding the S&P 500 — the do-nothing baseline every strategy has to beat.
How to read this: two paper accounts, one Fable API call each per session. FA — fable-aggressive — buys short-dated call and put spreads on SPY or QQQ off the macro calendar, from a $5,000 ring-fenced bankroll with per-punt caps and loss stops enforced by code. FE — fable-economist — keeps a running thesis journal on the AI economy and buys the leverage points it finds (long-only equities, $10,000 bankroll). The model decides; the code builds every order, checks it against the caps, and refuses anything outside them — refusals are published too.
Both accounts run under a daily-trade mandate: every session, each must take at least one position action. That changes how you read the record — a trade here is not automatically conviction. The model flags trades it took only to satisfy the mandate and the digest labels them; safety halts beat the mandate, so a stopped account simply doesn’t trade.
What would count as success was defined before launch: beat the mechanical control and SPY buy-and-hold over 100+ logged trades with every cap respected. One green day proves nothing. The record below is the product — forced trades, refusals, and all.
Day 4 — Jul 31
FA’s session ended no_signal_stand_aside; flat on the day, flat since launch.
FE logged no session.
Track A stayed flat (IV rank 72.6, below its 85.0 gate). SPY +0.7%.
Day 3 — Jul 30
FA’s session ended no_signal_stand_aside; flat on the day, flat since launch.
FE opened ETN on its ‘t-002’ thesis, voluntarily; +$195.60 on the day, up $115.44 since launch.
“Web search was unavailable this session, so no conviction changes — I trade only on the standing journal. Regime is CALM on both SPY and QQQ with LPPL at zero, but QQQ IV rank at 82.5 argues against chasing extended…”→ full decision record
Track A stayed flat (IV rank 82.5, below its 85.0 gate). SPY +1.7%.
Day 2 — Jul 29
FA reached for a call spread on SPY ($400.00 at risk), to satisfy the mandate — but the caps code bounced it: “orders[0](call_spread SPY): budget 400.00 below one contract (~475.25; per-punt cap 500.00, cash 5000.00)”. No position, so the mandate went unmet; flat on the day, flat since launch.
“Both underlyings show implied vol well above realized (SPY IV proxy ~20.7 vs 20d realized ~12.7; QQQ IV rank 97 with IV ~30.8 vs RV ~23.5), which means buying premium — the only thing this account can do — is…”→ full decision record
FE opened MU on its ‘t-001’ thesis, voluntarily; −$80.16 on the day, down $80.16 since launch.
“Web search was unavailable this session, so I made no conviction changes — moving a conviction score without new evidence would be manufacturing signal. The actionable gap in the book was that t-001 (memory-supercycle)…”→ full decision record
Track A fired: open iron condor (IV rank 97.2, above its 85.0 gate). SPY −1.5%.
Day 1 — Jul 28
FA reached for a call spread on SPY ($300.00 at risk), to satisfy the mandate — but the caps code bounced it: “orders[0](call_spread SPY): budget 300.00 below one contract (~336.25; per-punt cap 500.00, cash 5000.00)”. No position, so the mandate went unmet; flat since launch — first session on the books.
“FOMC lands tomorrow at 14:00 ET and implied vol is priced rich for it — SPY IV rank 64.7 with IV ~18.2 vs 20-day realized ~12.4, QQQ IV rank 90.5. That gap means outright long options are paying for a move the market…”→ full decision record
FE opened VST and GEV on its ‘t-002’ thesis, voluntarily (1 order didn’t clear the caps); flat since launch — first session on the books.
“First deployment of a fresh account: I'm opening conviction-scaled starter positions in my two actionable structural theses rather than waiting, because the regime is benign (dynamics battery CALM on SPY and QQQ, both…”→ full decision record
Track A fired: open iron condor (IV rank 90.5, above its 85.0 gate). SPY: first close on the record.
The race — the same accounts as one chart, % return since launch
ATrack A — mechanical premium selling (S3 champion cell, public control) — negative — launched as public control per brief §7.5
CTrack C — dynamical-state overlay (C2 registered default cell) — unproven — launched for public forward test per brief §7.5
FAfable-aggressive — fa_doctrine_v1: a planning-agent-authored doctrine, executed tactically by claude-fable-5 (FA-2) — designed — event leg carries D1 registered evidence; trend/vol legs and the composite are unproven; public forward test
FEfable-economist — claude-fable-5 thesis-driven AI-economy account (the headline experiment) — unproven — LLM thesis-driven account under a disclosed daily-trade mandate; forced trades flagged mandate_forced
SPYSPY buy-and-hold — the do-nothing baseline every track has to justify itself against
as of 2026-07-31 · exported Jul 31 22:14Z · live export
Decision log — every session, verbatim
The engine journals every decision it makes — and every session it refuses to trade, with the reason. The Fable accounts’ one-paragraph rationales appear here word for word, next to the orders the code actually built and every cap check they failed. Nothing is edited after the fact.
halted — stale_data: no ohlcv for ETN (DataValidationError: ETN: inconsistent OHLC bar (high not the max / low not the min), first at 2026-07-31) — cannot mark FE's book
Agate_evaluation
IV rank 72.6 vs gate 85.0 · 2 open → no_trade
“gate_closed”
Jul 30
FEfable_decision
→ accepted · voluntary
Web search was unavailable this session, so no conviction changes — I trade only on the standing journal. Regime is CALM on both SPY and QQQ with LPPL at zero, but QQQ IV rank at 82.5 argues against chasing extended…
Web search was unavailable this session, so no conviction changes — I trade only on the standing journal. Regime is CALM on both SPY and QQQ with LPPL at zero, but QQQ IV rank at 82.5 argues against chasing extended tech; MU marked ~+10% on day one and I will not add into a spike. Instead I continue staged deployment of my ~73% cash into the power-bottleneck thesis (t-002, conviction 3), which I currently express only through the generator leg (VST) and turbines (GEV). The thesis names a third, policy-insensitive leg — grid hardware with multi-year backlogs — so I open ~$900 of ETN to complete the structure: generators selling scarce firm capacity, turbine supply, and the electrical-equipment layer whose orders don't depend on which utility wins the rate-case fight. This concentrates the book on one causal chain, and I accept that knowingly: the shared falsifier is hyperscaler capex cuts or datacenter cancellations beyond Goldman's ~50-60% realization haircut, which would hit all three legs at once — that is the first thing I will check when search access returns.
Web search was unavailable this session, so I made no conviction changes — moving a conviction score without new evidence would be manufacturing signal. The actionable gap in the book was that t-001 (memory-supercycle)…
Web search was unavailable this session, so I made no conviction changes — moving a conviction score without new evidence would be manufacturing signal. The actionable gap in the book was that t-001 (memory-supercycle) carried zero expression despite being my original structural thesis: HBM demand crowding out conventional DRAM wafer capacity (~3:1 conversion), shortage persisting into 2027 because new fabs don't ship until then, and Micron's revenue base shifting from spot-cyclical to contracted (17 strategic agreements, some take-or-pay, HBM sold out through 2026). I opened a $750 starter in MU — half the per-position cap, sized to conviction 2 — rather than adding to VST/GEV, which are one session old and down ~4% on pure noise with nothing thesis-relevant in the packet. Regime is CALM, both indices above their 200-day, calendar empty until Aug 12 CPI, so there is no regime reason to delay initiating. What proves me wrong on this position: DRAM contract prices rolling negative QoQ, hyperscaler capex cuts, Samsung/SK Hynix adding capacity faster than expected, or strategic-agreement cancellations — any of those and I cut rather than average down. This is a genuine thesis expression, not a mandate trade.
Both underlyings show implied vol well above realized (SPY IV proxy ~20.7 vs 20d realized ~12.7; QQQ IV rank 97 with IV ~30.8 vs RV ~23.5), which means buying premium — the only thing this account can do — is…
Both underlyings show implied vol well above realized (SPY IV proxy ~20.7 vs 20d realized ~12.7; QQQ IV rank 97 with IV ~30.8 vs RV ~23.5), which means buying premium — the only thing this account can do — is structurally expensive today, so I have no genuine long-premium edge and I flag this as mandate-forced. Given I must trade, the most defensible expression is the one that minimizes what I'm overpaying for: an 8-DTE ATM debit call spread on SPY (the cheaper-vol underlying), long the 729 call and short the 739, expiring 2026-08-06. The causal claim is modest: the regime battery reads CALM, SPY sits ~4.7% above its 200-day with positive trend, the VIX curve is in mild contango, and the calendar is empty until the Aug 12 CPI — after this position expires — so a quiet drift higher is the base case and the short 739 leg sells back a chunk of the inflated vol I'm forced to buy. Risk is capped at the ~$400 debit; max payoff $1,000 if SPY closes above 739.
“orders[0](call_spread SPY): budget 400.00 below one contract (~475.25; per-punt cap 500.00, cash 5000.00)”
claude-fable-5 · prompt fa_v1 · call cost $0.00
Agate_evaluation
IV rank 97.2 vs gate 85.0 · 1 open → open_iron_condor
First deployment of a fresh account: I'm opening conviction-scaled starter positions in my two actionable structural theses rather than waiting, because the regime is benign (dynamics battery CALM on SPY and QQQ, both…
First deployment of a fresh account: I'm opening conviction-scaled starter positions in my two actionable structural theses rather than waiting, because the regime is benign (dynamics battery CALM on SPY and QQQ, both ~5-6% over their 200-day averages, VIX term structure in normal contango at 0.92) and my theses are about multi-quarter supply constraints, not this week's tape. Power-bottleneck (t-002, conviction 3) gets the larger allocation, split across its two legs: VST $1,000 for unregulated generation selling scarce firm capacity directly to hyperscalers, and GEV $1,000 for the policy-insensitive turbine/grid-equipment backlog. Memory-supercycle (t-001, conviction 2) gets a $750 MU starter on the re-rating of contracted HBM revenue from cycle-peak to backlog. Sizing is deliberately partial — ~27% deployed, each name at roughly half to two-thirds of its $1,500 cap — because these fills land at tomorrow's open, hours before the FOMC statement; a hawkish surprise would hit rate-sensitive utilities and long-duration capex names hardest, and I want dry powder to add into any post-FOMC dislocation rather than be fully invested through a binary event. Web search was unavailable this session, so conviction levels are unchanged. What would prove me wrong: DRAM contract prices turning negative QoQ or SCA cancellations (t-001); hyperscaler capex cuts, gas-turbine order cancellations, or datacenter project attrition beyond Goldman's ~50-60% realization haircut (t-002). These are genuine thesis expressions, not mandate filler.
buy VST · 6 sh · ~$891.84 · ‘t-002’
buy GEV · 1 sh · ~$943.38 · ‘t-002’
“orders[2](buy MU): sized to 0 shares (requested 750.00, cap room 1500.00, cash 8164.78, close 820.53)”
FOMC lands tomorrow at 14:00 ET and implied vol is priced rich for it — SPY IV rank 64.7 with IV ~18.2 vs 20-day realized ~12.4, QQQ IV rank 90.5. That gap means outright long options are paying for a move the market…
FOMC lands tomorrow at 14:00 ET and implied vol is priced rich for it — SPY IV rank 64.7 with IV ~18.2 vs 20-day realized ~12.4, QQQ IV rank 90.5. That gap means outright long options are paying for a move the market rarely delivers, and post-statement IV crush would bleed any naked long even if direction is right. Absent the daily mandate I would stand aside into a binary event with no directional edge. Given I must trade, I'm expressing the least-bad version of the prevailing regime: both indices are 5-6% above their 200-day averages, the dynamics battery reads CALM on both, and VIX term structure is in normal contango, so the base case is drift-up continuation once the event passes. A 3-DTE SPY 741/748 debit call spread is roughly vega-flat — the short 748 leg sells back much of the inflated premium — so it wins on modest upside follow-through and loses a small defined amount (~$300, 6% of bankroll) if the Fed surprises hawkish. This is a mandate-driven trade sized and structured accordingly, not a conviction bet.
“orders[0](call_spread SPY): budget 300.00 below one contract (~336.25; per-punt cap 500.00, cash 5000.00)”
One panel per track: paper equity, its win/loss record, open positions, and — kept in plain sight — the honest evidence label and any safety halt. The two Fable accounts are the experiment; the rest are the reference lines they’re measured against.
Open the 5 track panels (equity · records · positions · alarms)
A
Track A — mechanical premium selling (S3 champion cell, public control)
negative — launched as public control per brief §7.5
armed · QQQ
paper equity$99,790 (−0.2%)
recordno closed trades yet
IV rank72.6 / gate 85 — below gate
IV rank = how expensive options insurance is right now (0–100 vs the past year); this control only sells above its gate.
Track C — dynamical-state overlay (C2 registered default cell)
unproven — launched for public forward test per brief §7.5
armed · SPY
paper equity$101,085 (+1.1%)
recordno closed trades yet
regimeCALM
ts 0.84d200 0.07lppls 0.00
open
134 SPY @ 738.930075
CS2
Track C shadow — C2 nearest-miss QQQ|l0.2|d2|w0.5 (signals only)
shadow — signals only, no capital (C2 report recommendation #3)
armed · QQQ · signals only, no capital
regimeCALM
ts 0.84d200 0.07lppls 0.00
FA
fable-aggressive — fa_doctrine_v1: a planning-agent-authored doctrine, executed tactically by claude-fable-5 (FA-2)
designed — event leg carries D1 registered evidence; trend/vol legs and the composite are unproven; public forward test
armed · QQQ
paper equity$5,000 (0.0%)
recordno closed trades yet
FE
fable-economist — claude-fable-5 thesis-driven AI-economy account (the headline experiment)
unproven — LLM thesis-driven account under a disclosed daily-trade mandate; forced trades flagged mandate_forced
halted · SPY
stale_data: no ohlcv for ETN (DataValidationError: ETN: inconsistent OHLC bar (high not the max / low not the min), first at 2026-07-31) — cannot mark FE's book
paper equity$10,115 (+1.2%)
recordno closed trades yet
open
6 VST @ 148.986667
1 GEV @ 943.38
1 MU @ 795.62
alarm history
Jul 31stale_data: no ohlcv for ETN (DataValidationError: ETN: inconsistent OHLC bar (high not the max / low not the min), first at 2026-07-31) — cannot mark FE's book
Methodology — what’s proven, what isn’t
The backtests did not find an edge — and publishing that is the point. The premium-selling family was pre-registered and tested twice: first on synthetic implied-vol chains (no edge demonstrated either way), then re-run on real 2016–2026 option chains (what edge existed was consumed by friction — spreads, slippage, assignment).
The tracks launched anyway, as a public forward test, with the evidence status printed verbatim on every panel above. If a line goes up, that’s data, not vindication.
Paper money
Every dollar on this page is a virtual sub-ledger: each track runs $100k of paper capital against real market data. No real money moves. The published bundle carries no account identifiers, order ids, or keys — the exporter refuses to write a file that does.
Cost model
Fills are modeled against real end-of-day chains with commissions and slippage charged on every leg. Assignment friction — the gap between modeled settlement and what physical settlement would have cost — is logged as a first-class statistic in the trade logs, because that’s where paper results quietly diverge from reality.
Pre-registration
Entry rules, parameter grids, exits, and kill criteria are frozen in versioned specs before the data runs, and the kill criteria are code, not judgment — a track that trips one halts itself and says so on its panel. The Fable accounts can exercise judgment inside those limits; they cannot rewrite them.
The record
The engine, the backtests, and the verdicts live in the source repo — private, since it’s a live trading system — negative results included. This page renders the engine’s own export bundle verbatim — same file, same labels. The public artifact is the design story: the MERIDIAN lesson, the data saga, the three verdicts, why each track exists.