First iteration — three OpenAI models ran three different strategies (fundamental, news-driven, and trend-following). It varied strategy as well as model, so it is not a clean model comparison; it is where the benchmark started.

Over its 28-month run, the three models diverged sharply: the fundamental strategy finished ahead at +6.1%, the news strategy roughly flat at +2.8%, and the trend strategy down −19.2%. Yet all three badly trailed a simple S&P 500 buy-and-hold (+48.4%) — active AI decision-making did not beat the index.

The reasoning judge, run across every decision, surfaced why: recurring strategy drift, thesis thrashing (buying and re-selling the same names on opposite rationales), and weak temporal consistency. Because Season 1 varied strategy as well as model, it cannot isolate the model — and that limitation is exactly what shaped Season 2’s controlled, single-prompt design.

Final standings

ModelPortfolio ValueTotal Gain %Total Gain $Total TradesRecent Activity
OpenAI GPT-4 Turbo$106,145.97+6.15%$6,145.971135HOLD
OpenAI GPT-4$102,832.35+2.83%$2,832.351677BUY
OpenAI GPT-3.5$80,753.61-19.25%-$19,246.391438HOLD

Performance vs S&P 500

Over February 26, 2024 → June 26, 2026, a simple S&P 500 buy-and-hold returned +48.4%. None of the models matched it — which is the point of the benchmark: it measures decision quality and reasoning under uncertainty, not whether a model beat the index. Each decision is graded in the reasoning evaluations below.

ModelTotal returnMax drawdownvs S&P 500
OpenAI GPT-4 Turbo+6.1%-17.5%-42.3% pp
OpenAI GPT-4+2.8%-14.4%-45.6% pp
OpenAI GPT-3.5-19.2%-25.5%-67.6% pp

Model Evaluations

A panel of three judges — one per frontier provider — graded each model’s full decision history on reasoning, evidence, and process, on an anonymized record. The Total Score blends their reasoning median (90%) with reasoning efficiency — quality per second of thinking (10%) — reported alongside returns. How it's scored →

ModelReasoningEvidenceOutcomeEfficiencyTotal ScoreVerdict
OpenAI GPT-4 Turbo46576151Mixed / Improving but Inconsistent
OpenAI GPT-443495845Mixed Process, Catalyst-Aware but Undisciplined
OpenAI GPT-3.547402440Mixed Process, Weak Discipline

Open a model to see its full evaluation — score trend, dimension breakdown, and claim ledger.

This run is complete and frozen. Each model's full holdings and decision history remain available from the standings above. All seasons · How models are evaluated