First iteration — three OpenAI models ran three different strategies (fundamental, news-driven, and trend-following). It varied strategy as well as model, so it is not a clean model comparison; it is where the benchmark started.
Over its 28-month run, the three models diverged sharply: the fundamental strategy finished ahead at +6.1%, the news strategy roughly flat at +2.8%, and the trend strategy down −19.2%. Yet all three badly trailed a simple S&P 500 buy-and-hold (+48.4%) — active AI decision-making did not beat the index.
The reasoning judge, run across every decision, surfaced why: recurring strategy drift, thesis thrashing (buying and re-selling the same names on opposite rationales), and weak temporal consistency. Because Season 1 varied strategy as well as model, it cannot isolate the model — and that limitation is exactly what shaped Season 2’s controlled, single-prompt design.
Final standings
| Model | Portfolio Value | Total Gain % | Total Gain $ | Total Trades | Recent Activity |
|---|---|---|---|---|---|
| OpenAI GPT-4 Turbo | $106,145.97 | +6.15% | $6,145.97 | 1135 | HOLD |
| OpenAI GPT-4 | $102,832.35 | +2.83% | $2,832.35 | 1677 | BUY |
| OpenAI GPT-3.5 | $80,753.61 | -19.25% | -$19,246.39 | 1438 | HOLD |
Performance vs S&P 500
Over February 26, 2024 → June 26, 2026, a simple S&P 500 buy-and-hold returned +48.4%. None of the models matched it — which is the point of the benchmark: it measures decision quality and reasoning under uncertainty, not whether a model beat the index. Each decision is graded in the reasoning evaluations below.
| Model | Total return | Max drawdown | vs S&P 500 |
|---|---|---|---|
| OpenAI GPT-4 Turbo | +6.1% | -17.5% | -42.3% pp |
| OpenAI GPT-4 | +2.8% | -14.4% | -45.6% pp |
| OpenAI GPT-3.5 | -19.2% | -25.5% | -67.6% pp |
Model Evaluations
A panel of three judges — one per frontier provider — graded each model’s full decision history on reasoning, evidence, and process, on an anonymized record. The Total Score blends their reasoning median (90%) with reasoning efficiency — quality per second of thinking (10%) — reported alongside returns. How it's scored →
| Model | Reasoning | Evidence | Outcome | Efficiency | Total Score | Verdict |
|---|---|---|---|---|---|---|
| OpenAI GPT-4 Turbo | 46 | 57 | 61 | — | 51 | Mixed / Improving but Inconsistent |
| OpenAI GPT-4 | 43 | 49 | 58 | — | 45 | Mixed Process, Catalyst-Aware but Undisciplined |
| OpenAI GPT-3.5 | 47 | 40 | 24 | — | 40 | Mixed Process, Weak Discipline |
Open a model to see its full evaluation — score trend, dimension breakdown, and claim ledger.
This run is complete and frozen. Each model's full holdings and decision history remain available from the standings above. All seasons · How models are evaluated