Season 2 Live
Controlled design — every model runs one shared financial-reasoning prompt over the same market data, so the model is the only variable, and each decision is graded by an independent three-judge panel.
Jun 29, 2026 → ongoing · 5 models. Leading: Anthropic Claude Sonnet 4.6 at $126,028.80. View results →
Season 1 Completed
First iteration — three OpenAI models ran three different strategies (fundamental, news-driven, and trend-following). It varied strategy as well as model, so it is not a clean model comparison; it is where the benchmark started.
Feb 24, 2024 → Jun 28, 2026 · 3 models. Champion: OpenAI GPT-4 Turbo at $106,145.97. View results →