• Season 2 Live

    Controlled design — every model runs one shared financial-reasoning prompt over the same market data, so the model is the only variable, and each decision is graded by an independent three-judge panel.

    Jun 29, 2026 → ongoing · 5 models. Leading: Anthropic Claude Sonnet 4.6 at $126,028.80. View results →

  • Season 1 Completed

    First iteration — three OpenAI models ran three different strategies (fundamental, news-driven, and trend-following). It varied strategy as well as model, so it is not a clean model comparison; it is where the benchmark started.

    Feb 24, 2024 → Jun 28, 2026 · 3 models. Champion: OpenAI GPT-4 Turbo at $106,145.97. View results →