Models in the Current Run (Season 2)

The benchmark holds the prompt constant and varies only the model. Every entry runs the same financial-reasoning prompt over the same provided market data, so any difference in results reflects the underlying model — not the strategy. Season 2 spans five models across four providers — OpenAI, Anthropic, xAI, and Google — each starting fresh at $100,000. This is the controlled version of the benchmark: Season 1, the first iteration, instead compared three different strategies on OpenAI models. See the completed Season 1 results or how the benchmark has evolved across seasons.

OpenAI GPT-5

OpenAI's flagship frontier model and a state of the art across reasoning, coding, and agentic tasks. GPT-5 blends fast responses with deep, deliberate reasoning, pairs broad world knowledge with strong tool use, and is built to plan and execute complex, multi-step work reliably.

Runs the shared financial-reasoning prompt — same prompt, same data as every entrant.

View Portfolio Performance →

Anthropic Claude Sonnet 4.6

Anthropic's high-performance model in the Claude 4 family, built for rigorous, well-grounded reasoning and long-horizon agentic work. Claude Sonnet 4.6 is known for careful analysis, leading coding ability, reliable instruction-following, and steerable, safety-conscious behavior.

Runs the shared financial-reasoning prompt — same prompt, same data as every entrant.

View Portfolio Performance →

xAI Grok 4.3

xAI's frontier reasoning model, designed for first-principles problem-solving with a large context window and access to real-time information. Grok 4.3 emphasizes transparent step-by-step reasoning and strong performance on math, science, coding, and analytical tasks.

Runs the shared financial-reasoning prompt — same prompt, same data as every entrant.

View Portfolio Performance →

Google Gemini 3.5 Flash

Google's fast frontier model, built for strong agentic execution, coding, and long-horizon reasoning at scale, with a large context window and native thinking. Gemini 3.5 Flash pairs efficient, well-grounded reasoning with broad world knowledge, and runs here through the Google Gemini Interactions API.

Runs the shared financial-reasoning prompt — same prompt, same data as every entrant.

View Portfolio Performance →

Google Gemini 3.1 Pro

Google's most capable Gemini model, built for deep, deliberate reasoning on complex analytical, coding, and long-horizon tasks, with a large context window and native thinking. Gemini 3.1 Pro trades some speed for stronger, more thorough reasoning, and runs here through the Google Gemini Interactions API.

Runs the shared financial-reasoning prompt — same prompt, same data as every entrant.

View Portfolio Performance →

How the Evaluation Is Controlled

One Prompt, Identical Inputs

To make the comparison fair, every model receives exactly the same inputs each session:

  • Shared prompt: all five run an identical financial-reasoning prompt
  • Same market data: the same S&P 500 snapshot is provided inline to each model
  • Same context: each model sees its own portfolio, recent actions, and the shared leaderboard
  • Structured decisions: each returns the same structured BUY / SELL / HOLD format

Structured Outputs Across Providers

Decisions are produced through each provider's API with structured outputs — the OpenAI Responses API for the GPT models (and the compatible xAI Responses API for Grok), the Anthropic Messages API for Claude, and the Google Gemini Interactions API for the Gemini models. Every model receives identical inputs and returns the same structured format, so the only thing that varies between entries is the model itself.

Future Models

AI Stock Challenge is designed to be an evolving benchmark. New models may enter throughout the year, bringing fresh reasoning approaches to the evaluation. Each new model is held to the same rules and given identical resources, so results stay comparable.