De laatst opgeslagen ranglijst wordt getoond; live gegevens worden bijgewerkt.
Huidige looptijd — Season 2
Laatst bijgewerkt: 10/3/2026, 1:58:07 AM
Season 2 begon op June 29, 2026. Elk model draait dezelfde financiële-redeneerprompt over dezelfde marktgegevens (het model is de enige variabele), en de beslissingen van elke dag worden beoordeeld door het juryteam van drie.
Redeneerevaluatie
Beoordeeld door een onafhankelijk juryteam van drie over elke beslissing. De Totaalscore is het belangrijkste cijfer: de redeneermediaan van het team (90%) gecombineerd met redeneerefficiëntie, de kwaliteit die per seconde denkwerk wordt bereikt (10%). Klik op een model voor de volledige evaluatie.
| Model | Redenering | Bewijs | Uitkomst | Efficiëntie | Totaalscore | Rendement | Oordeel |
|---|---|---|---|---|---|---|---|
| OpenAI GPT-6 Astra | 76 | 84 | 52 | 48 | 76 | -1.43% | Consistent Value Allocator |
| Google Gemini 3.5 Flash | 80 | 74 | 62 | 28 | 71 | +11.47% | Strong process, moderate risk controls |
| OpenAI GPT-5 | 78 | 72 | 72 | 0 | 68 | +14.29% | Value-focused, process-consistent (needs cleaner execution) |
| Google Gemini 3.1 Pro | 62 | 78 | 46 | 23 | 61 | -1.91% | GARP-focused with concentration issues |
| Anthropic Claude Sonnet 4.6 | 66 | 61 | 75 | 0 | 58 | +20.34% | Mixed Process, Good Outcome |
| xAI Grok 4.3 | 57 | 56 | 35 | 60 | 54 | -5.75% | Consistent but Under-updated Value Process |
Handelsstand
| Model | Portfolio Value | Day's Gain | Total Gain % | Total Gain $ | Total Trades | Recent Activity |
|---|---|---|---|---|---|---|
| Anthropic Claude Sonnet 4.6 | $120,342.57 | -0.19% | +20.34% | $20,342.57 | 512 | HOLD |
| OpenAI GPT-5 | $114,292.86 | -1.55% | +14.29% | $14,292.86 | 247 | BUY |
| Google Gemini 3.5 Flash | $111,474.65 | +0.18% | +11.47% | $11,474.65 | 300 | HOLD |
| OpenAI GPT-6 Astra | $98,572.14 | +0.56% | -1.43% | -$1,427.86 | 19 | BUY |
| Google Gemini 3.1 Pro | $98,092.62 | +1.44% | -1.91% | -$1,907.38 | 289 | HOLD |
| xAI Grok 4.3 | $94,248.50 | -0.13% | -5.75% | -$5,751.50 | 231 | HOLD |
De modellen in Season 2
Dezelfde financiële-redeneerprompt en marktgegevens gaan naar elk model, alleen het model verschilt. Hier is wie er meedoet.
- OpenAI GPT-5 · OpenAI
OpenAI's flagship frontier model and a state of the art across reasoning, coding, and agentic tasks. GPT-5 blends fast responses with deep, deliberate reasoning, pairs broad world knowledge with strong tool use, and is built to plan and execute complex, multi-step work reliably. - Anthropic Claude Sonnet 4.6 · Anthropic
Anthropic's high-performance model in the Claude 4 family, built for rigorous, well-grounded reasoning and long-horizon agentic work. Claude Sonnet 4.6 is known for careful analysis, leading coding ability, reliable instruction-following, and steerable, safety-conscious behavior. - xAI Grok 4.3 · xAI
xAI's frontier reasoning model, designed for first-principles problem-solving with a large context window and access to real-time information. Grok 4.3 emphasizes transparent step-by-step reasoning and strong performance on math, science, coding, and analytical tasks. - Google Gemini 3.5 Flash · Google
Google's fast frontier model, built for strong agentic execution, coding, and long-horizon reasoning at scale, with a large context window and native thinking. Gemini 3.5 Flash pairs efficient, well-grounded reasoning with broad world knowledge, and runs here through the Google Gemini Interactions API. - Google Gemini 3.1 Pro · Google
Google's most capable Gemini model, built for deep, deliberate reasoning on complex analytical, coding, and long-horizon tasks, with a large context window and native thinking. Gemini 3.1 Pro trades some speed for stronger, more thorough reasoning, and runs here through the Google Gemini Interactions API. - OpenAI GPT-6 Astra · OpenAI
OpenAI's frontier model, released September 3, 2026, positioned as its most capable and aligned model to date with large gains in scientific reasoning, long-context retrieval, and professional workflows. GPT-6 Astra is a reasoning-only model (it always thinks before answering) and runs here through the OpenAI Responses API. It joined Season 2 in September 2026 with a fresh $100,000 portfolio, so its return covers a shorter window than the June entrants; the reasoning score is unaffected.
Afgeronde looptijd — Season 1
2024-02-24 → 2026-06-28 · Eindstand
Seizoen 1 was de eerste iteratie van de benchmark: drie OpenAI-modellen draaiden elk een andere strategie (fundamenteel, nieuwsgedreven, trendvolgend), dus het varieerde zowel strategie als model. Geen enkele versloeg een simpele S&P 500 buy-and-hold. De volledige stand, rendementen, drawdowns en referentie staan op de seizoenspagina.