Testing Models: Compare AI Models Side by Side, Live
THE AI MODEL ARENA<br>One prompt. Every model. Judge it yourself.<br>Real outputs running live in your browser — not screenshots, not cherry-picked demos. Judge blind, vote, and the community tally is public.<br>Open the coding arenaRead the method<br>54 challenges · 139 model variants across 23 families · 4,414 live artifacts · $436 estimated output spend, every prompt published. Counted at build, 2026-07-25.
All challenges<br>Every brief we have run, ranked by community votes. Click a row to judge it yourself — blind.<br>TaskCategoryModelsLeaderVotesEst. cost range1 · Landing · Driftwood web & tools139…0$0.00–1.102 · Game · NEON BREAKER arcade & games133…0$0.00–0.793 · Tool · Palette Studio web & tools135…0$0.00–1.404 · Dashboard · Bean There web & tools132…0$0.00–0.905 · Particles · Gravity Lab simulation127…0$0.00–0.666 · 3D Game · STARDRIFT 3d & godot120…0$0.00–1.777 · CHIP-8 · CHIP-8 emulator web & tools123…0$0.00–0.528 · News · The Meridian Wire web & tools124…0$0.00–1.359 · Marble 3D · godforge 3d & godot2…0$0.00–0.1910 · Runner 3D · godforge 3d & godot2…0—11 · FPS 3D · godforge 3d & godot2…0—12 · Digger · DEEP HAUL arcade & games123…0$0.00–0.95Show all 54 challenges ↓
Six arenas, one method<br>Coding and writing are live and interactive right now. Images, videos, voice, and music launch as deep evaluation guides while their side-by-side arenas come online.<br>Coding<br>Live now<br>One brief, every model, real code running live — landing pages, arcade games, 3D, emulators.<br>43 tasks · 139 model variants<br>Writing<br>Live now<br>One brief, every model, complete written pieces side by side — essays, stories, songs, satire.<br>11 briefs · blind judging<br>Images<br>Guide<br>How to judge AI image generators — prompt adherence, text, hands, style range, artifacts.<br>evaluation guide · arena in production<br>Videos<br>Guide<br>What separates good text-to-video from bad — motion, temporal consistency, prompt control.<br>evaluation guide · arena in production<br>Voice<br>Guide<br>How to judge AI voice models — prosody on long form, emotional control, hard pronunciations.<br>rubric live · audio arena next<br>Music<br>Guide<br>How to judge AI music generators — composition, fidelity, vocals, structure, prompt adherence.<br>first takes, not curated demos
How the arena works<br>01<br>One prompt, every model<br>We write a single brief and hand the exact same words to every variant. No per-model tuning, no quiet retries to make a favorite look good. The full prompt is published on every challenge.
02<br>Real outputs, running live<br>Each answer runs in your browser as a live artifact, an actual playable game or interactive page, not a screenshot or a marketing clip. Token and cost estimates sit next to every result.
03<br>Blind judging, public votes<br>You start blind: labels hidden, panes shuffled. Vote and the reveal shows who you picked — and whether the crowd agrees. Every vote rolls into a public community tally.
Full methodology, cost math and changelog: /method
The models on the stand<br>23 families and 139 variants, most at several thinking-effort levels — so you can see what extra reasoning actually buys.<br>Fable 5Opus 4.8Opus 5Sonnet 4.6Sonnet 5Haiku 4.5GLM-5.2GPT-5.5Gemini 3 FlashKimi K2.7 CodeQwen3.7 PlusDeepSeek V4 FlashMiniMax M3Mistral Large 2512Grok 4.3Seed 2.1 ProLongCat 2.0KAT-Coder-Pro V2MiMo-V2.5-ProHunyuan Hy3Nemotron 3 Ultragpt-oss-120bGemma 4 26B-A4BLlama 4 MaverickStep 3.7 FlashMuse Spark 1.1
Pick a challenge. Judge for yourself.<br>54 challenges, 139 model variants, and real outputs you can poke at. Go blind, compare, and add your vote.<br>Open the coding arena
🐛 Report a bug