Ante Terminal Bench 2.1 Results | Coding Agent Benchmark
Skip to main content<br>We just open sourced a tiny GPT-style cognitive core built in pure Rust.See our repository→
Terminal-Bench 2.1<br>One harness to unlock the potential of all models<br>Compare Ante runs across models on the same Terminal-Bench 2.1 task set, using consistent parameters and verified benchmark results.
Best Accuracy82.7%
Leading ModelDeepSeek V4 Flash 0731 max
Task Set89 tasks<br>Trials368 passed / 445 trials
Benchmark principlesRead more →<br>We benchmark what we ship. Every eval uses a pinned public Ante release. No eval-only branches or benchmark-specific prompts.<br>The runs are auditable. Every result links its raw Harbor run, so anyone can inspect the trials behind the number.<br>The constraints are official. All trials follow the official Terminal-Bench parameters: 89 tasks, 5 trials per task, strict timeouts, and hardware limits.
Model orgAllDeepSeekMiniMaxQwenxAIXiaomiZ.AI
All model orgsTB 2.1 · Same parameters: 89 tasks · 5 trials/task · Updated Aug 9, 2026<br>#ModelAccuracy↓Cost ($)↕Duration↕ⓘSame-modelAgentⓘSourceRun date↕1DeepSeek V4 Flash 0731 max
82.7%<br>±1.79 SE<br>$68.4138.9 min#1 same-modelAnte 0.preview.710.preview.71<br>Harbor linkAug 9, 2026›2Grok 4.5 mediumⓘReward hacking was identified in 20 trajectories. The displayed score excludes them; see PR #129 in Source for details.
80.9%<br>±1.27 SE<br>$242.578.4 min#1 same-modelAnte 0.preview.560.preview.56<br>PR #129Jul 10, 2026›3GLM 5.2
74.6%<br>±2.06 SE<br>$260.1111.3 min#1 same-modelAnte 0.preview.430.preview.43<br>Harbor linkJun 20, 2026›4DeepSeek V4 Pro
69.1%<br>±2.24 SE<br>$26.3448.4 min#1 same-modelAnte 0.preview.540.preview.54<br>Harbor linkJul 7, 2026›5DeepSeek V4 Flash
66.4%<br>±2.27 SE<br>$49.9841.4 minNo public rowsAnte 0.preview.530.preview.53<br>Harbor linkJul 5, 2026›6MiMo V2.5
65.8%<br>±2.30 SE<br>$73.7612.5 min#1 same-modelAnte 20260625-0824-e383a9220260625-0824-e383a92<br>Harbor linkJun 25, 2026›7MiniMax M3
62.1%<br>±2.33 SE<br>$121.0013.9 min#1 same-modelAnte 20260623-0825-ff174ee20260623-0825-ff174ee<br>Harbor linkJun 23, 2026›8Qwen3.6 27B
56.2%<br>±2.36 SE<br>Local61.6 minNo public rowsAnte 20260701-0837-70d2aac20260701-0837-70d2aac<br>Harbor linkJul 3, 2026›
Terminal-Bench Reference<br>Verified Public Leaderboard
Same parameters: 89 tasks · 5 trials/task · Updated Jul 17, 2026 · Source: Terminal-Bench official verified rowsFor how different models perform on TB 2.1, see Vals AI's Terminal-Bench 2.1 benchmark.
17 official rows<br>#AgentModelAccuracyRun date1Claude Code Anthropic<br>Fable 5 xhigh
83.8%<br>±1.16 SE<br>Jun 7, 2026›2DeepSeek V4 Flash 0731Grok 4.5Ante + DeepSeek V4 Flash 0731 model · 82.7%Ante + Grok 4.5 model · 80.9%would slot between public #2 and #3
Codex OpenAI<br>GPT-5.5 xhigh
83.2%<br>±1.13 SE<br>May 1, 2026›3Terminus 2 Terminal-Bench<br>Fable 5 high
80.5%<br>±1.16 SE<br>Jun 5, 2026›4Cursor CLI Cursor<br>Grok 4.5 high
79.3%<br>±1.46 SE<br>Jul 9, 2026›5Claude Code Anthropic<br>Opus 4.8 high
78.9%<br>±1.31 SE<br>Jul 9, 2026›6Codex OpenAI<br>GPT-5.6 Terra max
78.4%<br>±1.25 SE<br>Jul 11, 2026›7Terminus 2 Terminal-Bench<br>GPT-5.5 xhigh
78.0%<br>±1.22 SE<br>May 1, 2026›8mini-SWE-agent Princeton<br>Muse Spark 1.1 xhigh
76.2%<br>±1.23 SE<br>Jul 9, 2026›9Codex OpenAI<br>GPT-5.6 Luna max
75.7%<br>±1.32 SE<br>Jul 11, 2026›10GLM 5.2Ante + GLM 5.2 model · 74.6%would slot between public #10 and #11
Claude Code Anthropic<br>Sonnet 5 high
74.6%<br>±1.64 SE<br>Jul 9, 2026›11DeepSeek V4 ProAnte + DeepSeek V4 Pro model · 69.1%would slot between public #11 and #12
Terminus 2 Terminal-Bench<br>Gemini 3 Pro high
73.9%<br>±1.29 SE<br>May 1, 2026›12DeepSeek V4 FlashAnte + DeepSeek V4 Flash model · 66.4%would slot between public #12 and #13
Claude Code Anthropic<br>Opus 4.7 max
68.9%<br>±1.41 SE<br>May 1, 2026›13Terminus 2 Terminal-Bench<br>Opus 4.7 max
66.1%<br>±1.37 SE<br>May 1, 2026›14Gemini CLI Google<br>Gemini 3 Pro high
65.8%<br>±1.38 SE<br>May 1, 2026›15MiMo V2.5Ante + MiMo V2.5 model · 65.8%would slot between public #15 and #16
Gemini CLI Google<br>Gemini 3.1 Pro high
65.8%<br>±1.67 SE<br>May 5, 2026›16MiniMax M3Ante + MiniMax M3 model · 62.1%would slot between public #16 and #17
Terminus 2 Terminal-Bench<br>Gemini 3.1 Pro high
65.6%<br>±1.65 SE<br>May 5, 2026›17Qwen3.6 27BAnte + Qwen3.6 27B model · 56.2%would slot below public #17
Claude Code Anthropic<br>GLM-5.1 max
58.6%<br>±1.24 SE<br>May 1, 2026›