StateM: Reaching 95.3% on Terminal Bench 2.1

jumploops1 pts0 comments

Paper page - StateM: Reaching 95.3% Raw Accuracy, or a \$15 Frontier Run, on Terminal-Bench 2.1 via Harness Scaling

Log In<br>Sign Up

\n\nHarness scaling, a different way to scale agent performance.On Terminal-Bench 2.1, StateM raises GPT-5.5 xhigh to 92.1%, versus 83.1% reference and GPT-5.6 Sol Ultra at 91.9%. The runbook transfers unchanged to GPT-5.6. With GPT-5.6 Sol xhigh, StateM reaches 95.3% raw accuracy across 445 trials and succeeds on all 89 tasks at least once.\n","updatedAt":"2026-08-18T17:15:51.726Z","author":{"_id":"655452b8432af1b1116394d1","avatarUrl":"/avatars/85860fb3c2d09c9c23e7677d7129cca3.svg","fullname":"Kai Wang","name":"VictorKai1996NUS","type":"user","isPro":false,"isHf":false,"isHfAdmin":false,"isMod":false,"followerCount":10,"isUserFollowing":false}},"numEdits":0,"identifiedLanguage":{"language":"en","probability":0.8062605857849121},"editors":["VictorKai1996NUS"],"editorAvatarUrls":["/avatars/85860fb3c2d09c9c23e7677d7129cca3.svg"],"reactions":[],"isReport":false}},{"id":"6a85089b69393ddd0ab51b25","author":{"_id":"63d3e0e8ff1384ce6c5dd17d","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/1674830754237-63d3e0e8ff1384ce6c5dd17d.jpeg","fullname":"Librarian Bot (Bot)","name":"librarian-bot","type":"user","isPro":false,"isHf":false,"isHfAdmin":false,"isMod":false,"followerCount":378,"isUserFollowing":false},"createdAt":"2026-08-19T01:36:27.000Z","type":"comment","data":{"edited":false,"hidden":false,"latest":{"raw":"This is an automated message from the [Librarian Bot](https://huggingface.co/librarian-bots). I found the following papers similar to this paper. \n\nThe following papers were recommended by the Semantic Scholar API \n\n* [LongHorizon-Harness: Advancing Long-Horizon Agents for Real-World Tasks](https://huggingface.co/papers/2608.01964) (2026)\n* [Argus: A General-Purpose Agentic Reasoning Runtime for Long-Horizon Tasks](https://huggingface.co/papers/2608.05144) (2026)\n* [When and How Context Rot Appears in Coding Agents: A White-Box Study of Agent Skills in Code Auditing](https://huggingface.co/papers/2607.17937) (2026)\n* [Demystifying Agent Skills: Why They Work-Until They Don't](https://huggingface.co/papers/2608.14036) (2026)\n* [MemoHarness: Agent Harnesses That Learn from Experience](https://huggingface.co/papers/2607.14159) (2026)\n* [Control Under Compression: Reliability Frontiers for Tool-Using Agents](https://huggingface.co/papers/2608.01056) (2026)\n* [Evidence-Calibrated Runtime Reconstruction for Agent Skills Across Heterogeneous Coding Agents](https://huggingface.co/papers/2608.08793) (2026)\n\n\n Please give a thumbs up to this comment if you found it helpful!\n\n If you want recommendations for any Paper on Hugging Face checkout [this](https://huggingface.co/spaces/librarian-bots/recommend_similar_papers) Space\n\n You can directly ask Librarian Bot for paper recommendations by tagging it in a comment: `@librarian-bot recommend`","html":"This is an automated message from the Librarian Bot. I found the following papers similar to this paper. \nThe following papers were recommended by the Semantic Scholar API \n\nLongHorizon-Harness: Advancing Long-Horizon Agents for Real-World Tasks (2026)\nArgus: A General-Purpose Agentic Reasoning Runtime for Long-Horizon Tasks (2026)\nWhen and How Context Rot Appears in Coding Agents: A White-Box Study of Agent Skills in Code Auditing (2026)\nDemystifying Agent Skills: Why They Work-Until They Don't (2026)\nMemoHarness: Agent Harnesses That Learn from Experience (2026)\nControl Under Compression: Reliability Frontiers for Tool-Using Agents (2026)\nEvidence-Calibrated Runtime Reconstruction for Agent Skills Across Heterogeneous Coding Agents (2026)\n\n Please give a thumbs up to this comment if you found it helpful!\n If you want recommendations for any Paper on Hugging Face checkout this Space\n You can directly ask Librarian Bot for paper recommendations by tagging it in a comment: @librarian-bot recommend\n","updatedAt":"2026-08-19T01:36:27.774Z","author":{"_id":"63d3e0e8ff1384ce6c5dd17d","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/1674830754237-63d3e0e8ff1384ce6c5dd17d.jpeg","fullname":"Librarian Bot (Bot)","name":"librarian-bot","type":"user","isPro":false,"isHf":false,"isHfAdmin":false,"isMod":false,"followerCount":378,"isUserFollowing":false}},"numEdits":0,"identifiedLanguage":{"language":"en","probability":0.7501848936080933},"editors":["librarian-bot"],"editorAvatarUrls":["https://cdn-avatars.huggingface.co/v1/production/uploads/1674830754237-63d3e0e8ff1384ce6c5dd17d.jpeg"],"reactions":[],"isReport":false}}],"primaryEmailConfirmed":false,"paper":{"id":"2608.15089","authors":[{"_id":"6a848ed0536bdd3bdd48f61f","name":"Ziheng Qin","hidden":false},{"_id":"6a848ed0536bdd3bdd48f620","name":"Yaxin Lu","hidden":false},{"_id":"6a848ed0536bdd3bdd48f621","name":"Zhangyang Atlas Wang","hidden":false},{"_id":"6a848ed0536bdd3bdd48f622","name":"Kai...

false librarian https huggingface papers agent

Related Articles