Is Claude Code a bad harness? - by William Sun - generray
generray
SubscribeSign in
Is Claude Code a bad harness?<br>There aren’t many harness vs. harness benchmarks, but they don’t look good...
William Sun<br>Aug 20, 2026
Share
Disclaimer: This article was written and edited without the use of LLMs.<br>Last month, Databricks showed off some of their internal measurements on AI agent performance1. They found that Opus 4.8 on Claude Code was slightly better than Opus 4.8 on pi, but more than twice the price.
This goes against the conventional wisdom that there is a synergy from Anthropic developing both the Claude models and the Claude Code/Cowork harnesses. Theoretically, Anthropic should be able to tune the combination so that it’s an optimal mix of performance, latency and cost. No other harness should be beating Claude Code, especially one as simple as pi, right?<br>Harness vs. harness benchmarks are rare. Most benchmarks use mini-swe-agent as a neutral standard, or only look at OpenAI/Anthropic models on their respective company’s harnesses. A more rigorous, though older, comparison is from the TerminalBench 2.0 paper2:
A pretty difficult chart to read, but the results indicate that Claude Code is generally a little less performant and more expensive for the non-Haiku models.<br>The only other comparison I’ve found is from Artificial Analysis’s agentic coding benchmark3:
Now Claude Code has the worst performance, but the lowest cost. This just makes it all more confusing. Let’s break this down…<br>First, benchmarking agent performance is very difficult. All benchmarks are bad, but you do have to pick a least bad one to judge agents by, lest you fall victim to judging by vibes. As a demonstration of this difficulty: OpenAI criticized SWE-Bench, made SWE-Bench Verified to fix its issues, disavowed SWE-Bench Verified for SWE-Bench Pro, then claimed there were issues with SWE-Bench Pro. Harnesses are also updating all the time, so we’re trying to measure moving targets (not to mention Pi, where the purpose is for each user to write their own set of extensions on top of a minimal base). It’s also possible that Anthropic has some internal benchmark that they’ve been hillclimbing against and beating other harnesses in.<br>Second, is the purpose of the Claude Code harness actually to maximize one-shot, long horizon performance on software engineering tasks? Actually, I think Claude Code is mostly an effort by Anthropic to add differentiation at the application layer. In a world where switching a model is two clicks away in a model-agnostic harness, and 0 clicks if you use a model router, you really want to lock users in (or give investors the impression that you have locked users in). If we look at recent feature updates, we can see a lot of updates are for improving usability: adding extra configuration settings, making tokenmaxxing more accessible, or building out the long tail of integrations. Stuff like /buddy doesn’t make an ounce of difference to benchmarks, but they do improve the user experience. You could also be cynical and say that Claude Code makes it easier to juice revenue for a company that is in the middle of a funding round, like when the default reasoning setting increased from high to xhigh on Opus 4.7’s release.
So, to defy Betteridge’s Law, I do think that Claude Code is a bad harness for its core audience. Despite the inherent advantages Anthropic should have, there is an absence of external benchmarks that show a meaningful advantage on the Pareto frontier of cost vs. performance, even small. And for a company that’s<br>Supposed to be on the frontier of solving software engineering
Armed with infinite compute on the most advanced pre-release Anthropic models
Competing against open source projects and startups
This is a pretty disappointing showing.<br>… The caveat, of course, is that for consumers on subsidized plans, Claude Code is far cheaper than paying the API’s sticker price, making it easily the best harness for those people to use Anthropic models. And that’s how they get you!<br>1https://www.databricks.com/blog/benchmarking-coding-agents-databricks-multi-million-line-codebase
2https://arxiv.org/pdf/2601.11868#page=7
3https://artificialanalysis.ai/agents/coding-agents#:~:text=Harness%20Comparison,-Artificial%20Analysis%20Coding
Share
Previous
Discussion about this post<br>CommentsRestacks
TopLatestDiscussions
No posts
Ready for more?
Subscribe
© 2026 generray · Privacy ∙ Terms ∙ Collection notice<br>Start your SubstackGet the app<br>Substack is the home for great culture
This site requires JavaScript to run correctly. Please turn on JavaScript or unblock scripts