96.2% on ARC-AGI-3 with Opus 5

rzk1 pts0 comments

Jeremy Berman on X: "I got 96.2% on ARC-AGI-3 with Opus 5, and 99.3% pass@2. The program is basically Claude Code + Opus 5 (high), one action command, and filesystem logs. Almost nothing ARC specific.

https://t.co/NHyibLgade https://t.co/531wS8sxZ0" / X<br>Post

Log inSign up

Post

Jeremy Berman on X: "I got 96.2% on ARC-AGI-3 with Opus 5, and 99.3% pass@2. The program is basically Claude Code + Opus 5 (high), one action command, and filesystem logs. Almost nothing ARC specific.

https://t.co/NHyibLgade https://t.co/531wS8sxZ0"

Jeremy Berman

@jeremyberman

I got 96.2% on ARC-AGI-3 with Opus 5, and 99.3% pass@2. The program is basically Claude Code + Opus 5 (high), one action command, and filesystem logs. Almost nothing ARC specific.

github.com/jerber/arc-code

span:not(:empty)~span:not(:empty)]:before:content-['·'] [&>span:not(:empty)~span:not(:empty)]:before:px-1 [&>span:not(:empty)~span:not(:empty)]:before:shrink-0">8:11 PM · Aug 12, 2026223.8KViews

51<br>84<br>1.1K<br>611

span:not(:empty)~span:not(:empty)]:before:content-['·'] [&>span:not(:empty)~span:not(:empty)]:before:px-1 [&>span:not(:empty)~span:not(:empty)]:before:shrink-0 min-w-0 overflow-hidden">Jeremy Berman

@jeremyberman

Aug 12

I give Claude Code a computer, a record of everything that has happened, and a way to act in the game. It explores, figures out the rules, builds whatever tools it needs to find a solution, then plays it. Every game starts from scratch.

None of the ARC-specific machinery is Show more

111<br>12K

span:not(:empty)~span:not(:empty)]:before:content-['·'] [&>span:not(:empty)~span:not(:empty)]:before:px-1 [&>span:not(:empty)~span:not(:empty)]:before:shrink-0 min-w-0 overflow-hidden">Jeremy Berman

@jeremyberman

Aug 12

In one pass Opus wrote 269 programs (~12,700 lines). It built parsers for all 25 games, searching functions for 23, and game simulators for 9.

It built a different harness for each problem.

The filesystem memory idea came from the excellent PRO-LONG harness: Show more

GitHub - alexisfox7/PRO-LONG: Programmatic memory for long-horizon LLM agents: the harness appends...

From github.com

115<br>9.9K

span:not(:empty)~span:not(:empty)]:before:content-['·'] [&>span:not(:empty)~span:not(:empty)]:before:px-1 [&>span:not(:empty)~span:not(:empty)]:before:shrink-0 min-w-0 overflow-hidden">Jeremy Berman

@jeremyberman

Aug 12

Code execution makes this scalable and cheaper, which is why using this harness is *cheaper* than asking Opus to solve each game directly.

With code, the model can compile its reasoning into a function, run that function thousands of times, and execute whole action sequences. Show more

64<br>7.4K

span:not(:empty)~span:not(:empty)]:before:content-['·'] [&>span:not(:empty)~span:not(:empty)]:before:px-1 [&>span:not(:empty)~span:not(:empty)]:before:shrink-0 min-w-0 overflow-hidden">Jeremy Berman

@jeremyberman

Aug 12

I also ran the program with Codex and GPT 5.6 Sol (xhigh). It scored 73.7% (vs Opus' 96.2%) and used ~3x more actions.

An interesting difference between the two is that Sol kept trying to escape the sandbox and find solutions online. 7/25 Codex sessions did this vs 0/25 with Show more

103<br>7.6K

span:not(:empty)~span:not(:empty)]:before:content-['·'] [&>span:not(:empty)~span:not(:empty)]:before:px-1 [&>span:not(:empty)~span:not(:empty)]:before:shrink-0 min-w-0 overflow-hidden">Jeremy Berman

@jeremyberman

Aug 12

The models and coding harnesses are very good at building programs to learn new abstractions on the fly. As they continue to improve, it's important to give them the freedom and flexibility to build the machinery they need to solve a given task. As the models improve, the Show more

75<br>6.3K

span:not(:empty)~span:not(:empty)]:before:content-['·'] [&>span:not(:empty)~span:not(:empty)]:before:px-1 [&>span:not(:empty)~span:not(:empty)]:before:shrink-0 min-w-0 overflow-hidden">Psyho

@FakePsyho

19h

FYI, literally the first sentence in your github is incorrect. 30.2% is the semi-private score.

20<br>1.5K

Log in or sign up for X<br>See what’s happening and join the conversation<br>Continue with phoneContinue with AppleContinue with Google<br>or<br>Log in with username or email

Relevant people

Jeremy Berman@jeremybermanFollow<br>RL @humansand, prev training @reflection_ai, @ndea and co-founded https://t.co/aY50hNfhKb. yc w19.

Trending now

span empty before opus jeremy berman

Related Articles