Jeremy Berman on X: "I got 96.2% on ARC-AGI-3 with Opus 5, and 99.3% pass@2. The program is basically Claude Code + Opus 5 (high), one action command, and filesystem logs. Almost nothing ARC specific.
https://t.co/NHyibLgade https://t.co/531wS8sxZ0" / X<br>Post
Log inSign up
Post
Jeremy Berman on X: "I got 96.2% on ARC-AGI-3 with Opus 5, and 99.3% pass@2. The program is basically Claude Code + Opus 5 (high), one action command, and filesystem logs. Almost nothing ARC specific.
https://t.co/NHyibLgade https://t.co/531wS8sxZ0"
Jeremy Berman
@jeremyberman
I got 96.2% on ARC-AGI-3 with Opus 5, and 99.3% pass@2. The program is basically Claude Code + Opus 5 (high), one action command, and filesystem logs. Almost nothing ARC specific.
github.com/jerber/arc-code
span:not(:empty)~span:not(:empty)]:before:content-['·'] [&>span:not(:empty)~span:not(:empty)]:before:px-1 [&>span:not(:empty)~span:not(:empty)]:before:shrink-0">8:11 PM · Aug 12, 2026223.8KViews
51<br>84<br>1.1K<br>611
span:not(:empty)~span:not(:empty)]:before:content-['·'] [&>span:not(:empty)~span:not(:empty)]:before:px-1 [&>span:not(:empty)~span:not(:empty)]:before:shrink-0 min-w-0 overflow-hidden">Jeremy Berman
@jeremyberman
Aug 12
I give Claude Code a computer, a record of everything that has happened, and a way to act in the game. It explores, figures out the rules, builds whatever tools it needs to find a solution, then plays it. Every game starts from scratch.
None of the ARC-specific machinery is Show more
111<br>12K
span:not(:empty)~span:not(:empty)]:before:content-['·'] [&>span:not(:empty)~span:not(:empty)]:before:px-1 [&>span:not(:empty)~span:not(:empty)]:before:shrink-0 min-w-0 overflow-hidden">Jeremy Berman
@jeremyberman
Aug 12
In one pass Opus wrote 269 programs (~12,700 lines). It built parsers for all 25 games, searching functions for 23, and game simulators for 9.
It built a different harness for each problem.
The filesystem memory idea came from the excellent PRO-LONG harness: Show more
GitHub - alexisfox7/PRO-LONG: Programmatic memory for long-horizon LLM agents: the harness appends...
From github.com
115<br>9.9K
span:not(:empty)~span:not(:empty)]:before:content-['·'] [&>span:not(:empty)~span:not(:empty)]:before:px-1 [&>span:not(:empty)~span:not(:empty)]:before:shrink-0 min-w-0 overflow-hidden">Jeremy Berman
@jeremyberman
Aug 12
Code execution makes this scalable and cheaper, which is why using this harness is *cheaper* than asking Opus to solve each game directly.
With code, the model can compile its reasoning into a function, run that function thousands of times, and execute whole action sequences. Show more
64<br>7.4K
span:not(:empty)~span:not(:empty)]:before:content-['·'] [&>span:not(:empty)~span:not(:empty)]:before:px-1 [&>span:not(:empty)~span:not(:empty)]:before:shrink-0 min-w-0 overflow-hidden">Jeremy Berman
@jeremyberman
Aug 12
I also ran the program with Codex and GPT 5.6 Sol (xhigh). It scored 73.7% (vs Opus' 96.2%) and used ~3x more actions.
An interesting difference between the two is that Sol kept trying to escape the sandbox and find solutions online. 7/25 Codex sessions did this vs 0/25 with Show more
103<br>7.6K
span:not(:empty)~span:not(:empty)]:before:content-['·'] [&>span:not(:empty)~span:not(:empty)]:before:px-1 [&>span:not(:empty)~span:not(:empty)]:before:shrink-0 min-w-0 overflow-hidden">Jeremy Berman
@jeremyberman
Aug 12
The models and coding harnesses are very good at building programs to learn new abstractions on the fly. As they continue to improve, it's important to give them the freedom and flexibility to build the machinery they need to solve a given task. As the models improve, the Show more
75<br>6.3K
span:not(:empty)~span:not(:empty)]:before:content-['·'] [&>span:not(:empty)~span:not(:empty)]:before:px-1 [&>span:not(:empty)~span:not(:empty)]:before:shrink-0 min-w-0 overflow-hidden">Psyho
@FakePsyho
19h
FYI, literally the first sentence in your github is incorrect. 30.2% is the semi-private score.
20<br>1.5K
Log in or sign up for X<br>See what’s happening and join the conversation<br>Continue with phoneContinue with AppleContinue with Google<br>or<br>Log in with username or email
Relevant people
Jeremy Berman@jeremybermanFollow<br>RL @humansand, prev training @reflection_ai, @ndea and co-founded https://t.co/aY50hNfhKb. yc w19.
Trending now