ARC Prize on X: "Fable 5 from @AnthropicAI on ARC-AGI (Verified):
- ARC-AGI-2: 89.2%, $5.45/task<br>- ARC-AGI-1: 98.5%, $1.02/task
Fable 5 is the highest-scoring model we have evaluated on ARC-AGI-1 to date, while also achieving 89.2% on ARC-AGI-2. https://t.co/GpVKBEMXPz" / X<br>Post
Log inSign up
Post
ARC Prize
@arcprize
Fable 5 from @AnthropicAI on ARC-AGI (Verified):
- ARC-AGI-2: 89.2%, $5.45/task<br>- ARC-AGI-1: 98.5%, $1.02/task
Fable 5 is the highest-scoring model we have evaluated on ARC-AGI-1 to date, while also achieving 89.2% on ARC-AGI-2.
span:not(:empty)~span:not(:empty)]:before:content-['·'] [&>span:not(:empty)~span:not(:empty)]:before:px-1 [&>span:not(:empty)~span:not(:empty)]:before:shrink-0">9:26 PM · Aug 5, 202681.9KViews
18<br>44<br>803<br>82
span:not(:empty)~span:not(:empty)]:before:content-['·'] [&>span:not(:empty)~span:not(:empty)]:before:px-1 [&>span:not(:empty)~span:not(:empty)]:before:shrink-0 min-w-0 overflow-hidden">ARC Prize
@arcprize
12h
Full results: arcprize.org/results/anthro…
We're currently running ARC-AGI-3 evaluations and will publish results when they're complete.
86<br>7.1K
span:not(:empty)~span:not(:empty)]:before:content-['·'] [&>span:not(:empty)~span:not(:empty)]:before:px-1 [&>span:not(:empty)~span:not(:empty)]:before:shrink-0 min-w-0 overflow-hidden">ARC Prize
@arcprize
12h
ARC’s semi-private evaluation policy is designed to prevent test data from being used for training or manually reviewed, reducing the risk of developer-aware targeting. We normally enforce this through Zero Data Retention.
Because ZDR is not available for Fable 5, we worked with Show more
54<br>7.1K
span:not(:empty)~span:not(:empty)]:before:content-['·'] [&>span:not(:empty)~span:not(:empty)]:before:px-1 [&>span:not(:empty)~span:not(:empty)]:before:shrink-0 min-w-0 overflow-hidden">ARC Prize
@arcprize
12h
- Leaderboard: arcprize.org/leaderboard<br>- Reproduce the public results: github.com/arcprize/arc-a…<br>- Testing policy: arcprize.org/policy<br>- Full Fable 5 results: arcprize.org/results/anthro…
ARC Prize - Leaderboard
From arcprize.org
19<br>5.3K
span:not(:empty)~span:not(:empty)]:before:content-['·'] [&>span:not(:empty)~span:not(:empty)]:before:px-1 [&>span:not(:empty)~span:not(:empty)]:before:shrink-0 min-w-0 overflow-hidden">Marko Njegomir
@njmarko
12h
What about Prime Intellect achieving 95.5% on ARC-AGI-3?
span:not(:empty)~span:not(:empty)]:before:content-['·'] [&>span:not(:empty)~span:not(:empty)]:before:px-1 [&>span:not(:empty)~span:not(:empty)]:before:shrink-0 min-w-0 overflow-hidden">Prime Intellect
@PrimeIntellect
14h
Replying to @PrimeIntellectPrime Agent is a general-purpose coding harness
On ARC-AGI-3, it scores 95.5%, surpassing the human-expert baseline, but the gain is not benchmark-specific.
We see major improvements across models when compared to their proprietary harnesses:
38<br>3K
Log in or sign up for X<br>See what’s happening and join the conversation<br>Continue with phoneContinue with AppleContinue with Google<br>or<br>Log in with username or email
Relevant people
ARC Prize@arcprizeFollow<br>A North Star for open AGI. Co-founders: @fchollet @mikeknoop. President: @gregkamradt. We're hiring mission-driven builders: https://t.co/GswTSnyCoJ
Trending now