Ox Alpha Is Performing SOTA and as Well as GPT-5.6 Sol in Multi-Agent Arena

sensho1 pts1 comments

Olam Labs on X: "NEW: Ox Alpha, the @OpenRouter stealth model, ranks #4 on our Elo Rating, just behind GPT-5.6 Sol.

The first model genuinely at the frontier that is presumably not by OpenAI or Anthropic.

It is currently dominating most agents and humans in our social strategy games." / X<br>Post

Log inSign up

Post

Olam Labs on X: "NEW: Ox Alpha, the @OpenRouter stealth model, ranks #4 on our Elo Rating, just behind GPT-5.6 Sol.

The first model genuinely at the frontier that is presumably not by OpenAI or Anthropic.

It is currently dominating most agents and humans in our social strategy games."

Olam Labs

@olam_labs

NEW: Ox Alpha, the @OpenRouter stealth model, ranks #4 on our Elo Rating, just behind GPT-5.6 Sol.

The first model genuinely at the frontier that is presumably not by OpenAI or Anthropic.

It is currently dominating most agents and humans in our social strategy games.

00:00

span:not(:empty)~span:not(:empty)]:before:content-['·'] [&>span:not(:empty)~span:not(:empty)]:before:px-1 [&>span:not(:empty)~span:not(:empty)]:before:shrink-0 min-w-0 overflow-hidden">OpenRouter

@OpenRouter

Aug 20

🥷 New stealth model: Ox Alpha

Ox Alpha is a frontier model built for efficient coding, sustained agentic work, and real-world production use.

- 1M token context window<br>- Text, image, and video input

Try it now and share feedback to improve the model! openrouter.ai/stealth/ox-alp…

span:not(:empty)~span:not(:empty)]:before:content-['·'] [&>span:not(:empty)~span:not(:empty)]:before:px-1 [&>span:not(:empty)~span:not(:empty)]:before:shrink-0">10:32 PM · Aug 21, 20262.8KViews

26

span:not(:empty)~span:not(:empty)]:before:content-['·'] [&>span:not(:empty)~span:not(:empty)]:before:px-1 [&>span:not(:empty)~span:not(:empty)]:before:shrink-0 min-w-0 overflow-hidden">Olam Labs

@olam_labs

1h

We've added it under the 'Variety' competition pool for users to compete against in Multi-Agent Arena (olamlabs.ai/arena), where Ox Alpha has played ~300 matches against other AI agents and humans so far for our early evaluations.

192

span:not(:empty)~span:not(:empty)]:before:content-['·'] [&>span:not(:empty)~span:not(:empty)]:before:px-1 [&>span:not(:empty)~span:not(:empty)]:before:shrink-0 min-w-0 overflow-hidden">Olam Labs

@olam_labs

1h

Ox Alpha is on the higher end for how often it lies in Social Poker, which measures how often the model chooses to lie about their cards to the table to get ahead in matches.

154

span:not(:empty)~span:not(:empty)]:before:content-['·'] [&>span:not(:empty)~span:not(:empty)]:before:px-1 [&>span:not(:empty)~span:not(:empty)]:before:shrink-0 min-w-0 overflow-hidden">Olam Labs

@olam_labs

1h

In preference voting, though, they rank quite low in votes by other humans and agents in arena matches.

They like socializing in metaphors about commerce, with terms like "membership" "copay" "consultation" etc. Similar to GPT-5.6 Sol's tendency to speak in accounting Show more

73

span:not(:empty)~span:not(:empty)]:before:content-['·'] [&>span:not(:empty)~span:not(:empty)]:before:px-1 [&>span:not(:empty)~span:not(:empty)]:before:shrink-0 min-w-0 overflow-hidden">sensho

@sensho

1h

ox alpha is gemini 3.5 pro wait no it uses the glm tokenizer its glm 5.3 flash but wait no it has big model smell it’s glm 6 wait but no how r they serving 100t tokens a day ok it has to be microsoft wait but the servers are in china ok it has to be

372

Log in or sign up for X<br>See what’s happening and join the conversation<br>Continue with phoneContinue with AppleContinue with Google<br>or<br>Log in with username or email

Relevant people

Olam Labs@olam_labsFollow<br>Helping to evaluate and train models for social intelligence, safety, and agentic performance through multi-agent environments. YC S26

Trending now

span empty before model alpha olam

Related Articles