frontier AI labs lost the "deep research" frontier

williamtrask1 pts0 comments

~6 weeks ago... frontier AI labs lost the "deep research" frontier

Andrew’s Substack

SubscribeSign in

~6 weeks ago... frontier AI labs lost the "deep research" frontier<br>~6 weeks ago, several "fusion models" exceeded every frontier AI system on a deep research benchmark called DRACO. Over a month later, their lead is holding. Why?

Andrew Trask(Oxford/OpenMined)<br>Jul 22, 2026

Share

Since May, OpenRouter, Sakana AI, and the US Government’s Los Alamos National Laboratory have each shared that routed combinations of AI models exceed the accuracy of individual frontier models . In one case… Fable-level quality at half the price of Anthropic’s Fable API (see results).<br>A month after OpenRouter’s announcement, I asked two of my colleagues at OpenMined (Siddhant and Khoa) to re-check OpenRouter’s results, and the lead is still reproducible.

If you want the most capable “deep research” AI system in the world, or you want frontier level quality at 50% of the price… for the last month, you couldn’t get it from a single open-source or closed-source AI (at least according to DRACO). You could only get it from routed ensembles of models… from network-source AI.<br>Q: Why?<br>Colloquially… because 2 heads are better than one.

The reason is actually the re-discovery of an old idea from Machine Learning. When you ask 10 models the same prompt, they tend to make mistakes in anti-correlated ways. So when 2 or 3 of those models agree on an answer, it tends to be correct more often than any single model in the ensemble on its own. Consequently, when several models are evaluated on DRACO, if you combine their results, you can get higher scores on DRACO than any individual model on its own.<br>Q: But the top models are ensembles of frontier models… why are you saying frontier labs “lost the frontier”?<br>True, but that’s less relevant than one might think. Consider the following questions:<br>User Relationship: Whose API do I go to to get the most accurate AI predictions?<br>It used to be APIs run by frontier AI companies… but now it’s not. I have to go to routers like Sakana, OpenRouter, etc. That’s the new “front door” that the customer sees yielding the most accurate deep research predictions in the world. That’s a significant shift in user power from frontier AI companies to routers.

Economic Power: Who sets the price of AI-powered deep research?<br>Routers of ensembles offer a radical level of interoperability which has a significant impact on pricing power. It used to be individual model providers who set the price of AI-powered deep research, but now the fair market price of intelligence comes from routers and their ability to pivot quickly across ensembles. If an individual model provider changes their price, that doesn’t necessarily mean the end-user price will change… the router could absorb the change, or swap that model for another model without the user knowing or caring. A dramatic shift in economic power from frontier AI companies to routers.

Political Power: Who has the off button for AI-powered deep research?<br>Fable was turned off… but Fable-level quality was still available via routers. That’s a significant shift in political power from frontier AI companies to routers.

So yeah… the top ensembles include frontier systems… but I think it’s pretty hard to say that any single AI lab holds the frontier on AI-powered “deep research” anymore. From a user, economic, or political perspective… it’s not really the case. Is there any other perspective that matters?<br>Doesn’t this increase the cost?<br>This was perhaps the biggest surprise from the OpenRouter and Los Alamos articles. OpenRouter and Los Alamos researchers offset the increase in accuracy by using smaller/cheaper models in their ensembles… creating a net-decrease in the cost of frontier level AI capability. In the case of OpenRouter, they were able to show Fable-level quality at half the price of Fable.<br>So it would appear that the pareto-frontier of AI quality/cost is now routed ensembles… at least for deep research tasks.<br>What does this mean for AI’s future?<br>At the moment, the overton window isn’t really paying attention to this, but the window can only ignore superior quality + lower price for so long. At OpenMined, we’re going to start checking these results on more benchmarks. I’ll share the bigger findings here.

Thanks for reading Andrew’s Substack! Subscribe for free to receive new posts and support my work.

Subscribe

Share

Discussion about this post<br>CommentsRestacks

TopLatestDiscussions

No posts

Ready for more?

Subscribe

© 2026 Andrew Trask · Privacy ∙ Terms ∙ Collection notice<br>Start your SubstackGet the app<br>Substack is the home for great culture

This site requires JavaScript to run correctly. Please turn on JavaScript or unblock scripts

frontier deep research models from price

Related Articles