Try the Fast Models

meetpateltech1 pts0 comments

Try the very fast models – Nate Meyvis

Try the very fast models

27 Jul, 2026

Models are getting better. That means that the most powerful models today are much more powerful than the most powerful models of a month ago, but it also means that today's fast models today are a lot more powerful than the fast models of a month ago.

Even if you pay close attention to new models,1 it's easy to focus on power at the frontier and overlook the "93% as good and super fast" model. But that's a mistake for a lot of reasons, some more obvious than others:

If the fast model does something as well as the full-power model, you can save significant tokens or money by using the fast one.

Your instincts about what tasks benefit from the full-power model are probably not perfect. Mine certainly aren't, and it seems to me that they're worse than they were a month or two ago. Put another way: "93% as good" is better than it used to be, and I suspect that the shape of that 93% is harder to understand and predict.

Using different models to check each other's work is an important technique, and combining a full-power model with a fast one can be a great way to do this.

Saying "fast is different!" is one thing, and (if you're like me in this respect) actually experiencing is quite a different thing. When you fix a bug in 45 seconds instead of 3 minutes, do a data-centric task over thousands of rows in a few seconds, or get professional-quality dissertation feedback2 in 10 to 30 seconds, you see all sorts of new possibilities.

Many of us formed instincts about what to send to fast models when those models were much, much worse. Lots of day-to-day work is now amenable to fast-model assistance. For me, at least, starting all LLM work with fast models has been a useful way to reset my obsolete3 instincts.

I drafted the first 90% of this blog post last Friday morning and only finished it now. Everything above is significantly more true now than it was when I started writing it. So, this post is itself an example of LLM progress punishing (relatively, at least) normal-speed production.

This is a tricky subject to write about, just on an audience-relationship level. To a first approximation, there are two groups of readers: one group is paying little or no attention to the models powering the AI products they use, and the other is paying very close attention to them.↩

Again, by "professional-quality" I don't mean "as good as expert philosophers within their expertise," I mean "comparably useful to professionally trained philosophers outside their expertise but being careful and doing their homework."↩

That is, approximately two months old.↩

#future of work

#generative AI

#software

models fast model powerful power work

Related Articles