Frontier Models Don't Write So Good

seanmozeik1 pts0 comments

Frontier Models Don't Write So Good | Sean Lees Skip to content 22 July 2026<br>Share on X Share Copy link<br>GPT-3.5 ar­rived on the scene during the final year of my law degree. I used it as a study and re­search part­ner: it could reason across com­plex legal texts, ex­plain them to me, and help me cram more ef­fec­tively than ever before. Back then, fron­tier models weren't yet 'agen­tic'. Sure, they had a few tools in their har­nesses (e.g. web search), but pri­mar­ily they were text-based as­sis­tants:

user inputs text

model reasons

model outputs text

And they were damn good at that!

Beyond the Chatbot

It was clear that LLMs had po­ten­tial far beyond the Chat in­ter­face. The race was on to make models more 'agen­tic': able to do more in the real world. Tool call­ing needed to become more re­li­able, and models needed to be able to reason longer and try harder before giving up. This was a post-train­ing arms race, with each major model re­lease herald­ing a step-change in agen­tic ca­pa­bil­i­ties. In 2025, the models started get­ting good. Opus 4, in the nascent Claude Code, could sud­denly work across and main­tain an entire code­base. By late 2025, Opus 4.5 was good enough that pro­fes­sional soft­ware en­gi­neers were trust­ing it with real work. The GPT-5.6 models are re­lent­lessly agen­tic: they can work for days, call com­plex tools, and write great soft­ware. METR es­ti­mates that the du­ra­tion of tasks models can com­plete with a 50% suc­cess rate has dou­bled roughly every seven months since 2019.

Making models more agen­tic was half of the goal. The other half was making them more in­tel­li­gent. Nobody wants an in­de­fati­ga­ble as­sis­tant if it's stupid. Likewise, an eru­dite chat­bot isn't very useful if it's im­po­tent. Frontier labs' in­ex­orable drive to scale their mod­els' raw in­tel­li­gence has been in­spir­ing. They're com­mit­ting tril­lions in capex and hiring the world's smartest minds to work on my gen­er­a­tion's Manhattan Project. The gains kept coming through 2025, but I really started to feel the ac­cel­er­a­tion in 2026: GPT is solv­ing Erdős prob­lems; an Anthropic math­e­mati­cian has just used Fable 5 to dis­prove the Jacobian Conjecture. And this in­tel­li­gence isn't just ac­ces­si­ble to math­letes: even I used Fable to solve a re­search prob­lem I'd been tin­ker­ing with for eight years, let­ting me im­ple­ment a novel state-of-the-art loss­less audio codec (but that's a sep­a­rate ar­ti­cle!).

Smashing lossless audio compression benchmarks (with some help from Fable). Full songs come out average 6.3% smaller than FLAC's best setting, decoding up to ~29x faster than realtime for playback. Not sure what to do with the codec yet 🥸 If you're spending $$$ on lossless audio egress, DM me!<br>— Sean (@seanmozeik) 14 July 2026<br>Smashing lossless audio compression benchmarks (with some help from Fable). Full songs come out average 6.3% smaller than FLAC's best setting, decoding up to ~29x faster than realtime for playback. Not sure what to do with the codec yet 🥸 If you're spending $$$ on lossless audio egress, DM me!<br>— Sean (@seanmozeik) 14 July 2026

Claude Fable

Unfortunately, Fable also pushed me to in­ves­ti­gate a sus­pi­cion that had crys­tal­lized over 18 months of work­ing daily with fron­tier models. The first thing I did with Fable was point it at my notes. I main­tain an Obsidian vault filled with my life's work: ideas, re­search, brain dumps, jour­nal en­tries, and even song­writ­ing (I used to be a record pro­ducer!). I use my own hybrid search tool, talon, so agents can search and reason over this knowl­edge base with­out load­ing the whole thing into con­text. I'm ob­vi­ously in­ti­mately fa­mil­iar with these data: taken to­gether, they're my life's story! I'm also used to what agents do with it. So when I asked Fable:

"Analyze Sean's character and the state of his life based on this vault. Identify his weaknesses and blind spots. Write him a plan to help him realize his potential"

I al­ready knew with some ex­ac­ti­tude what a good answer looked like. I'd asked pre­vi­ous fron­tier models the same ques­tion, and I ru­mi­nate over it in my jour­nal every day.

Fable blew me away. It found con­nec­tions be­tween my dis­parate ideas that I could never make myself, though I can't dis­close more de­tails with­out leak­ing half my "Dear Diary" en­tries to the in­ter­net! Although I had yet to ask it to do any­thing new, the leap in in­tel­li­gence was un­de­ni­able: Fable had that sweet 'big model smell'. But all I had in front of me was the text it pro­duced, and that text fi­nally con­firmed the shift I'd been sens­ing for 18 months:

AI Slop

The writ­ing was shit. The prose was un­nat­u­ral, the struc­ture pre­dictable, and the vo­cab­u­lary full of in­fe­lic­i­ties. It was rid­dled with 'tics': phrases that rarely occur in human writ­ing but appear con­stantly in model output:

load-bearing; it's not just X, it's Y; you're absolutely right; I hear...

models fable good search text work

Related Articles