Frontier Models Don't Write So Good | Sean Lees Skip to content 22 July 2026<br>Share on X Share Copy link<br>GPT-3.5 arrived on the scene during the final year of my law degree. I used it as a study and research partner: it could reason across complex legal texts, explain them to me, and help me cram more effectively than ever before. Back then, frontier models weren't yet 'agentic'. Sure, they had a few tools in their harnesses (e.g. web search), but primarily they were text-based assistants:
user inputs text
model reasons
model outputs text
And they were damn good at that!
Beyond the Chatbot
It was clear that LLMs had potential far beyond the Chat interface. The race was on to make models more 'agentic': able to do more in the real world. Tool calling needed to become more reliable, and models needed to be able to reason longer and try harder before giving up. This was a post-training arms race, with each major model release heralding a step-change in agentic capabilities. In 2025, the models started getting good. Opus 4, in the nascent Claude Code, could suddenly work across and maintain an entire codebase. By late 2025, Opus 4.5 was good enough that professional software engineers were trusting it with real work. The GPT-5.6 models are relentlessly agentic: they can work for days, call complex tools, and write great software. METR estimates that the duration of tasks models can complete with a 50% success rate has doubled roughly every seven months since 2019.
Making models more agentic was half of the goal. The other half was making them more intelligent. Nobody wants an indefatigable assistant if it's stupid. Likewise, an erudite chatbot isn't very useful if it's impotent. Frontier labs' inexorable drive to scale their models' raw intelligence has been inspiring. They're committing trillions in capex and hiring the world's smartest minds to work on my generation's Manhattan Project. The gains kept coming through 2025, but I really started to feel the acceleration in 2026: GPT is solving Erdős problems; an Anthropic mathematician has just used Fable 5 to disprove the Jacobian Conjecture. And this intelligence isn't just accessible to mathletes: even I used Fable to solve a research problem I'd been tinkering with for eight years, letting me implement a novel state-of-the-art lossless audio codec (but that's a separate article!).
Smashing lossless audio compression benchmarks (with some help from Fable). Full songs come out average 6.3% smaller than FLAC's best setting, decoding up to ~29x faster than realtime for playback. Not sure what to do with the codec yet 🥸 If you're spending $$$ on lossless audio egress, DM me!<br>— Sean (@seanmozeik) 14 July 2026<br>Smashing lossless audio compression benchmarks (with some help from Fable). Full songs come out average 6.3% smaller than FLAC's best setting, decoding up to ~29x faster than realtime for playback. Not sure what to do with the codec yet 🥸 If you're spending $$$ on lossless audio egress, DM me!<br>— Sean (@seanmozeik) 14 July 2026
Claude Fable
Unfortunately, Fable also pushed me to investigate a suspicion that had crystallized over 18 months of working daily with frontier models. The first thing I did with Fable was point it at my notes. I maintain an Obsidian vault filled with my life's work: ideas, research, brain dumps, journal entries, and even songwriting (I used to be a record producer!). I use my own hybrid search tool, talon, so agents can search and reason over this knowledge base without loading the whole thing into context. I'm obviously intimately familiar with these data: taken together, they're my life's story! I'm also used to what agents do with it. So when I asked Fable:
"Analyze Sean's character and the state of his life based on this vault. Identify his weaknesses and blind spots. Write him a plan to help him realize his potential"
I already knew with some exactitude what a good answer looked like. I'd asked previous frontier models the same question, and I ruminate over it in my journal every day.
Fable blew me away. It found connections between my disparate ideas that I could never make myself, though I can't disclose more details without leaking half my "Dear Diary" entries to the internet! Although I had yet to ask it to do anything new, the leap in intelligence was undeniable: Fable had that sweet 'big model smell'. But all I had in front of me was the text it produced, and that text finally confirmed the shift I'd been sensing for 18 months:
AI Slop
The writing was shit. The prose was unnatural, the structure predictable, and the vocabulary full of infelicities. It was riddled with 'tics': phrases that rarely occur in human writing but appear constantly in model output:
load-bearing; it's not just X, it's Y; you're absolutely right; I hear...