On DeepSeek V4 Flash and Cheap Intelligence | Sid's Blog<br>🇸">
The cost of tokens and intelligence seems to be plunging, despite what my own internet bubble led me to believe was going to happen. Between DeepSeek V4 Flash going toe to toe with many SOTA models at a very, very small fraction of the cost and GPT-5.6 Luna getting a massive price cut, the narrative that intelligence would remain expensive, if not increase over time, is looking increasingly difficult to defend in my head.
Full disclaimer on my workloads though: my intelligence needs are very prosaic. I mainly use AI for code: a lot of C/C++ for embedded devices plus generic web endpoints and dashboards to ingest and present data. I also end up needing a ton of Swift. So not exactly (or exclusively) webslop but not cutting edge research work either. (I’m not using it to disprove the Jacobian conjecture, that’s for sure.)
The release of DeepSeek V4 Flash has upended tokenomics and has caught a lot of people off guard with its performance and cost. Artificial Analysis’ analysis shows it costing 1/100th the cost of Fable while being fairly competitive in various benchmarks. Yes, not a typo. 1/100th . 1/60th the cost of GPT-5.6 Sol, 1/80th the cost of Opus 5. Arena’s leaderboard paints the same picture.
Maybe it’s recency bias but never before could you do so much for so little. Intelligence that’s so cheap and so good that it’s too cheap to meter. I no longer find myself model switching with Claude Code just to protect my 5-hour, and weekly, quota.
And yes, it meanders around on long horizon tasks, is slower, and not very token efficient so I just end up spinning up way more subagents and have something else orchestrate and coordinate. Sure, it doesn’t have the taste of Opus, but those are areas where I can step in and fill in the blanks. The weaknesses are things I could live with and engineer around.
When the cost of a workflow, any workflow, drops from a few dollars to a few cents, many ideas that were previously only viable for high-value enterprises, or just untenable altogether, can suddenly become practical for everyday use. I’m not downplaying the impact of true frontier intelligence, but intelligence at mass scale is where things start to get really interesting.
Between DeepSeek V4 Flash, MiniMax H3, and Qwen 3.8 27B, this has been an incredible week in open source LLM history. I’ve been cautiously optimistic for the longest time but these last few weeks have nudged me deep into pure meliorism.
If you've reached this far, thank you for reading! :)
I thought retiring in my mid 30s after a few exits would be fun but I've just been bored and a bit undersocialized without morning Slacks and emails to wake up to. If you’re building something interesting and could use an extra set of hands to ship, or just want to say hi, feel free to reach out. My inbox is open.