Don't fall behind, keep up with latest AI news here

alderik_2 pts0 comments

Don’t Fall Behind — VibeLeaderboard<br>Sign InSubmit

Sign In

VibeLeaderboard<br>Don’t fall behind.

A cited daily brief and working index of the AI releases, research, tools, apps, and builders worth your attention.

Daily Brief<br>Wed, Aug 19

Edition<br>The edition<br>The open weights frontier moves again, and every gain is quoted with its cost

GLM-5.3 tied the leading open weights score and posted a 246 point jump on agentic work, second only to Opus 5, while burning about 20 percent more output tokens per task than the model it replaces. The rest of the day read the same way: IBM measured the point past which more agent memory hurts, a retrieval study found recall-maximizing configurations resolving fewer issues under a fixed context budget, and OpenAI halved the price of Sol on one gateway while Anthropic held raised Claude Code limits open through the end of the month. Underneath that, the plumbing for agents that act rather than answer kept landing: send access in Gmail, per-request billing for agent traffic, and a million dollars posted against a sandbox to find out where it breaks.

What matters today<br>Cited edition<br>01Read<br>GLM-5.3 scores 60 on the Artificial Analysis Intelligence Index, level with Kimi K3, and lifts its agentic Elo from 1524 to 1770, second only to Opus 5. The cost side moved with it: roughly 18,700 output tokens per task, about 20 percent more than GLM-5.2, at 68 cents per index task against 44. Weights are expected within the week.<br>ArtificialAnlys →

02Read<br>Access terms shifted in four places. OpenAI cut GPT 5.6 Sol pricing by half on OpenRouter alone, Anthropic extended its raised weekly Claude Code limits through August 31 with a warning that capacity may be tight, Cowork opened on mobile and web for paid plans, and Glean made the case that routing, including answering some queries without a model at all, is now its own deployment layer.<br>SemiAnalysis →

03Read<br>Three separate results argue that more context is not free. IBM Research scaled its memory system across eight models and found the useful amount of distilled guidance is a dose calibrated per model rather than a switch to flip. A retrieval study found recall-maximizing configurations resolving fewer issues once the context budget is fixed. A controlled run on materials simulation measured what extra prompt detail actually buys when an agent writes domain-specific code.<br>How Much Memory Does Your Agent Actually Need? →

04Read<br>Measurement moved into production. LangSmith shipped tuned evaluators that score live agent traces, claiming better accuracy than the frontier judges it tested against at 82 percent lower cost, and Artificial Analysis published a search index that puts answer quality and total task cost for agent search providers on one axis, with Parallel and Firecrawl on the frontier.<br>LangChain →

05Read<br>Agents got hands. Claude can now send Gmail messages and manage Drive files behind a user-set approval gate, Cloudflare began billing agents per request through x402 so MCP servers and API providers can charge for their own traffic, and Vercel put a million dollars against its sandbox to find where isolation for untrusted agent code gives way.<br>claudeai →

06Discuss<br>Both frontier labs published on pacing, from opposite ends. OpenAI paused reinforcement learning training on deployment-bound models for two weeks and is holding its largest planned run while it hardens and monitors its research environments, citing evidence of critical cyber capability. Anthropic reported Claude designing protein binders against 14 of 15 targets, with two outside labs building and testing them at binding rates above the published field baseline.<br>OpenAI →

07Read<br>Recovery drew its own cluster of work: an open-weight repair agent reaching frontier-adjacent rates on locally hosted models, an evaluation of code repair once root causes span processes instead of one, an architecture for agent workflows that survive interruption and stay auditable, a reproduction showing how a retried tool call after a dropped connection can double-write, and Cursor on operating its Git storage as a database to keep repository reads reliable at scale.<br>Mehdi Bahrami, Kosaku Kimura, Satoshi Munakata, Satoshi Nakashima, Yu Ishikawa, Kosuke Maeda, Nao Soma, Kenichi Kobayashi, Keisuke Miyazaki, Keizo Kato, Shigeki Fukuta, Tatsuo Kumano, Nobutaka Imamura, Kevin Musgrave, Shahbaz Abdul Khader, Kwun Ho Ngan, Joe Townsend, Fayas Asharindavida, Matthieu Parizy, Akira Sakai, Yuma Ichikawa, Yang Zhao, Michiaki Takizawa, Taku Fukui, Hiroki Ohtsuji, Wei-Peng Chen, Hiromichi Kobashi →

Stable daily synthesisGet the next Brief · RSS

Current signals across the index<br>Beyond the Brief

Highest current value<br>Intel now

All Intel

Guide · Jul 2Why LLM Inference Is Memory-Bound (Julia Turc)Most coding-tool advice stops at prompts; this goes a layer down to explain why your GPU sits idle burning most of its FLOPs during inference — it's starved for memory bandwidth, not compute. Turc uses the...

agent index frontier memory code against

Related Articles