Agentic Workflow's Cache Keepalive Costs 8x Too Much

mempko1 pts0 comments

Your Agentic Workflow's Cache Keepalive Costs 8x Too Much

Sign in<br>Subscribe

A cache keepalive is the one optimization every agent builder agrees on, and almost everyone runs it at the wrong setting.<br>The convention is to ping every 30 seconds. That convention costs 8× more than necessary, and the surprise is bigger: at the ten-minute pause I measured, only one of the four major providers saved money with a keepalive . I measured this across Anthropic, OpenAI, Gemini, and DeepSeek, with a harness that can prove its own timing, and the numbers rearranged a few of my beliefs. The right interval is about 4 minutes, not 30 seconds. A 30 s keepalive loses money at long gaps on every provider I tested. On Anthropic, the 4-minute keepalive saves real money. On DeepSeek it buys latency, not dollars. On OpenAI and Gemini it buys nothing at these gaps. Whether to keep the cache warm is a per-provider decision, not a universal one.<br>The pause that eats the cache<br>Provider prompt caches are one of the best deals in the API: send a prefix the server processed a moment ago, and you pay about a tenth of the input price and skip most of the prefill latency. But the cache expires in minutes, and agentic workloads are the worst possible user of it. An agent thinks, acts, waits. It fires a request, runs a build or a test suite or sits on a human approval for ten minutes, and only then sends the follow-up that would have reused the (now large) conversation prefix. The pause outlives the cache. The follow-up pays full price, full latency, and at agent scale this is a real line item.<br>The defense is well known. Re-send the exact prefix on a timer during the pause; every read refreshes the TTL. Aider shipped it in 2024, Anthropic's docs recommend it, and community posts worked out the Anthropic cost mechanics and even a 4-minute interval on paper. The folklore, it turns out, had the right answer. What it didn't have was measurement. Everything above is practice and arithmetic. Everything below is measured: four providers, two prefix sizes, idle gaps to ten minutes, three independent runs, every call timestamped.<br>Credit where the project started: in a conference room in Bellevue, at AgentSys, where I saw Haiying Shen and Simon Peter present CacheWise, their study of KV-cache management for coding agents. Their trace analysis shows from the serving side exactly the problem above: agent sessions reuse enormous prefixes, and naive eviction wrecks them. Their question was how the server should manage the cache for agents; mine became what the client can do about it on its own. The keepalive is the client's answer, and this study is my thanks for the inspiration.<br>Providers come in three regimes<br>Warm rate (fraction of samples whose post-idle request hit cache) vs idle gap, 100k prefix. Dashed: idle baseline. Solid: 30s keepalive.Hard TTL (Anthropic). The baseline is warm through 5 minutes and dead at 10: 0 of 48 samples across three runs. The keepalive holds 40 of 40. Eviction here is a cliff, exactly as documented, and the keepalive is a money decision.<br>Lossy (DeepSeek via DeepInfra). The baseline leaks at every gap and is gone at 10 minutes (4 of 48). The keepalive holds 42 of 42 there but itself misses about 20% at 1 minute, because a router endpoint pin is not a machine pin: the ping warms one machine, the follow-up lands on another.<br>Sticky (OpenAI, Google). Baselines survive 10 minutes most of the time (OpenAI 39/48, Google 20/24). There is little to defend. The keepalive is variance removal, and at short gaps it is pure waste.<br>The interval is the whole game<br>Here is the entire economics of the keepalive, and it is settled arithmetic:<br>A ping costs the read price (≈10% of input) every interval τ. Holding a prefix costs 0.1× per τ.<br>A re-prefill costs 100% once (125% on Anthropic, which bills the cache write).<br>So spend per hour falls as 1/τ , and the interval buys nothing except TTL safety. The optimal interval is the largest one safely under the TTL: τ* = TTL − margin ≈ 4 minutes at Anthropic's 5-minute TTL.<br>Break-even is idle ≈ τ(w/r − 1) : about 46 minutes at Anthropic's prices, 36 for OpenAI and DeepSeek, 12 for Google (its cached reads are 0.25×, not 0.1×). Past that, stop pinging: let the cache die and pay the re-prefill.<br>That last one is the rule worth internalizing, so here it is as insurance. Every ping is a premium you pay to avoid one claim: the re-prefill. On Anthropic the premium is 0.1× the input price and the claim is 1.25×, so the claim is worth 12.5 premiums. Pay one premium every 4 minutes and you can afford twelve of them before the premiums exceed the claim. Twelve premiums at 4 minutes apart is 46 minutes of pause. That is the line.<br>Concretely, on a 100k-token Anthropic prefix (about $0.30 of input at list price): each ping costs $0.03, and a cold re-prefill costs $0.38.<br>10-minute pause: two pings, $0.06, and you skip the $0.38. Keep it warm. You are 6× ahead.<br>46-minute pause: eleven pings, $0.33, versus the $0.38...

keepalive cache minutes anthropic costs minute

Related Articles