Does a context gate for search agents actually work? | bluenotebook.io Table of Contents
Table of Contents
At Berlin Buzzwords 2026, Lester Solbakken gave a talk titled “Better retrieval makes agents worse” 1Marginnote buzzwords1Lester Solbakken. “When better retrieval makes agents worse.” Berlin Buzzwords 2026. Talk recording. Lester builds Hornet.dev. ↩.<br>These two slides matched my experience building search for agents.
Distractors degrade accuracy nonlinearlyControl context admission, not just top-k
Lester’s argument was that agentic retrieval is context admission control .<br>An agent retrieves context to act. A false positive is not a wasted result on a page: it enters the context, gets re-read at every later step, and shapes the next action.<br>Classic retrieval optimizes recall. A search tool inside an agent has to optimize precision.
The talk inspired me to build a context gate inside the search tool: a second model call that decides, per retrieved document, what enters the agent’s context.<br>BrowseComp-Plus 2Marginnote browsecomp2Chen et al. 2025. “BrowseComp-Plus: A More Fair and Transparent Evaluation Benchmark of Deep-Research Agent.” ACL 2026. OpenAI’s BrowseComp queries (Wei et al. 2025, https://arxiv.org/abs/2504.12516) rehosted over a fixed 100K-document corpus with labeled gold and evidence documents, indexed with BM25 and Qwen3 embeddings. https://arxiv.org/abs/2508.06600 · dataset ↩ was the test-bed.<br>It ships with labeled evidence documents for every query, which makes the idea measurable.<br>The rest of this post measures it: six agent-and-gate configurations over the same 180 queries.<br>Counting its own tokens, a cheap gate cuts input 1.4×, at accuracy indistinguishable from the ungated baseline.<br>But on DeepSeek’s cache pricing the gate loses money, and it doubles per-query latency.
Why context accumulates
An agent with a search tool issues several search calls for a single question, and every result stays in the conversation until the final answer.<br>By search 8, the agent is re-reading the distractors from search 1 on every step.<br>Massive context windows make context stuffing the easy way out: leave everything in and trust the model to figure out the answer.<br>It mostly works, and it pays for that in tokens and latency at every step.
Prior work has measured what it also costs in accuracy.<br>The first-drop-of-ink paper Lester cites is worth a closer read.
accuracy vs share of hard distractors
Llama-3.1-8B, NQ, 128K context (Gao et al. 2026)<br>60% 70% 80% 90% 0255075100<br>hard distractors in the context (%)
first 10% →<br>0% hard distractors: 87% accuracy 87 1% hard distractors: 85.5% accuracy 2% hard distractors: 82% accuracy 3% hard distractors: 78% accuracy 5% hard distractors: 76% accuracy 10% hard distractors: 72.5% accuracy 72.5 20% hard distractors: 70.5% accuracy 40% hard distractors: 66% accuracy 60% hard distractors: 63.5% accuracy 80% hard distractors: 62.5% accuracy 90% hard distractors: 64.5% accuracy 100% hard distractors: 62% accuracy 62<br>the first drop, zoomed
same run, first tenth of the axis<br>70% 75% 80% 85% 90% 0123510<br>10% of the distractors, 58% of the damage<br>0% hard distractors: 87% accuracy 87 1% hard distractors: 85.5% accuracy 85.5 2% hard distractors: 82% accuracy 82 3% hard distractors: 78% accuracy 78 5% hard distractors: 76% accuracy 76 10% hard distractors: 72.5% accuracy 72.5<br>From Gao et al.<br>Left: accuracy against the share of hard distractors in that context.<br>Right: the shaded first 10%, zoomed in.<br>Of the 25 points the model loses in total, 14.5 are gone before the context is even a tenth distractors.
Gao et al. 3Marginnote firstink3Gao, Chen, and Huang. 2026. “The First Drop of Ink: Nonlinear Impact of Misleading Information in Long-Context Reasoning.” The paper behind the talk’s ink slide: a small fraction of hard distractors causes most of the degradation, and filtering gains come mainly from context-length reduction rather than distractor removal. https://arxiv.org/abs/2605.10828 ↩ pin the damage on hard distractors, documents close enough to the topic to pass for evidence.<br>The first such documents to enter the context do most of the harm.
In my ungated baseline run on BrowseComp-Plus, a single question costs the deepseek-v4-pro agent just over a million input tokens this way, 1,035K on average. The weaker flash agent averages 1.29M.<br>Either agent writes roughly 1/100 of what it reads.
0 20K 40K 60K 80K 100K<br>context the model re-reads at this step (K tokens)
summed over all 18 steps this schematic reads ~910K tokens. Measured flash-agent runs average 1.29M per query.<br>search 1 re-reads 8K tokens. 3K system, 0K stale distractors, 0K evidence, 5K new results · 38% of this step is dead weight<br>search 1 38% dead weight<br>search 2 re-reads 13K tokens. 3K system, 5K stale distractors, 0K evidence, 5K new results · 62% of this step is dead weight<br>search 2 search 3 re-reads 18K tokens. 3K system, 10K stale distractors, 0K evidence, 5K new results · 72% of this step is dead...