$0.26 DeepSeek V4 Flash 0731 ties $5.01 GPT-5.6 run on Agentic Memory Benchmark

howardme11 pts0 comments

ATM-Bench Leaderboard

EN<br>中文

Leaderboard<br>ATM-Bench Leaderboard

Unified results across oracle, agent, memory, and RAG systems

Submitted results on ATM-Bench, ATM-Bench-Hard, and the NIAH long-context stress test.<br>Use the tabs to switch boards, the chips to filter by system type, and click any column header to sort.

Project Page

arXiv

Code

Dataset

Submit Result

Last updated: 2026-06-01

← ATM-Bench<br>ATM-Bench<br>ATM-Bench-Hard<br>NIAH<br>Submit

Price vs. Performance

ATM-Bench-Hard agent runs: what each one scored against what it cost, at API list-price<br>equivalent. A line joins one model's reasoning-effort tiers; a diamond is a single configuration.

ATM-Bench<br>ATM-Bench-Hard

ATM-Bench-Hard-NIAH<br>Preview

ATM-Bench

- indicates the field has not been reported by the submitter.<br>Memory Model is the LLM used to construct the memory store;<br>Retriever is the embedding model used at query time.<br>Click any column header to sort; click a filter chip to narrow by system type.<br>When a caption model is not stated, it is Qwen3-VL-2B (the default).<br>* Memexa's QS is measured with a DeepSeek-V4-flash judge (its own answer model), not the gpt-5-mini judge used for every other row, so it is shown for reference and is not directly comparable; Recall is judge-independent and like-for-like.

ATM-Bench-Hard

- indicates the field has not been reported.<br>Coding agents use their default configuration unless the model label states a reasoning effort.<br>Results use our proposed SGM (Schema-Guided Memory) unless marked (w/o SGM).<br>When a caption model is not stated, it is Qwen3-VL-2B (the default).<br>* QS on the DeepSeek-V4-flash rows (Memexa and the A-Mem / MemPalace / HippoRAG2 re-runs) is measured with a DeepSeek-V4-flash judge (the system's own model), not the gpt-5-mini judge used for the other rows — shown for reference, not directly comparable. The three baseline re-runs are community-run at this LLM tier, not the original authors' official numbers.<br>† Memexa's Recall is reported on the fixed Qwen3-VL-2B captions (like-for-like with other rows), although this configuration answers from Qwen3.6-27B captions.<br>‡ Memexa's Total Tokens / Cost include the one-time memory build (33.9M tokens / ~$3.80 over all 11,034 items, amortizable across evaluations) plus the 31-question query (0.40M / ~$0.04); other rows' agent costs are per full run.<br>Cost is an API-equivalent estimate calculated from saved per-call token counters with Tokdash's bundled standard short-context rates; it does not represent Codex subscription charges.<br>§ OpenCode meters the DeepSeek V4 Flash Free endpoint at $0; the cost shown is the underlying model's list-price equivalent at the same rates as every other row.<br>¶ GPT-5.6 Luna and Terra costs follow OpenAI's official price reduction for these models, verified 2026-07-31: Luna $1.00/$6.00 to $0.20/$1.20 and Terra $2.50/$15.00 to $2.00/$12.00 per million input/output tokens. The runs are unchanged — same tokens, same QS — only the rate they are priced at.

ATM-Bench-Hard-NIAH

Preview · More submissions welcome

Needle-in-a-haystack sweep on ATM-Bench-Hard (31 questions): each question's gold evidence is hidden among k = 25 / 50 / 100 / 200 distractor memory items, under the SGM (Schema-Guided Memory) setting. Oracle = gold evidence only, no haystack; judge: gpt-5-mini. Approximate context depth is shown in each column header.

SGM

Raw (real images/video)

Why SGM, not raw? Raw (real images/video) edges out SGM at the Oracle ceiling. But that advantage collapses under realistic conditions: as the haystack fills with distractors, raw degrades and even fails (payload/context limits), and under agentic retrieval the gap is stark — every "w/o SGM" (raw) agent lands far below its SGM run. SGM is the representation that holds up once there is noise under realistic conditions.

Submit Your Result

We welcome new submissions across all three boards. To keep the leaderboard credible, please include<br>reproduction details (system type, harness, model + version, code or commit, total token cost when applicable).

Send a Pull Request

Fastest path: open a PR adding a row to the TRACKS array at the bottom of<br>leaderboard.html. Include a short description of the setup and a link to your run logs<br>or code in the PR body.

Edit leaderboard.html on GitHub

Open an Issue

Prefer not to send a PR? File an issue with your system type, harness, scores, and a reproduction<br>pointer. We will add the row on your behalf.

Open submission issue

Acknowledgement

This page was adopted from the<br>Nerfies project page, licensed under a<br>Creative<br>Commons Attribution-ShareAlike 4.0 International License. Many thanks to the<br>Academic<br>Project Page Template.

bench model memory hard leaderboard judge

Related Articles