Notion’s Knowledge Board<br>This is not a traditional benchmark. Benchmarks are stable and curated sets of difficult, fully specified tasks, which are designed to discover new capabilities as models advance. Notion’s Knowledge Board asks a different question: How do models handle the tasks users actually delegate to them?<br>Notion users rely on agents for all kinds of knowledge work: meeting follow ups and action items, inbox and calendar cleanup, support-ticket follow-ups, sales-pipeline updates, lead research, knowledge-base Q&A, recurring project reports, and more. Our Knowledge Board is a live, ongoing measurement of how well models in Notion do that work.<br>We developed an ensemble judge in collaboration with researchers at Anthropic and OpenAI. The chart below traces the path end to end: a small, random slice of live Notion traffic is split evenly across models, each outcome is scored by an ensemble of independent judges, and the surviving traces become the metrics shown throughout.<br>Many of the differences we measure are smaller than their confidence intervals, so we’ve deliberately avoided creating a leaderboard. The data on this site is presented to make that visible. Let us know how we can improve: [email protected].
All weekly Notion agent trafficSample: 7% of all tokensExperiment gateEven, random split of users across modelsM1M2M3M15...All models on “High” effort, or closest equivalentDropped: Opt-outs, model switches, harness errorsEnsemble judgeOpus 4.8Gemini 3.5 FlashGPT-5.5Majority vote: 2-of-3 decidesDropped: Unfair to judge, missing judge outputsSurviving tracesScore: 1Score: 0Product metricsCost, average resolution time,undo/copy ratesAll weekly Notion agent trafficSample: 7% of all tokensExperiment gateEven, random split of usersacross modelsM1M2M3M15...All models on “High” effort,or closest equivalentDropped: Opt-outs, model switches,harness errorsEnsemble judgeOpus4.8Gemini3.5 FlashGPT-5.5Majority vote: 2-of-3 decidesDropped: Unfair to judge,missing judge outputsSurviving tracesScore: 1Score: 0Product metricsCost, average resolution time,undo/copy rates<br>The charts below offer two ways to read the same data. The first compares models on one metric at a time. (Note that overlapping confidence intervals mean models tie; our Knowledge Board offers a live snapshot of model performance, not a leaderboard.) The second chart plots quality against cost, so you can spot the models that deliver the most resolved work for the least money or time.
Resolution rate$/taskResolution time
Model performance in Notion, based on percentage of knowledge work tasks resolved across CX, project management, and more.<br>93%94%95%96%97%98%99%Opus 5Kimi K3Opus 4.8GPT-5.6-SolGrok 4.5GPT-5.5Opus 4.7Sonnet 5GPT-5.4GLM-5.2GPT-5.6-TerraGPT-5.6-LunaSonnet 4.6Kimi K2.7 CodeDeepSeek V4 ProResolution rate by model, with 95% confidence intervals.ModelEstimate95% CI lower95% CI upperOpus 598.4%98.2%98.6%Kimi K398.2%98%98.4%Opus 4.898%97.7%98.3%GPT-5.6-Sol97.4%97.1%97.6%Grok 4.597.2%96.9%97.5%GPT-5.597.2%96.8%97.6%Opus 4.796.8%96.4%97.3%Sonnet 596.7%96.3%97.1%GPT-5.496.6%96.2%97%GLM-5.296.5%96.2%96.8%GPT-5.6-Terra95.5%95.2%95.9%GPT-5.6-Luna95.3%94.8%95.8%Sonnet 4.694.6%94.1%95.2%Kimi K2.7 Code94.1%93.5%94.7%DeepSeek V4 Pro93.7%93.1%94.2%
Resolution rate, 95% CI
$/taskResolution time
Cost per task in USD for models to resolve knowledge work tasks.<br>$0.00$0.25$0.50$0.75$1.00Resolution rate98%96%94%Opus 5Kimi K3Opus 4.8GPT-5.6-SolGrok 4.5GPT-5.5Opus 4.7Sonnet 5GPT-5.4GLM-5.2GPT-5.6-LunaGPT-5.6-TerraSonnet 4.6Kimi K2.7 CodeDeepSeek V4 ProResolution rate versus $/task by model, with 95% confidence intervals.ModelResolution rate estimate95% CI lower95% CI upper$/task estimate95% CI lower95% CI upperProbability of being on the frontierOpus 598.4%98.2%98.6%$0.87$0.86$0.8985%Kimi K398.2%98%98.4%$0.45$0.44$0.46100%Opus 4.898%97.7%98.3%$0.71$0.68$0.7413%GPT-5.6-Sol97.4%97.1%97.6%$0.64$0.61$0.670%Grok 4.597.2%96.9%97.5%$0.47$0.45$0.506%GPT-5.597.2%96.8%97.6%$0.62$0.58$0.650%Opus 4.796.8%96.4%97.3%$0.70$0.66$0.740%Sonnet 596.7%96.3%97.1%$0.34$0.32$0.3576%GPT-5.496.6%96.2%97%$0.38$0.36$0.4036%GLM-5.296.5%96.2%96.8%$0.17$0.16$0.19100%GPT-5.6-Luna95.3%94.8%95.8%$0.02$0.02$0.03100%GPT-5.6-Terra95.5%95.2%95.9%$0.30$0.29$0.310%Sonnet 4.694.6%94.1%95.2%$0.39$0.37$0.410%Kimi K2.7 Code94.1%93.5%94.7%$0.17$0.16$0.190%DeepSeek V4 Pro93.7%93.1%94.2%$0.28$0.27$0.300%
$/task
Methodology<br>Context<br>Our goal with this work is to give our customers a helpful resource for deciding which model best fits their work, along with a tool for understanding the trade-offs between cost, average resolution time, and performance.<br>Many rigorous benchmarks already exist, but none capture the full range of industries that Notion customers use Notion AI for, nor do they cleanly reflect the distribution of tasks that people delegate to the agent inside Notion.<br>So we built the Notion Knowledge Board. It evaluates models in Notion’s production environment, using real user...