Benchmark Radar
Skip to content
Daily briefing
Questions for today
Search
Scan date
Kind
Category
Source
Organization
Event
Clear filters
Matching observations
Show more results
Sources
Corpus totals
What the two layers say
Stated findings
Benchmarks to watch
Choose a reporting story
Start with signals that distinguish emerging instruments from established<br>standards and saturated conventions.
Reporting over time
Benchmark adoption frontier
All tracked benchmarks
Three readings on one time axis. Each orange diamond marks an organization's first<br>report of this benchmark, which is the only event that raises the cumulative count.<br>The rug beneath it puts one tick per dated model card, so later cards from an<br>organization already counted appear as gray ticks and leave the staircase flat. A card<br>with no publication date cannot be placed on the timeline and is absent from both<br>bands, though it is still counted in the totals above. A long flat run is<br>reporting saturation observed within this curated registry, not a claim about<br>benchmark score saturation.<br>The score track below it is a separate reading: every value that could be read<br>verbatim from a cited document, connected only where the instrument and protocol<br>are identical. A flat score tail usually means no newer number could be read, which<br>is why the gap is marked rather than drawn through.
Frontier milestones
What would make this a true Pareto frontier?
Comparable score observations need the benchmark version and split,<br>metric direction, model, harness or scaffold, reasoning budget, cost or<br>latency, publication date, and source. Only compatible configurations can<br>share a score frontier; this registry currently stores mentions, not those<br>measurements.
With those observations, a Harbor-style view can put cost or latency on the<br>x-axis and score on the y-axis, connect only nondominated observations, and<br>use a publication-time slider to reveal how the frontier moved.
Benchmarks by model card adoption
Search
Domain
Organization
Benchmark released
Clear filters
Each model card counts once per benchmark. A card reporting AIME in four<br>configurations counts the same as a card reporting it once, so a long appendix<br>cannot outweigh a different vendor. Organizations breaks the tie: the same count<br>from six vendors is a shared standard, from one vendor a house style.
Audit the counts<br>Model cards in the registry
The curated source list this ranking is computed from. Expand any card to see<br>every benchmark it reports, grouped the way the source document groups them, so<br>our data can be checked line by line against the original.
Select a node
Inspect a relationship
Topic, source, and organization nodes set the corresponding Today filter.<br>Artifact nodes set the date and title search.
New by domain
Daily evidence and attention volume
Category tags overlap. Each bar is an independent count, not a part of a stacked total.
Releases only
Excludes records re-announced as an update to something already surfaced.
Daily ledger
Source mix counts ranked evidence after scoring. Fetch health counts raw records<br>returned before scoring, so a source can be ok and still empty.
Date<br>Coverage (UTC)<br>Evidence<br>Source mix<br>Categories<br>Events<br>Attention<br>Fetch health
Dashboard unavailable
The validated data file could not be loaded.
Try refreshing, or inspect the latest daily Issue while the dashboard rebuilds.
Open daily Issues ↗