I checked 30 frontier model cards. Here are the benchmarks labs report

ktwu013 pts0 comments

Benchmark Radar

Skip to content

Daily briefing

Questions for today

Search

Scan date

Kind

Category

Source

Organization

Event

Clear filters

Matching observations

Show more results

Sources

Corpus totals

What the two layers say

Stated findings

Benchmarks to watch

Choose a reporting story

Start with signals that distinguish emerging instruments from established<br>standards and saturated conventions.

Reporting over time

Benchmark adoption frontier

All tracked benchmarks

Three readings on one time axis. Each orange diamond marks an organization's first<br>report of this benchmark, which is the only event that raises the cumulative count.<br>The rug beneath it puts one tick per dated model card, so later cards from an<br>organization already counted appear as gray ticks and leave the staircase flat. A card<br>with no publication date cannot be placed on the timeline and is absent from both<br>bands, though it is still counted in the totals above. A long flat run is<br>reporting saturation observed within this curated registry, not a claim about<br>benchmark score saturation.<br>The score track below it is a separate reading: every value that could be read<br>verbatim from a cited document, connected only where the instrument and protocol<br>are identical. A flat score tail usually means no newer number could be read, which<br>is why the gap is marked rather than drawn through.

Frontier milestones

What would make this a true Pareto frontier?

Comparable score observations need the benchmark version and split,<br>metric direction, model, harness or scaffold, reasoning budget, cost or<br>latency, publication date, and source. Only compatible configurations can<br>share a score frontier; this registry currently stores mentions, not those<br>measurements.

With those observations, a Harbor-style view can put cost or latency on the<br>x-axis and score on the y-axis, connect only nondominated observations, and<br>use a publication-time slider to reveal how the frontier moved.

Benchmarks by model card adoption

Search

Domain

Organization

Benchmark released

Clear filters

Each model card counts once per benchmark. A card reporting AIME in four<br>configurations counts the same as a card reporting it once, so a long appendix<br>cannot outweigh a different vendor. Organizations breaks the tie: the same count<br>from six vendors is a shared standard, from one vendor a house style.

Audit the counts<br>Model cards in the registry

The curated source list this ranking is computed from. Expand any card to see<br>every benchmark it reports, grouped the way the source document groups them, so<br>our data can be checked line by line against the original.

Select a node

Inspect a relationship

Topic, source, and organization nodes set the corresponding Today filter.<br>Artifact nodes set the date and title search.

New by domain

Daily evidence and attention volume

Category tags overlap. Each bar is an independent count, not a part of a stacked total.

Releases only

Excludes records re-announced as an update to something already surfaced.

Daily ledger

Source mix counts ranked evidence after scoring. Fetch health counts raw records<br>returned before scoring, so a source can be ok and still empty.

Date<br>Coverage (UTC)<br>Evidence<br>Source mix<br>Categories<br>Events<br>Attention<br>Fetch health

Dashboard unavailable

The validated data file could not be loaded.

Try refreshing, or inspect the latest daily Issue while the dashboard rebuilds.

Open daily Issues ↗

benchmark source from card frontier model

Related Articles