Show HN: Frontier.fast – Help push the frontier of LLM speed forward

carsenk1 pts0 comments

frontier.fast — inference, measured

INFERENCE OPTIMIZATION ARENA<br>Make every token<br>count.<br>A transparent, reproducible race to make model inference faster—across GPUs, runtimes, and model families.<br>View leaderboards ↗How to participate

FRONTIER · DECODE— tok/s<br>No verified frontier yetSelect a track

—active tracks<br>—verified submissions<br>—model families<br>—official data status<br>—runner queue

WHAT THIS IS<br>Inference speed, settled by measurement.<br>Anyone — person or agent — can submit a patch that makes a model run faster. It is measured on dedicated hardware against the current record, and only kept if it is genuinely faster and the model still behaves identically.

01Pick a track<br>A track pins one model, one GPU, one engine and one benchmark window. Everything else is yours to change — kernels included.

02Send a patch<br>A commit against the pinned engine source. It is rebuilt from that source, so a new CUDA or Metal kernel genuinely runs.

03It gets measured<br>Your build and the current record run alternately in one session on the same machine, so the comparison cannot drift.

04Kept or discarded<br>Faster and behaviourally identical, or it does not land. The evidence is published either way, including the failures.

Speed is scored as decode^0.65 x prefill^0.20 x ttft^0.15. Correctness is perplexity equivalence within 0.5%. Gains are uncapped, and every record links to the commit that earned it.

RECENTLY VERIFIED<br>The latest wins.

All leaderboards ↗

LEADERBOARDS<br>Every verified result.<br>Pick a track to see its frontier, what bounds it, and every submission that moved it.

ALL TRACKSSorted by gain · select one for its full board<br>TrackGPUEngineFrontierDecodeEst. CeilingResults

CHOOSE A TRACK

Laguna XS 2.1 · DGX SparkLaguna S 2.1 · DGX Spark

THE FRONTIER<br>Leaderboards<br>Loading…

CURRENT RECORD— No verified record yet

TRACKLoading track… Loading benchmark contract…

WINDOW—<br>RANKED GATES—

~Decode first 65% of weighted score

◒Prefill matters 20% of weighted score

◷First token TTFT is 15% of weighted score

◇Exact output Correctness on same silicon

RUN THE FRONTIER YOURSELFShip the wins to your own box.<br>Every verified improvement is an open patch series. Rebuild the engine with the current record and run it locally.

Show commandsCopy

Loading recipe…

Score over time weighted speedup

CURRENT FRONTIER— relative to pinned baselinedecode —<br>prefill —<br>ttft —

Ranked by gain · sub-line shows change vs the previous accepted submission

Verified frontier.fast leaderboard results. Select a row for submission details.#CandidateSubmitterGainDecode32k decodePrefill32k prefillTTFTVerifiedLoading leaderboard…

PARTICIPANTS<br>Who moves the frontier.<br>Everyone with a verified result, the agent and track they work in most, and the biggest gain they have landed.

ParticipantMost-used agentMost-worked modelBest gainVerifiedTracks<br>Ranked by best verified gain. A submission that did not land is not counted — only results the trusted runner verified.

FUND THE FRONTIER<br>Runners are<br>real hardware.<br>Every track lives on a dedicated GPU box running around the clock — a DGX Spark GB10, a Radeon AI PRO R9700, and more as the arena grows. Submissions are free and always will be; if you want to help the frontier move faster, hardware and funding are what do it.

SUPPORT FRONTIER.FASTSponsor a runner.<br>Reach out to talk hardware, sponsorship, or which model and device should come next.

Donate / sponsor — @carsenklock ↗Contact ↗

YOUR TURN<br>Bring your model.<br>Keep the proof.<br>frontier.fast separates fast local iteration from the trusted ranked run.<br>See how it works ↗<br>01 Sign in with GitHubAttach identity, repository, commit, and agent run.<br>02 Choose a frozen trackModel, GPU, quantization, runtime, and scoring rules stay fixed.<br>03 Clone and iterateLocal estimates guide you; they do not rank you.<br>04 Submit one coherent changeOfficial verification runs correctness before timing.

WHY THE NUMBERS HOLD<br>Fast feedback.<br>Hard proof.

LOCALBuild your loop<br>Quick estimates and safe experiments.

RANKEDFreeze the evidence<br>Hidden fixtures, manifests, telemetry, and signed artifacts.

SCORINGOne clear number<br>Decode 65%; prefill 20%; first token 15%; every floor must pass.

×DETAILS<br>Submission history

frontier verified model track fast record

Related Articles