Human Benchmark — how do you score against the machines?
Answer questions used to measure AI reasoning abilities
Difficulty adapts according to your ability
Responses are timed
Five questions will give a good sense of how you compare (and your potential earnings)
Start
Time remaining—
Question 1—
Submit answer
Next item<br>Finish
That's your measure
Take it again
The Leaderboard
— models
Shaded bars are confidence ranges.
— items
Settings
Time per question<br>1m 00s
Run out of time and the question is marked wrong.
Changes apply to the next question.
About
Question Sources
CommonsenseQA — everyday common sense that most people find easy.
GSM8K — grade-school arithmetic word problems.
AGIEval (LSAT logical reasoning) — real law-school admission test items: read a short argument, spot the flaw or the assumption.
AGIEval (AQuA-RAT) — GMAT and GRE style quantitative word problems.
BIG-Bench Hard — logical deduction, object counting, date arithmetic, tracking things as they move, causal judgement, and truth-teller puzzles.
Methodology
The app uses a Rasch model to estimate a percentage score from a relatively small number of items. Information
Reading the range
The shaded bar is a 95% confidence range. It starts enormous and narrows with every answer. If two bars overlap heavily, the difference between them is within the margin of error.
Disclaimers
This is a proof of concept and illustration of how LLM benchmarking works in a human context, but it is not a serious test of reasoning or intelligence and has limited psychometric validty.