Show HN: Human Benchmark – Compare your reasoning skills against AI models

twoslide1 pts0 comments

Human Benchmark — how do you score against the machines?

Answer questions used to measure AI reasoning abilities

Difficulty adapts according to your ability

Responses are timed

Five questions will give a good sense of how you compare (and your potential earnings)

Start

Time remaining—

Question 1—

Submit answer

Next item<br>Finish

That's your measure

Take it again

The Leaderboard

— models

Shaded bars are confidence ranges.

— items

Settings

Time per question<br>1m 00s

Run out of time and the question is marked wrong.

Changes apply to the next question.

About

Question Sources

CommonsenseQA — everyday common sense that most people find easy.

GSM8K — grade-school arithmetic word problems.

AGIEval (LSAT logical reasoning) — real law-school admission test items: read a short argument, spot the flaw or the assumption.

AGIEval (AQuA-RAT) — GMAT and GRE style quantitative word problems.

BIG-Bench Hard — logical deduction, object counting, date arithmetic, tracking things as they move, causal judgement, and truth-teller puzzles.

Methodology

The app uses a Rasch model to estimate a percentage score from a relatively small number of items. Information

Reading the range

The shaded bar is a 95% confidence range. It starts enormous and narrows with every answer. If two bars overlap heavily, the difference between them is within the margin of error.

Disclaimers

This is a proof of concept and illustration of how LLM benchmarking works in a human context, but it is not a serious test of reasoning or intelligence and has limited psychometric validty.

question reasoning human answer time items

Related Articles