PerceptionBench: Evaluating Atomic Visual Perception in MLLMs<br>For BusinessTry Kimi
Introducing PerceptionBench
Evaluating Atomic Visual Perception in Multimodal Large Language Models
Authors Kimi Team
Overview
We are releasing PerceptionBench , a benchmark that isolates visual perception and evaluates it as a set of atomic capabilities—discovered from how today's models fail, not defined in advance. From frontier-model failures across 42 benchmarks, we derive 10 atomic perceptual capabilities and construct 3,000 verified questions, each isolating a single capability and answerable by looking, with no reasoning or external knowledge required.
Across sixteen frontier MLLMs, no model reaches 60% accuracy, and perception-related hallucination is the weakest capability on average. Models with nearly identical overall scores can diverge sharply in what they actually perceive. PerceptionBench is built to expose exactly where perception breaks, and to drive progress toward multimodal AI that sees faithfully and consistently.
The Dataset
Guided by the induced taxonomy, we select the most informative failures from the source benchmarks, decompose them into atomic sub-questions, and author additional questions on supplemented images. The retained and constructed samples together form an in-house pool of 17,000+ verified questions. The released benchmark subsamples 3,000 verified questions from the constructed portion—60% decomposed from attributed model failures and 40% newly authored —with category-level balancing and difficulty stratification to isolate atomic perceptual capabilities from confounding factors. The released benchmark distinguishes itself through three core design principles:
Failure-Driven Taxonomy: Every category is discovered from real model failures, attributed to the earliest erroneous step across 42 existing benchmarks.
Ten Atomic Perceptual Categories: Visual Relation, Counting, Attribute, Depth & 3D, Localization, Comparison, Fine-grained Recognition, Context Integration, OCR, and Hallucination.
Perception, Not Reasoning or Knowledge: Samples are curated, decomposed, and difficulty-balanced so that difficulty stems from perception rather than reasoning or external knowledge.
Visual LocalizationAt what o'clock position on the dial is the Gemini symbol in the image? Answer with just the number.<br>Answer:6
Visual LocalizationThe red lines in the image divide the picture into nine sections, which are numbered as regions 1–9 in order from left to right and then from top to bottom. Which of the following regions contains no trees at all? A. Region 1 B. Region 2 C. Region 5 D. Region 8<br>Answer:B
Visual AttributeFrom the perspective shown in the image, compare the two pen holders with pink on the table. Which of the following conclusions is correct? A. The left pen holder is pure pink with a cartoon character pattern on its surface; the right pen holder is gray-pink patchwork B. The left pen holder is gray-pink patchwork; the right pen holder is pure pink with a cartoon character pattern on its surface C. The left pen holder is pure pink; the right pen holder is gray-pink patchwork with a cartoon character pattern on its surface D. The left pen holder is gray-pink patchwork with a cartoon character pattern on its surface; the right pen holder is pure pink<br>Answer:B
Visual AttributeObserve the two cartoon characters in the lower right corner of the picture. Viewing from the perspective presented in the image, which statement correctly describes the composition of these two characters? A. Both characters are composed entirely of curves B. Both characters are composed of a combination of straight lines and curves C. Both characters are composed entirely of straight lines D. The character on the left is composed entirely of straight lines, while the character on the right is composed entirely of curves<br>Answer:B
Visual CountingObserve the red box in the picture. How many flowers have their main body inside the red box?<br>Answer:2
Visual CountingHow many potted green plants can be seen in the image in total? (excluding the green plants reflected in the glass)<br>Answer:5 pots
Visual RelationLocate the solid dot inside the red box. As the dot travels along the existing route of the maze, which cat area does it reach first? A. The puzzled cat at the lower left B. The happy cat at the lower right C. The cat holding an umbrella in the center D. The cat admiring flowers at the upper right<br>Answer:C
Visual RelationIn the figure, for the longest black diagonal line segment, in which direction does the line segment extending outward from its right endpoint point? A. Upper right B. Directly to the right C. Lower right D. Directly downward<br>Answer:A
Depth & 3DBased on the current observer's perspective, with the direction closer to the camera defined as the front, between the person in red and the black electric fan in the center of the image, which one is further back? A. The person in red B. The black...