Little Dorrit Editor Benchmark Leaderboard
Top Model
Loading...
Best F1 Score
Loading...
Best Precision
Loading...
Best Recall
Loading...
Performance Leaderboard
Vendor
All vendors
Release Year
All years
Clear filters
Most models listed above were tested through OpenRouter,<br>a unified interface for LLMs. The GPT models were also tested on the api.openai.com<br>endpoint, so the OpenRouter versions are suffixed with "OR".
* Claude-family results use the same benchmark prompt, but the page images are<br>resized and recompressed as needed to stay below Anthropic's 5 MiB per-image limit.
The notes below are not comprehensive, but document some of our observations<br>during testing.
Claude-family models are marked with * on the leaderboard because<br>their inputs are resized and recompressed to fit Anthropic's 5 MiB<br>image limit before inference.
Grok 2 Vision 1212 and OpenAI o1-pro continuously fail with<br>errors in response indicating that a valid JSON was not returned.
Error: 'NoneType' object is not subscriptable
Phi 4 Multimodal Instruct and Qwen VL Plus repeatedly returned<br>unparseable JSON:
Error: Could not parse JSON from model response: Could not<br>extract valid JSON from response: ...
Detailed Model Performance
Further Contents
About the Benchmark
How It Works
The Task
Example Input
Expected Output
About the Benchmark
This benchmark evaluates the ability of multimodal language models to interpret handwritten editorial corrections in printed text.<br>Using annotated scans from Charles Dickens' "Little Dorrit," we challenge models to accurately capture human editing intentions.
Models are assessed on their ability to detect and interpret various types of editorial marks including insertions, deletions,<br>replacements, punctuation changes, capitalization corrections, and text italicization.
The benchmark consists of several document pages containing Dickens' original manuscript with editorial marks and corrections.<br>These documents represent the kind of annotations an editor might make when reviewing a text for publication, providing<br>a realistic test of how well AI models can understand human editing practices.
How It Works
Models are presented with scanned pages containing handwritten editorial<br>marks and are asked to identify each edit, its type, location, and the<br>text before and after the edit.
The Task
Each page in the benchmark consists of printed text overlaid with<br>handwritten editorial annotations. Models are tasked with detecting and<br>interpreting all such editorial corrections and outputting them in<br>structured JSON format.
For each correction, the model must identify:
Type of edit: One of insertion, deletion, replacement, punctuation, capitalization, or italicize
Original text: The text as it appeared before the edit
Corrected text: The intended version after applying the edit
Line number: The line on which the edit occurs. Use line 0 for titles or headings, and start counting full lines of body text from line 1
Page: A known identifier for the image (e.g., "001.png"), provided alongside the input and not extracted by the model
Models should infer the intent behind handwritten annotations using both<br>visual and textual cues. Common markup conventions include:
Insertions: Indicated by caret marks (^) or added words between lines
Deletions: Shown using strikethroughs or crossed-out text
Replacements: Circled, underlined, or bracketed text with substitutions nearby
Punctuation edits: Handwritten punctuation added, removed, or modified
Capitalization: Case changes marked explicitly or via notation
Italicize: Text that should be formatted in italics, typically indicated by underlining or special notation
This task combines fine-grained visual recognition with natural language<br>understanding and domain knowledge of editorial conventions. The goal is<br>not just OCR or layout detection, but true interpretation of handwritten<br>edits in context.
Example Input
Sample page from Little Dorrit with editorial marks.
Expected Output
For the example above, models should identify all ten editorial corrections, producing output like:
"image": "001.png",<br>"page_number": 5,<br>"source": "Little Dorrit",<br>"annotator": "pairsys",<br>"annotation_date": "2025-04-04",<br>"verified": true,<br>"edits": [<br>"type": "punctuation",<br>"original_text": "church bells",<br>"corrected_text": "church bells,",<br>"line_number": 2,<br>"page": "001.png"<br>},<br>"type": "punctuation",<br>"original_text": "wine bottles",<br>"corrected_text": "wine-bottles",<br>"line_number": 11,<br>"page": "001.png"<br>},<br>"type": "punctuation",<br>"original_text": "got through",<br>"corrected_text": "got, through",<br>"line_number": 14,<br>"page": "001.png"<br>},<br>"type": "punctuation",<br>"original_text": "iron bars fashioned",<br>"corrected_text": "iron bars, fashioned",<br>"line_number": 14,<br>"page": "001.png"<br>},<br>"type": "punctuation",<br>"original_text": "grating where",<br>"corrected_text": "grating, where",<br>"line_number": 17,<br>"page": "001.png"<br>},<br>"type": "punctuation",<br>"original_text": "outside...