AI Insurance Leaderboard | Coverage Cat
Coverage Cat AI Insurance Benchmark
Which AI models understand insurance work?
We evaluate frontier models on two separate insurance tasks: price estimation<br>for anonymized umbrella quote rows, and brokerage/agent task reasoning against<br>benchmark reference answers for underwriting, eligibility, and coverage questions.
Explore rankings<br>View example questions
Price leader<br>Grok 4.3
Brokerage/Agent leader<br>DeepSeek 3.2
Models tested
Total battles<br>59,619
Leaderboard
Price-estimation performance
Price-estimation results compare predicted premiums and uncertainty ranges against actual quote outcomes. These metrics are separate from the brokerage/agent task leaderboard.
Best quote Elo<br>Grok 4.3<br>1,687 Elo
Best coverage<br>ChatGPT 5.5<br>69.1% coverage
Model quote responses<br>19,873
2,839 unique quote cases
1,687
Grok 4.3
1,577
ChatGPT 5.5
1,547
Claude Opus 4.7
1,462
GLM 5
1,451
Kimi K2.5
1,406
Mistral Large
1,370
DeepSeek 3.2
45.4%
Grok 4.3
53.9%
ChatGPT 5.5
54.2%
Claude Opus 4.7
58.7%
GLM 5
44.4%
Kimi K2.5
44.3%
Mistral Large
49.1%
DeepSeek 3.2
45.6%
Grok 4.3
69.1%
ChatGPT 5.5
54.1%
Claude Opus 4.7
44.4%
GLM 5
48.9%
Kimi K2.5
41.0%
Mistral Large
39.9%
DeepSeek 3.2
49.5%
Grok 4.3
94.0%
ChatGPT 5.5
49.5%
Claude Opus 4.7
51.8%
GLM 5
62.3%
Kimi K2.5
65.7%
Mistral Large
54.1%
DeepSeek 3.2
2880.00
Grok 4.3
3311.46
ChatGPT 5.5
2393.84
Claude Opus 4.7
2685.73
GLM 5
3175.43
Kimi K2.5
2942.89
Mistral Large
2779.78
DeepSeek 3.2
Model comparison
Price-estimation ranking
Rank<br>Model<br>Elo<br>Win rate
Coverage<br>Quote MAPE<br>Winkler loss
Record
Grok 4.3<br>xAI
1,687<br>45.4%
45.6%<br>49.5%<br>2880.00
7717-9290-27
ChatGPT 5.5<br>OpenAI
1,577<br>53.9%
69.1%<br>94.0%<br>3311.46
9162-7831-41
Claude Opus 4.7<br>Claude
1,547<br>54.2%
54.1%<br>49.5%<br>2393.84
9212-7765-57
GLM 5<br>Z.ai
1,462<br>58.7%
44.4%<br>51.8%<br>2685.73
9985-7035-14
Kimi K2.5<br>Moonshot AI
1,451<br>44.4%
48.9%<br>62.3%<br>3175.43
7519-9420-95
Mistral Large<br>Mistral
1,406<br>44.3%
41.0%<br>65.7%<br>2942.89
7466-9402-166
DeepSeek 3.2<br>DeepSeek
1,370<br>49.1%
39.9%<br>54.1%<br>2779.78
8266-8584-184
Eval examples
Two different benchmark tasks
Price-estimation rows are scored against actual quote outcomes. Brokerage/agent task<br>rows are scored against reference answers and judged separately, so their leaderboard<br>should be read as answer-quality performance rather than premium-estimation performance.
Price-estimation examples
Estimate the annual premium and uncertainty range for an anonymized $1M California umbrella quote from a specific carrier.
Given state, carrier, coverage limit, and anonymized risk features, return calibrated P10/P50/P90 premium estimates.
Predict a quote range that contains the actual annualized premium without making the interval unnecessarily wide.
Brokerage/agent task examples
A household has a listed underwriting profile. Are they likely to be eligible with a specific umbrella carrier?
In Texas, how much more does moving from $1M to $2M of umbrella coverage typically cost with a named carrier?
A customer asks about coverage requirements or eligibility constraints. What should an assistant say, using the benchmark reference answer?
Methodology
Domain-specific, aggregate-only benchmarking
General AI benchmarks rarely measure whether a model can reason through the details that<br>matter in insurance: liability limits, carrier constraints, premium ranges, eligibility<br>rules, and uncertainty. This benchmark focuses on those workflows.
Price scoring
Quote rows compare each model's estimated annual premium and range against the actual<br>quote outcome. Coverage rewards calibrated ranges; MAPE rewards accurate point estimates;<br>Winkler loss penalizes ranges that miss the actual quote or are too wide.
Brokerage/agent task scoring
Brokerage/agent task rows compare model answers to benchmark reference answers with<br>AI judging. The public task view reports aggregate judge scores, pairwise Elo, and win<br>rate only.
How Elo works
For each shared scenario or question, every pair of model outputs is compared. Better<br>outputs win the local battle, ties split credit, and Elo updates model strength within<br>that benchmark section.
Data protection
Public results are aggregate-only. The page does not expose raw prompts, row identifiers,<br>model responses, judge reasoning, or any operational eval artifacts. The evals use<br>anonymized data on no-retention and no-logging platforms, so customer data is never<br>exposed even to model providers.