Which AI models understand insurance work?

botacode1 pts0 comments

AI Insurance Leaderboard | Coverage Cat

Coverage Cat AI Insurance Benchmark

Which AI models understand insurance work?

We evaluate frontier models on two separate insurance tasks: price estimation<br>for anonymized umbrella quote rows, and brokerage/agent task reasoning against<br>benchmark reference answers for underwriting, eligibility, and coverage questions.

Explore rankings<br>View example questions

Price leader<br>Grok 4.3

Brokerage/Agent leader<br>DeepSeek 3.2

Models tested

Total battles<br>59,619

Leaderboard

Price-estimation performance

Price-estimation results compare predicted premiums and uncertainty ranges against actual quote outcomes. These metrics are separate from the brokerage/agent task leaderboard.

Best quote Elo<br>Grok 4.3<br>1,687 Elo

Best coverage<br>ChatGPT 5.5<br>69.1% coverage

Model quote responses<br>19,873

2,839 unique quote cases

1,687

Grok 4.3

1,577

ChatGPT 5.5

1,547

Claude Opus 4.7

1,462

GLM 5

1,451

Kimi K2.5

1,406

Mistral Large

1,370

DeepSeek 3.2

45.4%

Grok 4.3

53.9%

ChatGPT 5.5

54.2%

Claude Opus 4.7

58.7%

GLM 5

44.4%

Kimi K2.5

44.3%

Mistral Large

49.1%

DeepSeek 3.2

45.6%

Grok 4.3

69.1%

ChatGPT 5.5

54.1%

Claude Opus 4.7

44.4%

GLM 5

48.9%

Kimi K2.5

41.0%

Mistral Large

39.9%

DeepSeek 3.2

49.5%

Grok 4.3

94.0%

ChatGPT 5.5

49.5%

Claude Opus 4.7

51.8%

GLM 5

62.3%

Kimi K2.5

65.7%

Mistral Large

54.1%

DeepSeek 3.2

2880.00

Grok 4.3

3311.46

ChatGPT 5.5

2393.84

Claude Opus 4.7

2685.73

GLM 5

3175.43

Kimi K2.5

2942.89

Mistral Large

2779.78

DeepSeek 3.2

Model comparison

Price-estimation ranking

Rank<br>Model<br>Elo<br>Win rate

Coverage<br>Quote MAPE<br>Winkler loss

Record

Grok 4.3<br>xAI

1,687<br>45.4%

45.6%<br>49.5%<br>2880.00

7717-9290-27

ChatGPT 5.5<br>OpenAI

1,577<br>53.9%

69.1%<br>94.0%<br>3311.46

9162-7831-41

Claude Opus 4.7<br>Claude

1,547<br>54.2%

54.1%<br>49.5%<br>2393.84

9212-7765-57

GLM 5<br>Z.ai

1,462<br>58.7%

44.4%<br>51.8%<br>2685.73

9985-7035-14

Kimi K2.5<br>Moonshot AI

1,451<br>44.4%

48.9%<br>62.3%<br>3175.43

7519-9420-95

Mistral Large<br>Mistral

1,406<br>44.3%

41.0%<br>65.7%<br>2942.89

7466-9402-166

DeepSeek 3.2<br>DeepSeek

1,370<br>49.1%

39.9%<br>54.1%<br>2779.78

8266-8584-184

Eval examples

Two different benchmark tasks

Price-estimation rows are scored against actual quote outcomes. Brokerage/agent task<br>rows are scored against reference answers and judged separately, so their leaderboard<br>should be read as answer-quality performance rather than premium-estimation performance.

Price-estimation examples

Estimate the annual premium and uncertainty range for an anonymized $1M California umbrella quote from a specific carrier.

Given state, carrier, coverage limit, and anonymized risk features, return calibrated P10/P50/P90 premium estimates.

Predict a quote range that contains the actual annualized premium without making the interval unnecessarily wide.

Brokerage/agent task examples

A household has a listed underwriting profile. Are they likely to be eligible with a specific umbrella carrier?

In Texas, how much more does moving from $1M to $2M of umbrella coverage typically cost with a named carrier?

A customer asks about coverage requirements or eligibility constraints. What should an assistant say, using the benchmark reference answer?

Methodology

Domain-specific, aggregate-only benchmarking

General AI benchmarks rarely measure whether a model can reason through the details that<br>matter in insurance: liability limits, carrier constraints, premium ranges, eligibility<br>rules, and uncertainty. This benchmark focuses on those workflows.

Price scoring

Quote rows compare each model's estimated annual premium and range against the actual<br>quote outcome. Coverage rewards calibrated ranges; MAPE rewards accurate point estimates;<br>Winkler loss penalizes ranges that miss the actual quote or are too wide.

Brokerage/agent task scoring

Brokerage/agent task rows compare model answers to benchmark reference answers with<br>AI judging. The public task view reports aggregate judge scores, pairwise Elo, and win<br>rate only.

How Elo works

For each shared scenario or question, every pair of model outputs is compared. Better<br>outputs win the local battle, ties split credit, and Elo updates model strength within<br>that benchmark section.

Data protection

Public results are aggregate-only. The page does not expose raw prompts, row identifiers,<br>model responses, judge reasoning, or any operational eval artifacts. The evals use<br>anonymized data on no-retention and no-logging platforms, so customer data is never<br>exposed even to model providers.

quote coverage model price grok deepseek

Related Articles