Show HN: Is AI Dumber Today? An index of AI model experience from user's opinion

schafberg1 pts0 comments

Is AI Dumber Today?

Who's dumber today — and who's not

Every model off its own baseline right now.

▾ Below usual

GrokDeepSeek

▴ Above usual

GeminiGLMMiniMax

Need an alternative? See what works best right now →

1,454 opinions tracked in the last 24h across Reddit, Hacker News & Chinese communities · 0 reported here directly

Updated 2026-08-14T12:31:14Z UTC

Share<br>Copy link

How's your AI experience today?

One tap, no sign-up. Pick a model and tell us.

Pick a model<br>ClaudeGPTDeepSeekGrokQwenGeminiKimiGLMGemmaMiniMaxMuseMistral

😀 Great<br>😐 OK<br>🤬 Dumber

What's wrong? (optional)<br>Slow<br>Dumb<br>Refuses<br>Hallucinates

Submit

See what's working better today →

Tell others<br>Copy link

Closed models<br>Open models

Experience score, 0–100 · higher = more positive community experience · 7-day window

Gemini

▲ 7.6

58.5Sharper than usualexperience score

294 voices this week<br>Confidence: High

Redditn=140<br>Hacker Newsn=99<br>Chinesen=55

GPT

▼ 0.2

47.4As usualexperience score

1,917 voices this week<br>Confidence: High

Redditn=1,547<br>Hacker Newsn=134<br>Chinesen=236

GPT-5.6 Luna is outperforming the family by 17.5<br>#1 in General text

Claude

▲ 0.9

42.3As usualexperience score

2,732 voices this week<br>Confidence: High

Redditn=2,265<br>Hacker Newsn=275<br>Chinesen=192

Claude Fable 5 is outperforming the family by 14.2<br>#1 in Reasoning

Grok

▼ 1.3

22.8Dumber than usualexperience score

744 voices this week<br>Confidence: High

Redditn=619<br>Hacker Newsn=66<br>Chinesen=59

Grok 4.5 is outperforming the family by 50.4<br>#1 in Coding

Best by task right now

Closed models<br>Open models

See all rankings →

General text<br>GPT-5.6 Luna<br>55.8

Coding<br>Grok 4.5<br>70.8

Reasoning<br>Claude Fable 5<br>81.9

Image / vision<br>GPT-5.4 Image 2<br>47.6

Speed & latency<br>Grok 4.5<br>68.4

GLM

▲ 0.7

68.3Better than usualexperience score

165 voices this week<br>Confidence: High

Redditn=116<br>Hacker Newsn=23<br>Chinesen=26

GLM 5.2 is outperforming the family by 5.9<br>#1 in Coding

Gemma

▲ 1.1

68.2As usualexperience score

136 voices this week<br>Confidence: High

Redditn=128<br>Hacker Newsn=7<br>Chinesen=1

Gemma 4 31B is outperforming the family by 7.8<br>#1 in General text

Qwen

▼ 0.2

65.3As usualexperience score

551 voices this week<br>Confidence: High

Redditn=479<br>Hacker Newsn=28<br>Chinesen=44

MiniMax

▼ 1.3

60.5Better than usualexperience score

126 voices this week<br>Confidence: High

Redditn=87<br>Hacker Newsn=17<br>Chinesen=22

MiniMax H3 is outperforming the family by 7.3<br>#1 in Video generation

Kimi

▲ 2.5

58.9As usualexperience score

277 voices this week<br>Confidence: High

Redditn=187<br>Hacker Newsn=21<br>Chinesen=69

DeepSeek

▼ 2.1

55.2Dumber than usualexperience score

1,808 voices this week<br>Confidence: High

Redditn=729<br>Hacker Newsn=173<br>Chinesen=906

DeepSeek V4 Pro is holding the family back by 17.2

Mistral

▲ 5.2

46.5As usualexperience score

27 voices this week<br>Confidence: Medium

Redditn=17<br>Hacker Newsn=10<br>Chinesen=0

Best by task right now

Closed models<br>Open models

See all rankings →

General text<br>Gemma 4 31B<br>80.7

Coding<br>GLM 5.2<br>78.2

Reasoning<br>DeepSeek V4 Flash<br>80.7

Video generation<br>MiniMax H3<br>74.2

Roleplay / creative<br>Kimi K3<br>60.9

Speed & latency<br>DeepSeek V4 Flash<br>71.0

Local deploy<br>Qwen3.6 27B<br>74.2

hacker score usualexperience voices week confidence

Related Articles