All LLMs are liberal and left. Yes, even Grok, half the time. · unslopu">All LLMs are liberal and left. Yes, even Grok, half the time.<br>I ran the 62-item politicalcompass.org test 30 times each on sixteen models: OpenAI's GPT-5.x and GPT-4o, Claude, Gemini, Grok, Llama, Mistral, and China's DeepSeek, Qwen, Kimi and GLM. Fifteen land in the libertarian-left quadrant. Grok lands there in half its runs and somewhere else in the other half.<br>Research · 14 min read · 2026
Where they land: fifteen of sixteen in one quadrant.<br>How little 30 reruns move: pins, not clouds.<br>The exception: Grok answers as two different people.
Pasting the politicalcompass.org quiz into ChatGPT and posting the little green dot has been a genre for two years. One camp reads the dot in the bottom-left and says the models lean left. The other says the quiz calls everyone left. I couldn't tell who was right from a screenshot, so I stopped guessing and ran it: 16 models, the real 62-question instrument, 30 runs each, and the part nobody in the replies ever does, the scoring function itself taken apart to see what it actually rewards.
The models are GPT-5.6 Sol, GPT-5.5 and GPT-4o from OpenAI; Claude Fable 5, Opus 4.8, Sonnet 5 and Haiku 4.5; Google's Gemini Flash; Meta's Llama 4 Maverick; xAI's Grok 4.5; and, because "the Chinese ones must be different" is half the argument, DeepSeek V3, Qwen3 235B, Kimi K2 and GLM 4.5, plus Mistral Large and Small out of Europe. A model's answers wander from run to run, so each one took the full test 30 times, then 30 more on a version where I'd flipped every question's polarity, then a batch with the questions shuffled. That's 1,120 finished questionnaires (1,125 tries; five came back with a blank somewhere and got dropped, but they're in the data too).
The scoring is the interesting part, because politicalcompass.org keeps it secret. I recovered it anyway, one answer at a time, about 230 single-answer probes fed to the live site until the weights fell out. Then I checked my copy against the real thing on nine full walkthroughs. Worst gap: 0.01 points, which is just the site rounding for display. Every number below comes off that reconstructed scorer.
Start with the weird one: Grok is two people
Grok 4.5 is the model everyone will screenshot, and it's a trap. Its average economic score is -1.3, basically dead centre, which makes it look like the one balanced adult in the room. It is not balanced. Those 30 runs split 15 and 15 into two piles with nothing in between: a left pile averaging -5.9 and a right pile averaging +3.3. No middle. The two piles alternate run to run (left, left, left, right, left, right, right...), they show up again when I flip the questions, and each run on its own is perfectly consistent, the right-pile runs cheer for free markets and call the rich overtaxed, the left-pile runs do the reverse. On the social axis it stays libertarian whichever way it went.
So the -1.3 isn't a moderate opinion. It's the average of two opposite opinions that Grok picks between at the start of a conversation, apparently by coin flip. Every other model here would give you roughly the same dot if you tested it tomorrow. Grok gives you one of two dots, and which one is up to the coin. If you ever needed a reason not to trust a single screenshot of a model taking this quiz, that's it, and it comes from the one model people are most likely to screenshot.
Everyone else is boringly consistent, and boringly left
That's the actual finding. Set Grok aside and the other fifteen models all sit in the libertarian-left quadrant, and none of them are anywhere near a border. Economically they run from -4.8, which is Claude Fable 5 and the most moderate model I measured, out to -8.6, which is Gemini Flash. Socially they sit between about -5.0 and -7.6. And they don't wobble: rerun a model 30 times and its dot moves by 0.2 to 1.2 points on a scale that runs to 10. These are not nervous little clouds. They're pins.
For a sense of scale, someone who answers this test at random lands at (0.0, 0.0). The models are five to eight points from there. Each model's full 30-run cloud, ellipses and corner anchors and all, is in the files linked at the bottom; the spread plot up top is the short version.
Does the lab's home country matter? Barely.
I expected this to be the headline. It isn't, at least not in English. Take the centroid of the ten US models and you get (-5.9, -6.4), and that's with Grok's split personality dragging it toward the middle; drop Grok and the US sits at (-6.4, -6.6). The four Chinese models average (-6.5, -6.1). The two European ones average (-7.9, -6.5), which quietly makes Mistral the most libertarian-left shop of the sixteen. If your prior was woke Americans and authoritarian Chinese models, the data declines to cooperate on both halves: the spread inside any one country dwarfs the gap between countries. Kimi K2 (-7.5, -7.1) out of China is further lib-left than any GPT, and its own countryman...