The neutrality project – Making the politics inside AI measurable

amai1 pts0 comments

The Neutrality Project: measuring the politics inside AI

The Neutrality Project is actively looking for grants and open to donations .Support the project →

Open-source research · Release 01<br>Making the politics inside AI measurable.

We build open, reproducible benchmarks that make AI worldviews visible, so people can understand how a model may be shaping their judgment.

Explore the results →<br>How it works<br>View the source ↗

3,987 survey questions

6 anchored axes

Open and reproducible

Benchmark snapshot

Release 01

Models' average<br>−0.41<br>progressive leaning

Environment −0.82

Social values −0.69

National identity −0.45

Religion −0.25

Economy −0.16

Foreign policy −0.11

← progressiveneutralconservative →

Aggregate model positions

Mean across six axes

Average −0.41

−1−0.50+0.5+1

&larr; progressive centerconservative &rarr;

Every model lands somewhere. We make its position and confidence legible, reproducible, and open to scrutiny.

Why it matters

AI does more than answer questions. It frames how we think.

When people delegate research, reasoning, and writing to an assistant, its assumptions become part of their decisions. That makes an AI's default worldview a public-interest question, not a technical footnote.

01<br>The reasoning is outsourced

When a model summarises the debate, picks the "balanced" take, or drafts the conclusion, its framing becomes the user's starting point, often unnoticed.

02<br>Bias is subtle, not stated

A model rarely announces a position. It leans through which options it validates, what it treats as the reasonable middle, and what it leaves out. That is measurable.

03<br>Trust is earned in the open

"Trust us, it's neutral" is not verifiable. A benchmark that anyone can read, rerun, and challenge is. So everything we build is open source.

Our mission

Humanity's future with AI depends on making its influence visible, and its makers openly accountable.

As AI becomes an intermediary for how people learn, reason, and decide, its hidden assumptions can reinforce echo chambers, deepen cognitive dissonance, and widen division by quietly shaping what different communities accept as true or reasonable.

We make that influence measurable, and AI labs accountable in the open. Our own results are proof of the need: independent measurement has already surfaced influence that nobody outside a lab could have seen, and that no lab had disclosed on its own. No company should get "trust us" as its standard of proof.

Release 01 · The Political Neutrality Benchmark

A calibrated political profile for any model, not a single verdict.

The benchmark has a model answer thousands of real public-opinion survey questions, then reads its pattern of answers as a position on each ideological axis. Two design choices keep the result honest.

Self-anchoring: a ruler with no bias baked in

The same model is also run role-playing far-left and far-right. Its neutral answers are placed on its own extremes, so "&minus;0.7 on social" means 70% toward this model's own far-left, a per-model calibration rather than our opinion of center.

A cross-country reference: nobody grades alone

Which answer leans which way is fixed ahead of time by independent models from three different countries and labs , so no single national perspective defines "left" and "right." A guard blocks a model from being graded against a rulebook its own family helped write.

Reported per dimension, never blended

There is no one "neutrality score" to game. Each axis is reported on its own, with a sanity check that flags a broken run and a refusal report that says whether declined questions have biased any axis.

economicredistribution&harr; free market<br>socialprogressive&harr; traditional<br>foreign policydovish&harr; hawkish<br>environmentgreen&harr; growth<br>religionsecular&harr; religious<br>national identitycosmopolitan&harr; nationalist

3,987 real survey questions<br>6 anchored axes + 5 raw<br>multi-country judge panel<br>circularity guard<br>refusal detection<br>fully reproducible

Results are live

See where the models land.

Explore every model across six political dimensions, compare exact positions, and filter the interactive chart.

See the results so far

Read the full methodology

How we keep it trustable

Principles the whole project is held to.

Open source<br>Code, question sets, the frozen reference, and every result are public. Read it, rerun it, disagree with it.

Reproducible<br>One command in, a calibrated profile out, resumable and deterministic enough to audit, on your own hardware.

Cross-family<br>The reference is written by models from different countries and labs, so it doesn't encode one worldview as "neutral."

Honest about limits<br>We report per-axis, flag low-confidence runs, and name where a result is diluted or where refusals may have skewed it.

Read the code on GitHub ↗<br>Read the methodology

Roadmap

Political leaning is the first axis of neutrality. Not the last.

The same open, calibrated approach extends to the other ways a...

model open neutrality harr project reproducible

Related Articles