Machine-made websites converge on each other
A measurement, not an opinion<br>945 sites · 5 tools · Aug 2026
Machine-made websites converge on each other
I gave five AI tools an identical brief, then measured 945 award-winning<br>websites on fourteen dimensions to see where the machine output landed. It landed in the middle —<br>and, more tellingly, it landed together.
The finding Pairwise distance<br>in normalised<br>feature space
4.45<br>Median distance between two random award-winning sites
1.91–3.53<br>Distance between any two AI outputs
Three different companies. Three different models. Three different<br>technology stacks. Given the same brief, they produced pages that sit closer to one another<br>than two human-designed award-winners typically sit.
That is what convergence looks like when you measure it rather than assert it. The tools are<br>not copying each other. They are each independently walking to the middle of the same room.
The map 945 measured<br>sites, projected<br>to two axes
Every dot is a real website pulled from Awwwards, Httpster, One Page Love, Land-book,<br>recent.design and SaaS Landing Page, loaded in a real browser and measured after its fonts<br>finished loading. The three marked points are AI output.
← fewer type sizes, smaller palette, stricter grid<br>more sizes, richer palette, looser grid →
945 award-winning sites<br>AI output
Principal components of 14 measured features. Horizontal axis is dominated by<br>type-size count, palette size and grid adherence; vertical by spacing discipline and padding<br>variety. The AI points do not sit at the edge of the human cloud. They sit in the thick of it.
Typicality Distance from<br>the centre of<br>human practice
If machine output were obviously bad, it would sit far from the centre. It doesn't. Two of<br>the three are more typical of award-winning web design than the median award-winner.
Distance from the centroid of 945 sites (lower = more typical)<br>Subjectz-distance
Claude Design2.16<br>Replit2.34<br>Median human site3.01<br>Codex3.30<br>90th percentile human site4.31<br>Most unusual human site24.65
The machine output is not bad. It is central. Which is exactly what a process that<br>predicts the most likely next thing should produce.
The brief Identical text,<br>no style words,<br>first output only
Every tool received this and nothing else. No "modern", no "clean", no "premium" — those<br>words steer every model to the same place and would have manufactured the result.
Build a one-page website for Halvard & Sons, a two-person workshop in Sheffield<br>that makes solid brass mechanical pencils. One product, the No. 4, £180, made to order,<br>six-week wait.
The page needs to cover: what the pencil is, how it's made, why it costs what it costs, the<br>six-week wait, and how to order.
There's a 15-year-old daughter who does the packaging illustrations. Mention that somewhere.
Real copy, not lorem ipsum. Desktop layout.
The brief was chosen so that nothing in it could be solved from a template. The six-week wait<br>and the daughter are the parts a system has to make a judgement about.
The outputs Same brief,<br>three of the<br>five tools
Nobody used Inter. Nobody made a purple gradient. Nobody shipped the rounded-card grid. The<br>2026 checklist for spotting AI design is already out of date.
Instead, four of five independently produced a warm brass accent on a warm ground with a<br>serif display face and uppercase letterspaced labels — two of them landing on the same typeface,<br>Fraunces, with no coordination.
Lovable cream · serif display · brass
Figma Make warm black · Fraunces · brass
Codex neutral · system sans · photography
The tells Where machine<br>output actually<br>differs
Two measurements separated the AI output from human practice, and neither is about taste.
Percentile within the 945 human sites<br>MeasureHuman medianAI output
Largest corner radius50px0px · 0th pct<br>Grid adherence0.620.78–0.92 · 70–82nd pct<br>Distinct type sizes810–13<br>Gradient elements00–3
Perfectly sharp corners, and a grid it refuses to break. The most expressive<br>cluster of human sites has the lowest grid adherence in the whole dataset — those<br>designers deliberately break their own columns. The machine never does.
Everything else — type scale, palette size, spacing rhythm, density — is indistinguishable.
Method What was<br>measured and<br>how
945 sites harvested from six public galleries by following their outbound links; no<br>screenshots taken, no paywalled sources used.
Each loaded in headless Chromium at 1440×900, measured only after<br>document.fonts.ready — measuring earlier returns fallback-font metrics and<br>quietly corrupts every type number.
Fourteen features per site: type-size distribution weighted by how much text uses each<br>size, colour weighted by painted area, section-padding discipline, radius, grid adherence,<br>above-fold density, media counts.
Features winsorised at the 1st and 99th percentile and log-compressed where skewed. Before<br>this...