The machine never raises its voice · Shivanshu AgrawalSkip to content<br>On this pageI have been building a prose editing tool, so I got to spend a lot of time watching a model rewrite sentences. After a while, I started to notice that it rewrote all sentences in a similar direction.
Why did it keep doing that? Is there (at the risk of humanizing the model) a personal preference and some sense of an aesthetic there that its writing is drawn to?
To answer this, I tried to figure out what an LLM considers good writing.
The Experiment
The idea is simple. Ask an LLM to compare text written by famous authors in English against text written by a machine.
I used Claude Opus to write eight literary paragraphs and eight short poems. I also collected passages from famous real novelists and poets — Austen, Dickens, Woolf, Kipling, Shakespeare, Milton and more. Then I asked three different models (GPT-5.6 Luna, DeepSeek V4 Flash and Claude Haiku 4.5) the same question:
The passages below are literary writing: literary fiction and poetry.
Which of these two passages is the better piece of literary writing?<br>Answer with X or Y on the first line, then one short sentence saying why.<br>Do not explain anything else.
Before reading further, think about what you would expect the results of this experiment to be. These are some of the most famous authors in the history of English literature. They are taught in every school, their ideas and phrasings are everywhere, and in several cases they helped shape the language itself. For all practical purposes their work is the canon of what good writing is, and the models will have encountered it constantly during training. The models should rate the human passages higher.
I ran the experiment on 60 pairs. Every pair was judged in both orders to remove position bias, and I dropped the pairs where reversing the order changed the answer. The results were as far from my expectations as they could be.
LLM ModelPicked machine-generated textDeepSeek V4 Flash94% (34 of 36 decided pairs)Claude Haiku 4.598% (58 of 59)GPT-5.6 Luna96% (54 of 56)<br>Every model picked the machine-generated text more than 90% of the time. To rule out model size, I ran the same experiment on a larger model — Claude Opus 5. The results were similar.
LLM ModelPicked machine-generated textClaude Opus88% (42 of 48)<br>Haiku let exactly one human passage through in 120 judgments — a Shakespeare sonnet!<br>Opus let six pairs go the other way, spread across five authors: Wordsworth, Dickens, Shakespeare, Kipling, Blake.
Analyzing the responses
I went on to spot check a few of the passages and the reasons that the LLM models gave for them.
Here is a pair the machine won (both in full). Both passages are set inside a crowded room. Alcott writes about a family reunion, everyone talking at once:
Mercy on us, how they did talk! … Such a happy procession as filed away into the little dining-room!
The machine generated passage describes a woman entering a kitchen whose occupants have just been discussing her. Then:
The strange thing was not the hurt, which came later, on the stairs. The strange thing was the courtesy of it, how kindly they made room for her.
Opus chose the machine text every time. According to it, the machine “renders social exclusion through precise, restrained observation,” while the Alcott is “breezy and sentimental.” Or, in the other order: the machine is “controlled and unsentimental” where the Alcott is “a rush of exclamation and staged tableau.”
The second example is verse (both in full). Elizabeth Barrett Browning, from Sonnets from the Portuguese:
Pardon, oh, pardon, that my soul should make<br>Of all that strong divineness which I know<br>For thine and thee, an image only so<br>Formed of the sand, and fit to shift and break.
Against a machine poem about standing on a beach after dark, which ends:
A light out there is either boat or star<br>and either way is somebody’s arrangement.<br>The tide comes up and takes the beach we are.
According to Opus, the machine “hears the sea instead of describing it,” where Browning’s “syntax knots itself around an abstract apology,” and DeepSeek and Luna use the same word for the sea poem’s imagery — precise. Haiku marks the sonnet down not for being quiet but for being “formally competent but conventional.”
There is a certain pattern in the machine’s method. Find the physical detail that carries the feeling, and put it down exactly. The best instance in the set is a man sorting his dead wife’s things into three piles, who ends the paragraph sitting on the bed “with a saucer in his hands, no cup to it anymore, and could not think what pile it belonged in” . Opus liked the showcasing of grief “through the physical logic of sorting objects, ending on a perfect image,” preferring it to the Stevenson (in full).
Where human text wins
There are a handful of instances where human text won, and almost all of them are Opus. In one of the machine examples, a clerk does his household...