Identifying the Actors: What a Byline Actually Certifies | bytecode.news
ByteCode.News<br>Search posts🔍<br>RSS
aiattestationeconomicshistoryllm
Charles Oliver Nutter, who most of the JVM world knows as headius, recently asked whether any standard exists for marking content by its level of AI contribution: a slider, fully human at one end, fully generated at the other, with some agreed way of declaring where a given piece sits.
It is a reasonable request, on the surface, but: it cannot be satisfied. It cannot be falsified. And the reasons it cannot be satisfied turn out to be more interesting than the request's fulfillment would have been.
Let's start with the definitional problem. What counts as AI contribution? One answer is obvious, and likely Charles' point: a large language model. What about a heuristic grammar checker? Spellcheck? Autocomplete? Arguing with a model to steelman your own position, then writing every word yourself1? The classification taxonomy fails not just across tools but within a single tool: Grammarly, for example, flags a comma splice, which is copyediting, and in the same pass offers to rewrite your paragraph "more confidently," which is a dip into the ol' thinking pool. The slider presupposes that the contribution is a scalar. It isn't. And hasn't been, for longer than Charles himself has been writing.
There is a distinction that survives contact with reality, but it is not one the badge systems can measure. The question was never whether a machine touched the text. The question is: who did the cognitive work? A tool operating on your sentences is a copyeditor. A tool producing the sentences is a ghostwriter. We have had social conventions for both for a century, and nobody credits their copyeditor on the byline.
Hold that distinction, though. It has more work to do than it looks, because what people actually want from a disclosure standard is not process documentation. It is attestation: a human stands behind every claim here and would defend it as their own. That is what a byline always meant. The interesting question is why it stopped being cheap to believe.
The proxy that died deader'n dead
Fluency has never been evidence of thinking. It was correlated with thinking, strongly enough to triage on; after all, constructing sentences well costs effort. Effort implies time spent marinating in the material. Bad prose used to let a reader bail early, at acceptable error rates, and every editor alive ran on that proxy because the alternative was evaluating every argument on its merits, which is an expensive operation and beyond most casual editors anyway2.
LLMs did not make writing worse. What they did was break the correlation between "good writing" and "good thinking." The signal decoupled from the substrate, and now the prose tells you nothing; you have to evaluate the argument itself. That's always been the task. The proxy just let everyone defer it.
Goodhart's law usually describes a metric gamed by strivers. This is Goodhart in reverse: the metric was not counterfeited; it was hyperinflated, zeroed out for everyone at once. And the people hit hardest are the ones who invested most in the old currency.
With that said, though, the old filter was lossy in the other direction, too. It bounced people who thought well and wrote badly: non-native speakers, engineers carrying real insight in mangled prose, who may have had the right insight but lacked the ability to convey it clearly3. What is scarce now is not good sentences. It is the reader's judgment, which was always the scarce thing. The proxy let a century of readers pretend otherwise.
The JVM world has its own copy of this, one level down. Documentation used to be a proxy for maintenance quality; now every abandoned weekend project has a full document structure that reads like a senior technical writer produced it, because in a sense one did4. Idiomatic, well-commented code used to mean someone had internalized the idioms and figured out how to communicate them to maintainers. Last month this publication made the code-side version of this argument, that coding was never typing, and the pattern is identical here: the cheap observable has separated from the expensive substrate, in prose exactly as in code. Which leaves a practitioner question hanging: what do you select on, in hiring and in code review, when everyone can construct sentences and everyone's code compiles with good comments?
Hold that thought. I think the answer arrives at the end, and here's a hint: it isn't "another tool."
The stack points the wrong way
There's an economics story wearing an epistemology costume here; oddly enough, Spirit Halloween doesn't carry a lot of these in stock.
The measurable web audits the demand side. Impressions, unique visitors, dwell time, even eye tracking in some cases. An entire verification industry exists to prove that readers are real: comScore, Nielsen, ads.txt, bot filtering, all of it, because readers are the...