Human vs. AI vs. Human and AI: Who Does Better Work?

rafaelaziz1 pts0 comments

Human vs. AI vs. Human + AI: Who Actually Does Better Work? | Rafael Research

Book a call

← Research<br>Applied AI Research

Human vs. AI vs. Human + AI: Who Actually Does Better Work?

What eight controlled studies — found through a structured search,<br>not a shortlist — actually show about human, AI, and human+AI<br>performance, once you check the numbers instead of the narrative.

Rafael Research · August 2026

There's a comfortable assumption running through most corporate AI<br>strategy right now: a skilled person plus an AI tool will outperform<br>either one alone. It's intuitive, it's reassuring, and it lets<br>organizations roll out AI everywhere without asking hard questions<br>about where, specifically, it helps.

It's also not what the evidence shows.

We identified eight controlled empirical studies — through a<br>structured literature search described below, not an arbitrary<br>shortlist — that compare at least two of unassisted human<br>performance, AI-alone performance, and human-plus-AI performance<br>on a well-defined task, using real accuracy, speed, or quality<br>measurements rather than survey sentiment. Only a subset of the<br>eight contain all three conditions in a single design; the<br>remainder provide controlled two-arm evidence (most often<br>human-alone vs. human+AI, or human-alone vs. AI-alone) that helps<br>test whether the broader pattern generalizes across tasks — which<br>study includes which conditions is noted in each section below<br>and in the summary table. They span clinical diagnosis,<br>management consulting, customer support, professional writing,<br>and software engineering. These are not eight versions of the<br>same experiment. They differ in design, sample, and what they<br>measure: most are randomized controlled trials, one is a<br>staggered field rollout with a randomized pilot; some measure a<br>single 20-minute task, one follows real production work across<br>three companies over months. Read individually, each study<br>answers a narrow question about its own task and population. Read<br>together, a pattern emerges that is not "AI helps." It's closer<br>to a fault line: AI's advantage is real, large, and reproducible<br>on some tasks, and reverses into a disadvantage — or simply<br>disappears — on others that look, to a human, more or less the<br>same. A large independent meta-analysis of the broader literature<br>finds a closely aligned pattern, which is the strongest evidence<br>this isn't an artifact of which eight studies we happened to<br>pick.

Methodology: how these eight studies were selected

To avoid presenting a hand-picked set as more authoritative than it<br>is, we ran a structured — not a formal systematic-review-protocol —<br>literature search across Google Scholar, Semantic Scholar, PubMed,<br>SSRN, NBER, arXiv, the ACM Digital Library, and IEEE Xplore for<br>controlled studies published in 2022 or later that met all of the<br>following: (1) compare at least two of unassisted human<br>performance, AI-alone performance, and human+AI performance; (2)<br>use a real or realistic work task, not a survey or self-reported<br>opinion; (3) report a quantitative outcome — accuracy, speed, a<br>graded quality score, or an error/completion rate, not<br>satisfaction; (4) are peer-reviewed and published, or are a<br>working paper from an established research institution or lab,<br>not a vendor blog post or marketing study; and (5) cover clinical<br>diagnosis, knowledge work, customer service, writing, or software<br>engineering, with a broader scan across education, forecasting,<br>hiring, translation, legal work, and classic human-AI "centaur"<br>research to check for major work we might otherwise miss. This is<br>a transparent, criteria-driven search, not a formal systematic<br>review — it does not follow a pre-registered PRISMA-style<br>protocol, log exact search strings or per-database result counts,<br>or use a second independent screener, and readers who need that<br>standard of evidence should treat it accordingly.

That search surfaced roughly two dozen candidates. Most were<br>excluded for a specific, checkable reason: some measured<br>self-reported time allocation or satisfaction rather than task<br>performance (a large Microsoft 365 Copilot field study of over<br>7,000 workers, NBER Working Paper 33795, measured how workers<br>reallocated their time, not whether their output improved); some<br>were observational rather than randomized, which weakens causal<br>claims (a study of 72,000+ GitHub pull requests found AI-assisted<br>PRs merged faster, but without random assignment); and a few strong<br>two-arm studies (AI-alone vs. human-alone, no combined condition)<br>in medicine and forecasting were kept as corroborating context<br>rather than counted as primary sources, to avoid overweighting any<br>single domain. The eight studies below are the ones that survived<br>that screen with the strongest designs available — eight primary<br>studies in total, including one earlier preprint (Peng et al.,<br>2023) that we retained specifically as a labeled historical<br>predecessor to a later, stronger study on the same question,<br>rather than as independent...

human work eight studies performance alone

Related Articles