The null result in OpenAI's enterprise AI paper — The Working Model
Hype-o-meter
The null result in OpenAI's enterprise AI paper
14 August 2026<br>6 min read
OpenAI published a working paper this week with Columbia and Wharton, built on ChatGPT Enterprise telemetry through March 2026. The findings everyone will quote are in the abstract. The one worth your time is a coefficient in Table 2 that the paper reports and then moves past.
It's a good paper. It's also a vendor publishing data about its own customers, with two co-authors on its payroll, so read it the way you'd read anything else in that category.
The headline findings are what you'd expect. Usage grew sevenfold. Adopters skew large. Task use is broad rather than concentrated. All plausible, none of it especially surprising.
What I keep coming back to is elsewhere.
Among firms that have already adopted, larger headcount predicts lower usage per employee. Fewer messages per head, fewer weekly active users per head, fewer tokens per head, all statistically significant. This is presented, reasonably, as a scaling effect. Big companies dilute.
But there are four columns in that table, and the fourth is messages per weekly active user.
Table 2 — effect of lagged log employment on usage, among adopters
Messages per active week, per employee<br>−0.266 (0.019) ∗∗∗
Weekly active users per employee<br>−0.032 (0.003) ∗∗∗
Output tokens per employee<br>−0.667 (0.048) ∗∗∗
Messages per weekly active user<br>−0.002 (0.008) n.s.
Standard errors clustered by firm in parentheses; ∗∗∗ denotes significance at 1%. Three significant negatives and one null. The null is the one that tells you what to do.
Nothing. Flat. A standard error four times the size of the estimate. A 200,000-person company's engaged AI users are as engaged as a 500-person company's engaged AI users.
One honest caveat before I lean on this. That fourth regression has an R² of 0.100, against 0.36 to 0.48 for the other three. It's a noisier specification, so the null is weaker evidence than a null in a tightly fitted model would be. But the point estimate is essentially zero, not merely imprecise, and the sign flips nothing.
Which means the large-enterprise usage gap has nothing to do with how people use the tool once they're using it. It's entirely a question of how many people are using it at all.
If your problem is engagement, you buy training. If your problem is penetration, none of that touches the constraint.
I don't think this is a small distinction, because the two readings send you to completely different places. If your problem is engagement, you're buying better prompts, better training, better internal evangelism, better use case libraries. That's the entire enablement industry right now. If your problem is penetration, you should be looking at who never logged in a second time and why.
It's easy to measure the first thing and assume it explains the second. This paper is decent evidence that it doesn't.
What the data can't tell you is why the non-returners don't return, and that's the question the null result makes urgent. The dataset has no denominator for who was offered a seat and declined it, no record of the second session that never happened. Whatever is happening there is happening before any of the enablement machinery gets a chance to work, which is a fairly awkward finding for an industry that has organised itself almost entirely around the post-onboarding phase.
The second thing worth sitting with is what predicts adoption in the first place.
The researchers tested three kinds of accumulated intangible capital, all measured in fiscal 2021, well before any of these firms bought anything: R&D stock per employee, capitalised software per employee, and SG&A stock per employee. SG&A is the crude proxy here for organisational and managerial capability. It came out strongest and most robust. Physical capital intensity ran the other direction, negatively associated with adoption once you control for scale.
So the firms that moved first weren't the ones with the best engineering or the deepest software estate. They were the ones that had already spent years building the overhead function that redesigns how work happens. The unglamorous layer. The layer that gets cut first in a downturn and that no one puts in an earnings call.
I find this uncomfortable and probably correct. It's consistent with everything we know about general purpose technologies, and it's consistent with the pattern where the technically strongest organisation in a given industry is frequently not the one that adapts fastest. Capability to absorb is a separate asset from capability to build, and we're not very good at measuring or funding it.
It also suggests a fairly grim near-term dynamic. If the firms best positioned to absorb this are also the largest and most valuable ones, and the models...