AI Operations · established evidence
The AI Productivity Case: What a Peer-Reviewed Study of 5,000 Support Agents Actually Shows
The strongest single piece of evidence on AI productivity in a real workplace is a study of 5,179 customer-support agents whose employer rolled out a generative-AI assistant in stages. Access to the assistant raised the number of issues an agent resolved per hour by about 14 to 15 percent on average. The average hides the finding that matters: the gain was roughly 34 percent for the newest and lowest-skilled agents and close to zero for the most experienced. The tool did not replace anyone. It raised the floor by spreading the working habits of the best agents to everyone else, and it improved customer sentiment and staff retention at the same time. This is a case for treating AI as a productivity multiplier for a team, weighted toward its newer members, rather than as a substitute for it. This piece explains what the study measured, what it did not, and how far its results travel.
The study, and why it carries more weight than most
Most claims about AI and productivity rest on vendor case studies, self-reported surveys, or short demonstrations. Brynjolfsson, Li and Raymond published something different: a field study of the staggered rollout of a generative-AI conversational assistant to 5,179 customer-support agents at a large software firm, released as NBER Working Paper 31161 in 2023 and later peer-reviewed and published in the Quarterly Journal of Economics in 2025.
Two features give it unusual authority. First, the outcome measured was a hard operational number, issues resolved per hour, drawn from the firm's own logs rather than from what workers said about themselves. Second, because the assistant was introduced to different groups at different times rather than all at once, the researchers could compare agents who had the tool against otherwise similar agents who did not yet have it, which is far closer to a controlled comparison than a simple before-and-after. That design is why this study, rather than the louder claims around it, is the reference point for what generative AI does to knowledge work.
The headline number, and the distribution beneath it
On average, access to the assistant raised productivity, measured as issues resolved per hour, by about 14 to 15 percent. That figure is worth stating plainly because it is neither trivial nor miraculous. It is a meaningful operational improvement of the kind that compounds over a year, and it is a long way from the wholesale replacement of labor that both enthusiasts and alarmists tend to describe.
The more instructive result is how unevenly that average was distributed across the workforce.
Novices gained the most
The productivity increase was concentrated among the least experienced and lowest-skilled agents, where it reached roughly 34 percent. For a new hire, the assistant materially closed the gap between a first week on the job and competent handling of a support queue.
Experts gained almost nothing
For the most experienced and highest-skilled agents, the measured effect was close to zero. The people who already knew how to resolve issues efficiently had little to learn from a tool that surfaced approaches they had themselves developed. This asymmetry is not a footnote; it is the central finding, and it points directly at what the tool was actually doing.
The mechanism: generative AI at work spread tacit knowledge
The authors' explanation for the pattern is precise and testable. The assistant, trained on the firm's own record of successful resolutions, effectively captured the tacit best practices of its most able agents and made them available to everyone else in the moment of a customer conversation. It functioned less like a machine doing the work and more like an always-available senior colleague suggesting how a strong agent would handle the case.
That mechanism explains the distribution. If the value of the tool is the diffusion of expert practice, then the people with the most to gain are precisely those furthest from that practice, and the people with the least to gain are the experts who supplied it. It also reframes what such a system is for. Its economic contribution in this study was not to remove humans from the loop but to raise the productivity of the humans who were least productive, narrowing the spread between the best and the rest.
Beyond throughput: sentiment and retention moved too
Productivity was not the only outcome that shifted. The same study reported that the assistant improved customer sentiment, measured in the tone of interactions, and was associated with higher worker retention. Both effects are consistent with the mechanism. Customers reach resolution faster and with fewer poor exchanges when a struggling agent has better guidance, and agents who feel supported rather than exposed in difficult conversations are more likely to stay.
For any operator weighing an AI deployment, these secondary results matter as much as the headline. A change that speeds up work while degrading the customer experience or burning out staff is not a durable gain. Here the evidence pointed the other way, which is part of why the study is cited as a case for augmentation rather than mere speed.
Augmentation vs automation: what the wider evidence says
The support-agent study is one data point in a larger and more cautionary picture, and the accurate reading places it inside that context rather than treating it as a promise.
Anthropic's Economic Index, built from usage telemetry across its business customers, found that when firms embed the model through an API, roughly 77 percent of transcripts show full-task-delegation patterns and only about 12 percent show iterative, collaborative augmentation, the reverse of the ratio seen in consumer chat use. In other words, businesses tend to reach for automation by default even though the best-documented workplace gains in the support study came from augmentation. That gap between how AI is deployed and where its measured value has actually appeared is the practical risk.
Two further lines of evidence explain why default automation of a customer-facing role is fragile. In the peer-reviewed tau-bench benchmark, frontier LLM agents succeeded on fewer than half of realistic customer-service tasks, and consistency was worse still: asked to repeat the identical task eight times, the same agent succeeded only about a quarter of the time. And Klarna, after publicly replacing large parts of its support workforce with an AI assistant in 2024, resumed hiring human agents by 2025 and moved to a hybrid model, with its leadership stating that customers should retain the option to reach a person. The throughput gain in the support study is real; the reliability required to remove the human entirely is not yet demonstrated in the same rigorous way.
The limits: one large firm is not your front desk
The result should be held with clear caveats.
The study observed a single large software firm's contact center with thousands of agents handling a high, repetitive volume of text-based issues. That is close to an ideal setting for the mechanism to work: enough historical data to learn from, standardized tasks, and a large enough novice population for the floor-raising effect to show up in the aggregate. A five-person med-spa front desk or a solo-legal intake process shares some of these features and lacks others, and no study surfaced here replicates the finding at small-business scale.
Extending an enterprise contact-center result to a small team is therefore a reasonable inference, not a proven transfer, and it should be labeled as such. The 14 to 15 percent figure is a finding about one workforce, not a number any business should expect to reproduce. What travels more safely than the specific percentage is the shape of the result: that generative AI in this setting helped least-experienced workers most, worked by diffusing expert practice, and delivered its value alongside humans rather than instead of them. That shape is a hypothesis a small business can test on its own workflows, with its own baseline, rather than a promise to be taken on faith.
What a small business should take from this
The operational reading of the evidence is narrower and more useful than the headline. It suggests aiming an AI deployment at the parts of the work where the least experienced people lose the most time, drafting and researching and structuring a response, while keeping a human on the decisions that carry risk. It suggests measuring a genuine before-and-after on your own numbers rather than importing someone else's percentage. And it suggests treating full, unattended automation of a customer-facing role as an unproven step to be earned through evaluation, not assumed at the outset, given the benchmark and real-world reliability evidence above.
None of this argues against using AI. It argues for a specific way of using it: as a multiplier on a team, weighted toward its newer members, with the human kept in the loop where it matters and the effect measured rather than asserted. That is a defensible position drawn from the strongest evidence available, and it happens to be the opposite of the reflex the wider deployment data shows most businesses following.
The evidence
Key findings, with their sources
-
Access to a generative-AI assistant raised customer-support productivity, measured as issues resolved per hour, by about 14 to 15 percent on average across 5,179 agents.
established Brynjolfsson, Li & Raymond, "Generative AI at Work", NBER Working Paper 31161 (2023); Quarterly Journal of Economics 140(2), 889-967 (2025).
-
The gain was roughly 34 percent for the least experienced and lowest-skilled agents and close to zero for the most experienced.
established Brynjolfsson, Li & Raymond, "Generative AI at Work", NBER 31161 / QJE 140(2), 2025.
-
The assistant also improved customer sentiment and was associated with higher worker retention; its mechanism was diffusing the tacit best practices of the most able agents to everyone else.
established Brynjolfsson, Li & Raymond, "Generative AI at Work", NBER 31161 / QJE 140(2), 2025.
-
When businesses embed AI through an API, roughly 77 percent of transcripts show full-task-delegation automation versus about 12 percent collaborative augmentation, the reverse of consumer chat use.
established Anthropic, "The Anthropic Economic Index", usage telemetry across business customers, 2025 to 2026 editions.
-
Frontier LLM agents succeeded on fewer than half of realistic customer-service tasks, and on repeating the identical task eight times succeeded only about a quarter of the time.
established Yao et al., "tau-bench: A Benchmark for Tool-Agent-User Interaction in Real-World Domains", arXiv:2406.12045, 2024 (Sierra Research).
-
After replacing support staff with an AI assistant in 2024, Klarna resumed hiring human agents in 2025 and moved to a hybrid human-plus-AI model.
established Klarna corporate disclosures; Fast Company and CX Dive reporting, 2025.
-
No published study replicates the 14 to 15 percent support-agent result at small-business scale; applying an enterprise contact-center finding to a small team is an inference, not a proven transfer.
contested Assessment of the current literature surfaced for this piece; the Brynjolfsson, Li & Raymond study observed a single large software firm.
Calibration
What is proven, what is promising, what is unproven
| Evidence tier | Tactics | What the evidence says |
|---|---|---|
| established | Treat the 14 to 15 percent average, the roughly 34 percent novice gain, the near-zero expert effect, and the sentiment and retention improvements as findings of a peer-reviewed field study. | Brynjolfsson, Li & Raymond, NBER 31161 / QJE 140(2), 2025. |
| established | Treat the automation-over-augmentation deployment ratio, the sub-50 percent agent task success and poor repeat consistency, and the Klarna reversal as real context on why full automation of a support role is not yet reliable. | Anthropic Economic Index; tau-bench (arXiv:2406.12045); Klarna disclosures and reporting. |
| contested | Treat any specific percentage carried over to a small business front desk as an untested inference; measure your own before-and-after instead of importing the number. | No small-business replication of the support-agent study was found in this research pass. |
Reference
Glossary
- Augmentation
- A pattern in which AI assists a human doing the work, suggesting or drafting while the person decides and acts, rather than performing the task unattended.
- Automation (full-task delegation)
- A pattern in which AI is handed a whole task to complete on its own with little or no human involvement in each step.
- Staggered rollout
- Introducing a tool to different groups at different times, which lets researchers compare users against not-yet-users and approximate a controlled experiment.
- Tacit knowledge
- The hard-to-write-down working know-how that skilled people accumulate through experience. The support study's mechanism was making this expert know-how available to less experienced staff.
- Human in the loop
- A design in which a person reviews or approves an AI system's output before it takes effect, especially for work that is client-facing, financial, or otherwise carries real risk.
Straight answers
Frequently asked questions
What did the study of 5,000 support agents actually find?
It found that giving customer-support agents access to a generative-AI assistant raised issues resolved per hour by about 14 to 15 percent on average, with the largest gains, around 34 percent, among the least experienced agents and almost no effect on the most experienced. It also reported better customer sentiment and higher retention. The work is Brynjolfsson, Li and Raymond, published as NBER 31161 and in the Quarterly Journal of Economics in 2025.
Does this prove AI will replace customer-service jobs?
No. The study measured AI working alongside agents and raising their output, not replacing them, and its value came from helping less experienced staff. Separate evidence points the other way on full automation: benchmarked AI agents succeed on fewer than half of realistic support tasks, and Klarna reversed a heavy automation push and returned to a hybrid model. The reading is a case for augmentation, not replacement.
Why did new employees benefit so much more than experienced ones?
Because the tool worked by capturing the tacit best practices of the firm's best agents and making them available in the moment. People far from that expertise, the newest hires, had the most to gain, while the experts who already used those practices had little to learn from it. The tool raised the floor rather than the ceiling.
Will my small business see the same 14 to 15 percent gain?
You should not assume so. The result comes from one large contact center with thousands of agents and a high volume of repetitive text-based issues, and no study has replicated it at small-business scale. The number is a finding about that workforce, not a promise for yours. What travels is the shape of the result: aim AI at where inexperienced staff lose time, keep a human on risky decisions, and measure your own before-and-after.
What is the difference between augmentation and automation here?
Augmentation means AI assists a person who stays in control of the work; automation means AI is handed a whole task to run on its own. The support study documented gains from augmentation. Usage data shows most businesses reach for automation by default, which is where the reliability evidence says the risk sits, so the sequence and the human checkpoints matter.
Provenance
Sources
- Brynjolfsson, E., Li, D. & Raymond, L., "Generative AI at Work", NBER Working Paper 31161 (2023); Quarterly Journal of Economics 140(2), 889-967 (2025) (established, peer-reviewed)nber.org
- Anthropic, "The Anthropic Economic Index", business-customer usage telemetry, 2025 to 2026 editions (established)
- Yao, S. et al., "tau-bench: A Benchmark for Tool-Agent-User Interaction in Real-World Domains", arXiv:2406.12045, 2024, Sierra Research (established)arxiv.org
- Klarna corporate disclosures on its AI customer-service reversal; Fast Company and CX Dive reporting, 2025 (established)
- MIT NANDA, "The GenAI Divide: State of AI in Business 2025", 2025 (established, directional; survey and case-study based)nanda.media.mit.edu
- Stanford HAI, "The 2026 AI Index Report", hai.stanford.edu/ai-index/2026-ai-index-report (established)hai.stanford.edu
Every figure above is attributed to a real, dated source and tagged with its evidence tier. Where a claim could not be verified to a primary source, it is not stated as fact.