AI Operations · established evidence
How Businesses Use AI: Anthropic's Own Usage Data Shows They Delegate, They Don't Chat
How businesses use AI turns out to be different from how people use it, and the gap is measurable. Anthropic publishes an Economic Index built from usage telemetry across more than 300,000 business customers, and when companies embed its models through an API rather than typing in a chat window, about 77 percent of the transcripts show full-task delegation, handing over a bounded job and taking back a result, versus only about 12 percent that show iterative collaboration. That is close to the opposite of the consumer chat pattern, where people mostly work alongside the model. The practical lesson is not that AI is about to replace staff. It is that businesses are already treating AI as something you assign a defined task to, not a coworker you brainstorm with, and that reframes what an "AI employee" can realistically be asked to do.
The finding: delegation, not conversation
Most public claims about "how businesses use AI" rest on surveys, where owners report what they think they do. Anthropic's Economic Index is different in kind: it is derived from actual usage telemetry across more than 300,000 business customers, so it records behavior rather than intention. That distinction matters, because self-report and observed use often diverge.
The headline pattern is an inversion. When a business connects a model through an API, roughly 77 percent of the transcripts fall into what Anthropic labels an "automation" or full-task-delegation pattern: a defined task goes in, a completed result comes back, with little or no turn-by-turn steering. Only about 12 percent show the "augmentation" pattern of iterative, collaborative back-and-forth. In consumer chat use the ratio runs the other way, dominated by people working alongside the model. The same underlying system is being used in two structurally different ways depending on who is holding it.
Read carefully, the number describes a behavior, not a capability ceiling. It tells us how businesses choose to deploy AI when they wire it into their own systems, which is by delegating discrete tasks. It does not claim those tasks are performed flawlessly, and it does not say delegation is the correct default for every job. Those are separate questions the rest of this piece takes up next.
How businesses use AI is not how consumers use it
The reason the two patterns diverge is architectural. A person in a chat window is present for every turn; they can correct, reframe, and nudge in real time, which naturally produces collaboration. An API call embedded in a booking flow or a back-office pipeline runs unattended: there is no human sitting inside the loop to co-write, so the sensible design is to hand the model a scoped task and consume its output. The interface shapes the behavior.
The composition of that delegated work is also shifting, not static. Anthropic's data show the share of API usage tied to office and administrative-support tasks rising from a low base of about 3 percentage points to roughly 13 percent between August and November 2025, and sales-workflow API use roughly doubling between November 2025 and February 2026. The direction of travel is toward exactly the operational, repetitive back-office work that small and mid-size businesses run all day.
For an owner, the useful translation is this. The question is not "should I chat with an AI about my business," which is what most first exposure looks like. The question is "which bounded, repeatable tasks in my operation can be handed off and collected back reliably," which is what the businesses generating this telemetry are actually doing.
What "delegation" actually means: a bounded task, not a coworker
The word "delegation" invites a misleading mental image, the AI as a new hire you onboard and then trust broadly. The usage data supports a narrower reading. What businesses delegate is a task with a defined input and a defined output, not open-ended responsibility. That is a meaningful constraint, and it maps onto a distinction Anthropic itself draws in its engineering guidance.
In "Building Effective Agents", Anthropic separates workflows, where models and tools are orchestrated through predefined code paths, from agents, where the model dynamically directs its own process and tool use. Its explicit recommendation is to reach for the simplest pattern that works and to add open-ended autonomy only when flexibility is genuinely required, because autonomy trades latency, cost, and predictability for capability. Full-task delegation, in that framing, is usually a workflow: a scoped job on rails, not a free-roaming agent.
Why the "AI employee" metaphor breaks
A human employee generalizes. You can hand them an unfamiliar problem and they will improvise, ask, and escalate. The delegation pattern in the usage data is the opposite: it works precisely because the task is bounded and repeatable. Expecting a delegated system to behave like a general-purpose colleague is where most disappointment starts, and it is a category error, not a model shortcoming.
The realistic expectation, then, is a fleet of narrow, well-defined delegations, appointment reminders, intake capture, record sync, first-draft replies, each doing one job dependably, rather than a single synthetic staff member who does everything. That is what the businesses in Anthropic's dataset are building.
Delegation raises the reliability bar, it does not lower it
Handing a task off entirely sounds like less work, and in day-to-day terms it is. But it moves the burden of correctness to the front, because no human is watching each turn to catch a mistake. That is exactly where the strongest countervailing evidence sits.
The tau-bench benchmark from Sierra Research tested tool-using agents against simulated users and real tool APIs under realistic customer-service policies. Even frontier, GPT-4-class agents completed fewer than half of the realistic tasks, and consistency was worse than the raw success rate implies: run the identical task eight times and the same agent succeeded only about a quarter of the time. Unreliability under repetition, not occasional error, is the finding that matters most for anything you intend to delegate and walk away from.
This does not contradict the usage data; it qualifies it. Businesses are delegating at scale, and delegation is the right pattern for many tasks, but the reliability numbers say the sensible unit of delegation is a bounded task with a human approval gate wherever the action is irreversible or costly. The lesson from the field is to delegate the task and keep a person on the trigger for anything that sends a message to a customer or spends money, which is precisely how well-built automation is scoped.
Augmentation still has the best-documented payoff
The 12 percent augmentation pattern is easy to dismiss as the minor case. It is worth resisting that, because the single best-documented workplace productivity result to date is about augmentation, not full delegation.
Brynjolfsson, Li and Raymond studied the staggered rollout of a generative-AI conversational assistant to more than 5,000 customer-support agents at a software firm, published in the Quarterly Journal of Economics in 2025. Access to the assistant raised productivity, measured as issues resolved per hour, by roughly 14 to 15 percent on average, with the largest gains, about 34 percent, concentrated among novice and lower-skilled workers, and near-zero effect on the most experienced. The mechanism was diffusion: the tool spread the tacit know-how of the best workers to everyone else. That is augmentation working, a human and a model side by side.
The practical reading is that automation and augmentation are two different tools for two different jobs, and confusing them is a common and expensive mistake. Delegate the bounded, repeatable task. Augment the human in the judgment-heavy, variable one. One caveat: the productivity study was run in an enterprise contact center, and no small-business replication exists in the literature reviewed here, so treating a five-person front desk as guaranteed to see the same gain is a reasonable inference, not a proven transfer.
The adoption gap: delegating is common, running it in production is not
One more reality check keeps the delegation story in proportion. The usage data show that businesses embedding AI delegate heavily. Separate evidence shows how few businesses have anything running dependably in production, which is a different measurement.
MIT NANDA's 2025 study of enterprise deployments found that about 95 percent of generative-AI pilots produced no measurable profit-and-loss impact, and attributed the gap to organizational adoption failure rather than model quality, with purchased or partnered tools succeeding far more often than internal builds. Stanford HAI's 2026 AI Index found 88 percent of organizations use AI in at least one function but fewer than 10 percent have fully scaled it in any single one, and that 74 percent now name inaccuracy as their top AI risk. McKinsey's 2025 global survey put agentic scaling at roughly 10 percent per function even among organizations that describe themselves as scaling.
Held together, the picture is coherent rather than contradictory. Delegation is the dominant intent, and the composition is moving toward back-office operations, but the distance between "we delegated a task to AI" and "it reliably runs in production" is where most value is currently lost. That gap is an operational and governance problem, not a reason to avoid delegation.
You still own what you delegate
Delegation transfers the work. It does not transfer the responsibility, and the record on this is now settled enough to state plainly.
In Moffatt v. Air Canada (2024), a Canadian tribunal held the airline liable for incorrect information its website chatbot gave a customer, rejecting the argument that the chatbot was a separate legal actor and ruling that a business owns what is on its site whether the words come from a static page or a bot. In the United States, the Federal Trade Commission's Operation AI Comply and the finalized DoNotPay order penalized unsubstantiated claims about what an AI service could do. The through-line is direct: a delegated task performed by AI is still, legally and reputationally, the business's own act.
That is not a reason to avoid delegation. It is the reason to gate it. The businesses generating the usage telemetry are delegating tasks; the durable ones are delegating tasks whose outputs are reviewed, bounded, and reversible where it counts.
How to read this for your own operation
The evidence points to a specific, testable posture rather than either hype or fear. First, expect to delegate bounded tasks, not to hire a synthetic generalist; the usage data show that is what businesses actually do. Second, put the effort at the front, because delegation moves the reliability burden to design time and the benchmark evidence says consistency, not one-off success, is the hard part. Third, keep augmentation in the toolkit for the judgment-heavy work, since that is where the strongest productivity evidence sits. Fourth, gate anything irreversible, because you own the output regardless of who produced it.
None of that can be settled in the abstract. Which tasks in a given operation are bounded and repeatable enough to delegate, and where the human gates belong, is an empirical question about that specific business. The frontier capability keeps compounding, METR measures the length of tasks agents can complete autonomously roughly doubling every seven months, but the readiness of any single operation to delegate safely is the variable an owner can actually control, and it is worth measuring before buying anything.
The evidence
Key findings, with their sources
-
When businesses embed Claude via API, about 77% of transcripts show full-task-delegation ("automation") patterns versus only about 12% collaborative "augmentation", close to the opposite of consumer chat use.
established Anthropic, "The Anthropic Economic Index" (usage telemetry across 300,000+ business customers), 2026 report editions.
-
Office and administrative-support task share of API usage rose from about 3 percentage points to roughly 13% between August and November 2025, and sales-workflow API use roughly doubled between November 2025 and February 2026.
established Anthropic, "The Anthropic Economic Index," 2026 report editions.
-
Frontier, GPT-4-class agents completed fewer than 50% of realistic customer-service tasks, and repeating an identical task eight times, the same agent succeeded only about 25% of the time.
established Yao et al., "tau-bench: A Benchmark for Tool-Agent-User Interaction in Real-World Domains", arXiv:2406.12045, 2024 (Sierra Research).
-
A generative-AI conversational assistant raised support-agent productivity by about 14 to 15% on average, with a 34% gain concentrated among novice workers and near-zero effect on the most experienced, an augmentation result, not full automation.
established Brynjolfsson, Li & Raymond, "Generative AI at Work", Quarterly Journal of Economics 140(2), 2025 (NBER WP 31161, 2023).
-
About 95% of enterprise generative-AI pilots produced no measurable P&L impact, attributed to organizational adoption failure rather than model quality.
established MIT NANDA, "The GenAI Divide: State of AI in Business 2025", 2025.
-
88% of organizations use AI in at least one business function but fewer than 10% have fully scaled it in any single one; 74% now cite inaccuracy as their top AI risk.
established Stanford HAI, "The 2026 AI Index Report".
-
The length of tasks agents complete autonomously at 50% reliability has been doubling roughly every 7 months; the longer-horizon extrapolation is a forecast, not observed fact.
contested METR, "Measuring AI Ability to Complete Long Tasks", 2025-03-19.
Calibration
What is proven, what is promising, what is unproven
| Evidence tier | Tactics | What the evidence says |
|---|---|---|
| established | The 77% delegation versus 12% augmentation inversion in business API use, and the shift toward office and sales tasks | Anthropic Economic Index, direct usage telemetry across 300,000+ business customers, 2026 editions |
| established | Delegated agentic tasks are unreliable under repetition, so bounded scope and human gates are load-bearing | tau-bench (arXiv:2406.12045, 2024); MIT NANDA State of AI in Business 2025 |
| established | Augmentation, not full delegation, carries the best-documented productivity gain, largest for novices | Brynjolfsson, Li & Raymond, Quarterly Journal of Economics 140(2), 2025 |
| contested | The rate at which agents will handle ever-longer tasks autonomously | METR time-horizon measurement is established; the multi-year extrapolation is an explicit forecast |
Reference
Glossary
- Automation (delegation) pattern
- Usage where a defined task is handed to a model and a completed result is taken back, with little turn-by-turn steering. The dominant pattern in business API use in Anthropic's data.
- Augmentation pattern
- Usage where a human and a model work together iteratively, correcting and refining across turns. The dominant pattern in consumer chat, and the mode behind the best-documented productivity study.
- Workflow vs agent
- A workflow orchestrates models and tools through predefined code paths; an agent lets the model direct its own process and tool use. Most reliable delegation is built as a workflow, not an open-ended agent.
- Human-in-the-loop gate
- A required human approval step before an AI-produced action takes effect, placed wherever the action is irreversible or costly, such as sending a customer message or spending money.
- Economic Index
- Anthropic's published analysis of anonymized usage telemetry across its customer base, used to observe how AI is actually used rather than how people say they use it.
Straight answers
Frequently asked questions
What does Anthropic's Economic Index actually measure?
It analyzes anonymized usage telemetry across more than 300,000 business customers to see how AI is used in practice, not how people report using it. Its finding on business API use is that about 77% of transcripts are full-task delegation and about 12% are collaborative augmentation, close to the opposite of the consumer chat pattern.
Does "77% delegation" mean AI is replacing employees?
No. The figure describes how businesses choose to deploy AI when they wire it into their systems, by handing off bounded, repeatable tasks. It is a statement about task design, not headcount. The realistic picture is a fleet of narrow delegations, each doing one defined job, rather than a synthetic generalist employee.
Is it safe to delegate a whole task to AI and walk away?
Only for tasks that are bounded and reversible, and even then with a review step for anything that acts in the world. Benchmark evidence such as tau-bench shows even strong agents are inconsistent under repetition, so well-built delegation keeps a human approval gate on any action that sends a message or spends money.
How is a business using AI different from me using ChatGPT?
When you type in a chat window you are present for every turn, so you naturally collaborate. An AI embedded in a booking flow or back-office pipeline runs unattended, so it is designed to take a scoped task and return a result. The interface changes the behavior, which is why business usage skews to delegation and personal use skews to conversation.
What should a small business delegate to AI first?
Start with the bounded, repeatable back-office and follow-up steps that already run to a fixed pattern, intake capture, appointment reminders, review requests, and record sync across the tools you use, with a human approving anything customer-facing. That is exactly the operational work Anthropic's data show business API use moving toward.
Provenance
Sources
- Anthropic, "The Anthropic Economic Index" (usage telemetry across 300,000+ business customers), 2026 report editions, anthropic.com/economic-index (established)
- Anthropic, "Building Effective AI Agents", 2024-2025, anthropic.com/research/building-effective-agents (established)
- Yao, S. et al., "tau-bench: A Benchmark for Tool-Agent-User Interaction in Real-World Domains", arXiv:2406.12045, 2024 (Sierra Research) (established)arxiv.org
- Brynjolfsson, E., Li, D. & Raymond, L., "Generative AI at Work", NBER Working Paper 31161 (2023); Quarterly Journal of Economics 140(2), 889-967 (2025) (established)nber.org
- MIT NANDA, "The GenAI Divide: State of AI in Business 2025", 2025 (established, single influential survey/case-study report)nanda.media.mit.edu
- Stanford HAI, "The 2026 AI Index Report", hai.stanford.edu/ai-index/2026-ai-index-report (established)hai.stanford.edu
- McKinsey & Company, "The State of AI: Global Survey 2025", November 2025 (established)
- METR, "Measuring AI Ability to Complete Long Tasks", 2025-03-19, metr.org (measurement established; long-horizon extrapolation contested)
- Moffatt v. Air Canada, 2024 BCCRT 149 (British Columbia Civil Resolution Tribunal, 2024-02-14) (established)canlii.org
- U.S. Federal Trade Commission, Operation AI Comply and the In re DoNotPay final order, 2024-09 through 2025-01-16 (established)
Every figure above is attributed to a real, dated source and tagged with its evidence tier. Where a claim could not be verified to a primary source, it is not stated as fact.