The AI productivity impact has been measured in controlled experiments on specific tasks, and it is real: customer support agents resolved 14% more issues per hour and developers completed 26% more tasks. But gains depend on the task and the worker's experience, and individual speed-ups do not reach company profit unless workflows are redesigned.
The question on every leadership agenda
Some of your staff are almost certainly using AI already, to draft emails, summarise documents or fix code. The question that reaches the board is always the same: is this producing measurable productivity, or is it just a habit that feels useful? Vendor slides full of confident percentages rarely answer it, because they do not say what was measured, for whom, or under what conditions.
This article starts with primary field and experimental research rather than opinion surveys: studies with a control group, a defined sample and a defined task. We then look at the evidence on why those individual gains so often disappear at company level. Calculating return on investment is a separate exercise, which we cover step by step in our guide to measuring AI project ROI.
AI productivity impact: what the controlled experiments measured
The strongest evidence comes from experiments in which two groups do the same work and differ only in access to AI. The table below sets five of the most prominent studies side by side. Each row applies to one task and one sample; reading any single percentage as "AI raises productivity by X" would be a mistake.
| Study | Who, which task | Measured result |
|---|---|---|
| NBER — Generative AI at Work | 5,179 customer support agents | Issues resolved per hour up 14% on average |
| GitHub Copilot experiment, arXiv (2023) | 95 professional developers writing an HTTP server | Task finished 55.8% faster (71 vs 161 minutes) |
| MIT — Cui et al., three field experiments | 4,867 developers at three companies | Completed tasks up 26.08% |
| HBS–BCG experiment (2023) | BCG consultants, 18 realistic tasks | 12.2% more tasks, 25.1% faster, over 40% higher quality |
| MIT — Noy and Zhang (2023) | 444 professionals, workplace writing | Time down 37%, grades up 0.45 standard deviations |
Two patterns stand out. Single-task, lab-style experiments (Copilot, writing) show large speed gains, while months-long field studies in real workplaces (NBER, Cui et al.) show more modest but still meaningful effects. And none of the studies measures "the whole job"; each measures one specific, repeated task.
Who gains, and where the effect reverses
The most consistent finding is that gains are unevenly spread. In the NBER study the improvement reached 34% for novice and low-skilled agents, with minimal impact on experienced ones. In the HBS–BCG experiment consultants below the average improved by 43% and those above it by 17%. AI acts as a way of passing the know-how of experienced staff on to newcomers, which is why the split of tasks between people and tools needs deliberate design, as we discuss in human-AI collaboration at work.
The other side of that experiment gets less attention. On a task deliberately chosen to lie outside AI's capability frontier, consultants using AI were 19 percentage points less likely to reach the correct answer than those working without it. So the tool does not speed up every job; on some, it lowers quality with confident but wrong answers. We look at how to reduce that risk in our article on reducing AI hallucinations.
The practical lesson is to look for early gains in knowledge-heavy tasks with long onboarding curves. In customer service, that means an assistant that suggests replies and summarises cases for the human agent, described in our piece on agent assist for customer service.
Why individual gains rarely reach the bottom line
The speed-ups seen in experiments do not show up on the balance sheet to the same degree. Time savings at worker level are real but limited: in a November 2024 survey analysed by St. Louis Fed researchers, US workers using generative AI saved an average of 5.4% of their hours, or 2.2 hours in a 40-hour week. Across all workers, including non-users, the figure falls to 1.4% of total hours.
In McKinsey's "The state of AI in 2025" survey, only 39% of respondents attributed any level of enterprise EBIT impact to AI, and most of them said less than 5% of EBIT was attributable to AI use. (McKinsey — The state of AI in 2025)
An NBER study by Humlum and Vestergaard, linking adoption surveys to Danish administrative records, rules out effects larger than 2% on earnings and recorded hours two years after ChatGPT's launch; what changes is how tasks are reorganised. The executive view is similar: IBM's 2025 CEO study found only 25% of AI initiatives delivered the expected ROI. And in Microsoft and LinkedIn's 2024 Work Trend Index, 59% of leaders worried about quantifying AI's productivity gains.
The reason for the gap is simple: two hours saved by one employee vanish unless the workflow routes them to other work. In the same McKinsey survey, the roughly 6% of "high performers" attributing 5% or more of EBIT to AI were nearly three times as likely as others to have fundamentally redesigned their workflows. BCG's 2024 research likewise reports that AI leaders put 70% of their resources into people and processes.
Six steps to see the productivity effect in your own company
The design of these experiments offers a method you can reuse internally. The steps below are not a spreadsheet; they are a framework for setting up the measurement:
- Define the task narrowly. Not "speed up the sales team" but "write the first draft of a quotation email": a task with a clear start and end. Every experiment above worked this way.
- Measure the baseline. Before handing out the tool, record time per task and error or rework rates for a few weeks. Without a baseline, any gain you claim later is open to dispute.
- Set up a comparison group. Ideally one team uses the tool while another doing the same work does not; failing that, compare the same team before and after.
- Measure quality as well as speed. The HBS–BCG experiment showed speed can cost accuracy on tasks outside AI's frontier. Sample outputs and have a person score them.
- Split results by experience. Gains may be large for newcomers and small for experts; an average on its own misleads.
- Plan where the saved time goes. If freed hours are not directed to backlog, more customers or quality checks, they will never appear in the accounts.
Use this checklist to choose which task to start with:
| Task characteristic | Gains likely | Needs care |
|---|---|---|
| Type of output | Draft, summary, classification, translation | Calculation or judgement with one right answer |
| Verification | A person can check it quickly | Spotting errors requires deep expertise |
| Knowledge source | Company documents are accessible | Information is scattered or out of date |
| User | Newcomer, someone searching for information | Expert who already does it very quickly |
| Volume | Repeated dozens of times a day | Done a few times a month |
If the task follows fixed rules, you may not need AI at all; we discuss that distinction in AI versus rule-based automation. To score and rank candidate tasks department by department, use our AI use case prioritisation matrix.
What public case studies add
Beyond the experiments, companies publish their own results; these are not independently audited and often rest on user self-reports. According to a customer story Microsoft published in December 2024, participants in a 300-person Copilot trial at Vodafone reported saving four hours per person per week, and the company is extending use to 68,000 employees.
A narrower, more measurable example comes from manufacturing. In a Microsoft customer story from November 2024, Eaton says that across 1,000 standard operating procedures drafted with Copilot's help, time per procedure dropped from one hour to ten minutes (an 83% reduction). Both cases echo the experimental lesson: the gain is clearest on a repeated writing task with a known format.
How we approach this at Digital Bridge
We start with measurement, not productivity claims. In the discovery phase of our digital transformation consultancy, we map your teams' repeated tasks, the time they take and where the knowledge behind them sits. The thinking behind that needs analysis is set out in our article on software requirements analysis.
We then design a pilot around a single task that fits the checklist above. For knowledge-heavy work we build an enterprise LLM assistant that answers from your own documents and cites its sources; the technical basis is explained in RAG for enterprise LLMs. The baseline and comparison method are agreed in writing before the tool goes live.
So the gain is not lost, we place the tool inside the screens people already use, such as ERP, CRM or email, through AI integration and, where needed, system integrations. Where part of the task is rule-based, we handle that step with RPA process automation and keep AI for the steps that need judgement. We also recommend a clear company AI acceptable use policy so staff know which tools they may use with which data.
At the end of the pilot we report plainly which task and which group of employees saw gains, and where they did not. If you decide to scale, we tie that step into your digital transformation roadmap. For the general order of play, read where to start with AI in business.
Next step
This week, list three writing, summarising or information-search tasks that your teams repeat at least ten times a day, and note roughly how long each takes today. That list is the starting point for a measurable pilot. Send it to us through our contact page and we will use the discovery call to agree which task to test first. For more on the topic, browse our artificial intelligence articles.