Phone: 0 (552) 380 25 25  |  Weekdays 10:00–18:00 · Technical support 24/7

🇹🇷 TR

Digital Bridge Blog

Artificial Intelligence

AI Productivity Impact: What Controlled Experiments and Field Studies Actually Show

In controlled experiments, the AI productivity impact ranged from 12% to 56% depending on the task. See who gains most and why it rarely reaches profit.

10 min read  · Digital Bridge Engineering Team
AI Productivity Impact: What Controlled Experiments and Field Studies Actually Show

The AI productivity impact has been measured in controlled experiments on specific tasks, and it is real: customer support agents resolved 14% more issues per hour and developers completed 26% more tasks. But gains depend on the task and the worker's experience, and individual speed-ups do not reach company profit unless workflows are redesigned.

The question on every leadership agenda

Some of your staff are almost certainly using AI already, to draft emails, summarise documents or fix code. The question that reaches the board is always the same: is this producing measurable productivity, or is it just a habit that feels useful? Vendor slides full of confident percentages rarely answer it, because they do not say what was measured, for whom, or under what conditions.

This article starts with primary field and experimental research rather than opinion surveys: studies with a control group, a defined sample and a defined task. We then look at the evidence on why those individual gains so often disappear at company level. Calculating return on investment is a separate exercise, which we cover step by step in our guide to measuring AI project ROI.

AI productivity impact: what the controlled experiments measured

The strongest evidence comes from experiments in which two groups do the same work and differ only in access to AI. The table below sets five of the most prominent studies side by side. Each row applies to one task and one sample; reading any single percentage as "AI raises productivity by X" would be a mistake.

StudyWho, which taskMeasured result
NBER — Generative AI at Work5,179 customer support agentsIssues resolved per hour up 14% on average
GitHub Copilot experiment, arXiv (2023)95 professional developers writing an HTTP serverTask finished 55.8% faster (71 vs 161 minutes)
MIT — Cui et al., three field experiments4,867 developers at three companiesCompleted tasks up 26.08%
HBS–BCG experiment (2023)BCG consultants, 18 realistic tasks12.2% more tasks, 25.1% faster, over 40% higher quality
MIT — Noy and Zhang (2023)444 professionals, workplace writingTime down 37%, grades up 0.45 standard deviations

Two patterns stand out. Single-task, lab-style experiments (Copilot, writing) show large speed gains, while months-long field studies in real workplaces (NBER, Cui et al.) show more modest but still meaningful effects. And none of the studies measures "the whole job"; each measures one specific, repeated task.

Who gains, and where the effect reverses

The most consistent finding is that gains are unevenly spread. In the NBER study the improvement reached 34% for novice and low-skilled agents, with minimal impact on experienced ones. In the HBS–BCG experiment consultants below the average improved by 43% and those above it by 17%. AI acts as a way of passing the know-how of experienced staff on to newcomers, which is why the split of tasks between people and tools needs deliberate design, as we discuss in human-AI collaboration at work.

The other side of that experiment gets less attention. On a task deliberately chosen to lie outside AI's capability frontier, consultants using AI were 19 percentage points less likely to reach the correct answer than those working without it. So the tool does not speed up every job; on some, it lowers quality with confident but wrong answers. We look at how to reduce that risk in our article on reducing AI hallucinations.

The practical lesson is to look for early gains in knowledge-heavy tasks with long onboarding curves. In customer service, that means an assistant that suggests replies and summarises cases for the human agent, described in our piece on agent assist for customer service.

Why individual gains rarely reach the bottom line

The speed-ups seen in experiments do not show up on the balance sheet to the same degree. Time savings at worker level are real but limited: in a November 2024 survey analysed by St. Louis Fed researchers, US workers using generative AI saved an average of 5.4% of their hours, or 2.2 hours in a 40-hour week. Across all workers, including non-users, the figure falls to 1.4% of total hours.

In McKinsey's "The state of AI in 2025" survey, only 39% of respondents attributed any level of enterprise EBIT impact to AI, and most of them said less than 5% of EBIT was attributable to AI use. (McKinsey — The state of AI in 2025)

An NBER study by Humlum and Vestergaard, linking adoption surveys to Danish administrative records, rules out effects larger than 2% on earnings and recorded hours two years after ChatGPT's launch; what changes is how tasks are reorganised. The executive view is similar: IBM's 2025 CEO study found only 25% of AI initiatives delivered the expected ROI. And in Microsoft and LinkedIn's 2024 Work Trend Index, 59% of leaders worried about quantifying AI's productivity gains.

The reason for the gap is simple: two hours saved by one employee vanish unless the workflow routes them to other work. In the same McKinsey survey, the roughly 6% of "high performers" attributing 5% or more of EBIT to AI were nearly three times as likely as others to have fundamentally redesigned their workflows. BCG's 2024 research likewise reports that AI leaders put 70% of their resources into people and processes.

Six steps to see the productivity effect in your own company

The design of these experiments offers a method you can reuse internally. The steps below are not a spreadsheet; they are a framework for setting up the measurement:

  1. Define the task narrowly. Not "speed up the sales team" but "write the first draft of a quotation email": a task with a clear start and end. Every experiment above worked this way.
  2. Measure the baseline. Before handing out the tool, record time per task and error or rework rates for a few weeks. Without a baseline, any gain you claim later is open to dispute.
  3. Set up a comparison group. Ideally one team uses the tool while another doing the same work does not; failing that, compare the same team before and after.
  4. Measure quality as well as speed. The HBS–BCG experiment showed speed can cost accuracy on tasks outside AI's frontier. Sample outputs and have a person score them.
  5. Split results by experience. Gains may be large for newcomers and small for experts; an average on its own misleads.
  6. Plan where the saved time goes. If freed hours are not directed to backlog, more customers or quality checks, they will never appear in the accounts.

Use this checklist to choose which task to start with:

Task characteristicGains likelyNeeds care
Type of outputDraft, summary, classification, translationCalculation or judgement with one right answer
VerificationA person can check it quicklySpotting errors requires deep expertise
Knowledge sourceCompany documents are accessibleInformation is scattered or out of date
UserNewcomer, someone searching for informationExpert who already does it very quickly
VolumeRepeated dozens of times a dayDone a few times a month

If the task follows fixed rules, you may not need AI at all; we discuss that distinction in AI versus rule-based automation. To score and rank candidate tasks department by department, use our AI use case prioritisation matrix.

What public case studies add

Beyond the experiments, companies publish their own results; these are not independently audited and often rest on user self-reports. According to a customer story Microsoft published in December 2024, participants in a 300-person Copilot trial at Vodafone reported saving four hours per person per week, and the company is extending use to 68,000 employees.

A narrower, more measurable example comes from manufacturing. In a Microsoft customer story from November 2024, Eaton says that across 1,000 standard operating procedures drafted with Copilot's help, time per procedure dropped from one hour to ten minutes (an 83% reduction). Both cases echo the experimental lesson: the gain is clearest on a repeated writing task with a known format.

How we approach this at Digital Bridge

We start with measurement, not productivity claims. In the discovery phase of our digital transformation consultancy, we map your teams' repeated tasks, the time they take and where the knowledge behind them sits. The thinking behind that needs analysis is set out in our article on software requirements analysis.

We then design a pilot around a single task that fits the checklist above. For knowledge-heavy work we build an enterprise LLM assistant that answers from your own documents and cites its sources; the technical basis is explained in RAG for enterprise LLMs. The baseline and comparison method are agreed in writing before the tool goes live.

So the gain is not lost, we place the tool inside the screens people already use, such as ERP, CRM or email, through AI integration and, where needed, system integrations. Where part of the task is rule-based, we handle that step with RPA process automation and keep AI for the steps that need judgement. We also recommend a clear company AI acceptable use policy so staff know which tools they may use with which data.

At the end of the pilot we report plainly which task and which group of employees saw gains, and where they did not. If you decide to scale, we tie that step into your digital transformation roadmap. For the general order of play, read where to start with AI in business.

Next step

This week, list three writing, summarising or information-search tasks that your teams repeat at least ten times a day, and note roughly how long each takes today. That list is the starting point for a measurable pilot. Send it to us through our contact page and we will use the discovery call to agree which task to test first. For more on the topic, browse our artificial intelligence articles.

Let us look at your case

Tell us about your process; after a needs analysis we send a written proposal with scope, phases and cost.

Request a Quote +90 552 380 25 25
Questions we hear most often

Frequently Asked Questions

How much does AI increase productivity?

There is no single figure. In controlled experiments the result varied by task: customer support agents resolved 14% more issues per hour, developers across three companies completed about 26% more tasks, and in a 2023 single-task coding experiment speed rose by more than 55%. These numbers belong to specific tasks and samples; only a pilot with a proper baseline shows the effect in your own company.

Which employees benefit most from AI?

In the experiments, the largest gains went to newcomers and below-average performers. In the customer support study, novice agents improved by 34% while experienced agents saw minimal change. That makes AI a strong tool for shortening onboarding and spreading the know-how of experienced staff across a team, rather than a way to make top performers dramatically faster.

Does AI improve productivity on every task?

No. In the 2023 Harvard and BCG experiment with consultants, on a task chosen to lie outside AI's capability frontier, those using AI were 19 percentage points less likely to reach the correct answer than those without it. For tasks with one right answer, where errors are hard to spot, human review and a system that cites its sources are essential.

If employees save time, why doesn't it show up in profit?

Saved time is lost unless it is redirected to other work. Surveys show that companies reporting meaningful profit impact from AI are a minority, and those capturing value tend to have redesigned their workflows. Rolling out a tool is not enough; you need to plan in advance where freed hours will go and how the result will be measured.

How do you measure AI's productivity impact in your own company?

You do not need a large dataset. For one narrowly defined task, a few weeks of baseline timings, an error or rework rate and, ideally, a comparison group that does not use the tool are enough. Scoring a sample of outputs for quality, not just speed, and splitting the results by experience level stops a single average from misleading you.

Have a different question? Ask Us

Talk to an Engineer

Tell us what you need to solve. We'll come back with a written proposal.