Digna Legi guide / AI productivity
Stop Calling AI Activity Productivity
Seven AI productivity metrics can rise while the business stands still. Use an evidence chain that connects activity to accepted work and captured value.

The dashboard says the AI programme is working.
Weekly users are up. Employees are sending more prompts. The model produces more code, copy, specifications, and support replies. Suggestion acceptance is rising. Inference is getting cheaper. A survey says people save several hours a week.
Every number may be correct. The conclusion may still be false.
Ask what changed after the model returned its answer. Did more work survive review? Did the queue get shorter? Did defects fall? Did the company remove a cost, avoid a loss, earn revenue, or move scarce capacity to a better use?
If nobody can answer, the dashboard has measured contact with a tool—not productivity.
AI productivity is a causal claim about a changed system, not a count of activity near a model.
That distinction matters because the most available metrics sit closest to the technology and furthest from the result used to justify the investment. The distance between the reported number and the claimed outcome is filled with assumptions about review, rework, adoption, bottlenecks, quality, and organisational follow-through.
The seven numbers below are not fraudulent or useless. They become lies when they are asked to prove more than they observed.
Productivity is a chain that can break
An AI investment has to cross four boundaries before it becomes a business result.
- Exposure
- A person used a model, feature, or agent on a defined task.
- Task effect
- The complete task became faster, better, more reliable, or newly possible after prompting, checking, and correction.
- Workflow effect
- Accepted throughput, total lead time, queues, rework, reliability, or exception handling improved across the process.
- Captured value
- Revenue, cost, loss, risk, retention, or deliberately reassigned capacity changed because the workflow changed.
Call the number of unproven boundaries between a metric and the decision it is used to justify measurement distance.
A prompt count used to operate a product has little measurement distance. The same count used to defend an enterprise AI budget has enormous distance. A task-time experiment sits closer to productivity, but it still does not establish that the surrounding workflow improved or that the organisation captured the gain.
This is the synthesis the dashboard usually hides: AI can succeed at one boundary and fail at the next. More use can produce no task improvement. A faster task can feed a slower review queue. A faster workflow can create capacity that nobody redeploys. A business result can occur without enough evidence to attribute it to AI.
The right question for every reported metric is therefore:
What must be true between this number and the outcome we are claiming?
The longer the answer, the weaker the metric is as decision evidence.
The evidence is not contradictory
AI productivity research appears to tell incompatible stories.
An NBER study of 5,179 customer-support agents found that access to an AI assistant increased issues resolved per hour by 14 percent on average. Novice and lower-skilled workers improved by 34 percent; the effect on experienced and higher-skilled workers was minimal.
A Harvard and BCG field experiment with 758 consultants found gains in speed, completion, and human-rated quality for tasks inside the model's capability frontier. On a task outside that frontier, participants using AI were less likely to reach the correct answer.
In a 2025 randomised study of experienced open-source developers, METR observed 16 developers completing 246 tasks in mature repositories. They expected early-2025 AI tools to make them 24 percent faster and later believed the tools had made them 20 percent faster. Measured completion time went the other way: AI-allowed tasks took 19 percent longer. METR warns against generalising the result to most software work, and its later experiment could not produce a reliable current estimate after participation and measurement changed.
These results do not average into a universal AI productivity percentage. Together, they point to a more useful conclusion: the effect belongs to a particular combination of person, task, model, tools, verification burden, and workflow. Change the combination and the result can change direction.
The variation is not noise around the answer. It is the deployment answer.
This is why an organisation should distrust any metric that erases task boundaries, worker experience, review cost, or downstream constraints. The number may be accurate at its own layer while being useless for the decision above it.
The first lie: exposure is treated as improvement
Three popular metrics describe what happened at the interface. None establishes what happened to the work.
- 1. Weekly active users
- Adoption shows that people opened or used a product. It can rise because the tool is valuable, mandatory, novel, built into the default workflow, or politically difficult to resist. You Cannot Mandate an AI Transformation names the organisational mistake: distributing a tool does not redesign the system expected to absorb it. Use adoption to manage rollout and support. Do not rename it return.
- 2. Generated output
- AI makes production cheap, which makes production volume a worse proxy for value. More code, campaigns, reports, tests, tickets, or specifications can increase the burden on every downstream reviewer. A team that doubles the number of product requirements without increasing engineering capacity or decision quality has created inventory, not productivity. Count accepted outputs meeting a stable definition of done and advancing a valuable outcome. Tokens, Hours, Points, and Other Curious Proxies supplies the broader warning: once a convenient count becomes the goal, the organisation learns to produce the count.
- 3. Suggestion acceptance rate
- Acceptance proves that a person chose to keep an output at one moment. It does not prove that the output survived review, avoided correction, worked in production, or helped a customer. The consultant experiment demonstrates why this boundary matters: reliance can look locally useful even when the task sits outside the model's capability frontier. Measure first-review acceptance, correction distance, reversal, escaped defects, downstream use, and verification time.
These metrics remain useful as operational telemetry. Product teams need to know whether a feature is used, which outputs are accepted, and where users abandon it. The error begins when telemetry is promoted into evidence of organisational productivity without validating the next boundary.
Exposure tells you where to investigate. It does not tell you what the investigation will find.
The second lie: local efficiency is treated as economics
Two numbers turn a fast interaction into an implausibly precise financial claim.
- 4. Self-reported hours saved
- Asking people how much time they saved measures perceived assistance. It does not measure elapsed time. People notice the instant first draft and forget the scattered minutes spent preparing context, checking facts, repairing output, resolving exceptions, and coordinating the result. The METR experiment matters here not because it proves that AI slows developers, but because perceived and observed speed diverged enough to reverse the sign of the result.
- 5. Cost per token or licence
- The model bill is visible; the system cost is distributed. The organisation also pays for integration, permissions, context preparation, training, evaluations, review, rework, failure recovery, governance, maintenance, and additional work pushed into downstream queues. Comparing a model call with a worker's salary is rarely like for like. One price buys generated output. The other may include judgment, accountability, communication, and exception handling.
Consider a deliberately hypothetical support workflow. The arithmetic is illustrative, not a research result.
- Baseline
- One thousand cases take 18 minutes each from intake to first resolution. Eight percent reopen and take another 12 minutes. Total work: 316 hours.
- AI report
- Drafting becomes seven minutes faster. The tool claims 117 hours saved.
- Whole workflow
- Review expands, first resolution takes 14 minutes, and 14 percent of cases reopen. Total work: 261 hours.
- Observed result
- The workflow saved 55 hours, not 117.
- Captured result
- If the organisation ships no additional valuable work, removes no cost, avoids no hiring, and deliberately reallocates none of the capacity, the 55 hours remain potential value. They are not a realised financial return.
The AI ROI guide develops the full calculation. The judgment here is narrower: no saved hour should enter a financial headline until the organisation can name the changed consequence.
Potential capacity is not money. It becomes value only when the operating system captures it.
The third lie: portable numbers are treated as local truth
The final two metrics travel well between presentations. That portability is exactly what strips away the deployment decision.
- 6. Average productivity gain
- An average can be statistically accurate and operationally evasive. The support study's 14 percent mean concealed a much larger gain for novices and little effect for experienced workers. A company-wide percentage also mixes routine and exceptional tasks, clean and ambiguous inputs, and errors with radically different consequences. Report effects by task type, experience, input quality, workflow stage, and consequence of failure. The distribution tells you where to scale, assist, narrow, or stop.
- 7. Benchmark score
- Benchmarks answer controlled questions about capability. They can compare models, reveal regressions, and track technical progress. Production adds incomplete context, drifting inputs, permissions, failing tools, disagreement between reviewers, latency, escalation, and errors compounded across steps. How to evaluate LLMs before production argues for evaluations that resemble the environment of use; Your Agent Isn't Dumb. Your Tools Are. shifts attention to the interfaces and feedback surrounding the model.
DORA's 2025 research describes AI as an amplifier of an organisation's existing strengths and weaknesses. That is consistent with the wider evidence: capability reaches the business through a delivery system. A higher benchmark score matters, but it cannot tell you whether that system will convert capability into a reliable outcome.
Pair offline evaluation with live measures of accepted throughput, total lead time, rework, reliability, and business effect. A benchmark belongs near the beginning of the evidence chain, not at the end.
Replace the dashboard with an evidence chain
A useful AI productivity report should be readable from left to right.
- Exposure
- Who used which model, tools, context, and workflow—and on what tasks?
- Task effect
- What happened to completion, hands-on time, accuracy, and first-pass quality after prompting, checking, and correcting?
- Workflow effect
- What happened to accepted throughput, total lead time, queues, rework, reliability, and exceptions?
- Captured value
- What revenue, cost, loss, risk, retention, or deliberately reassigned capacity changed because the workflow changed?
Every row needs a baseline, a comparison condition, uncertainty, and segmentation. Every arrow between rows needs evidence. If exposure rises without task improvement, the rollout spread but the work did not improve. If the task improves without the workflow, the bottleneck moved. If the workflow improves without captured value, the organisation created optionality and has not yet realised a return. If a business measure changes without a credible mechanism, the result may be real but should not automatically be attributed to AI.
This chain is less flattering than an adoption chart. It is also more useful. It shows where the claimed gain disappeared and which operating decision follows.
Our view: remove activity from the productivity headline
Stop putting weekly users, prompt volume, generated output, suggestion acceptance, self-reported hours, token price, averaged gains, and benchmark scores in the executive productivity headline.
Keep them in the operating dashboard. Label them by what they actually measure. Use them to find adoption friction, compare models, manage capacity, or choose where to investigate. But do not let them cross an untested boundary and become evidence of return.
For investment decisions, require three things:
- A measured change in the complete task or workflow, including verification and rework.
- A credible comparison showing what would probably have happened without the intervention.
- A named mechanism through which the organisation captured or will deliberately capture the gain.
Then make a decision rather than announcing a percentage.
- Scale
- The complete workflow improves, quality holds, the effect survives segmentation, and the capture mechanism is credible.
- Narrow
- The gain belongs to particular people, task types, or risk bands.
- Redesign
- A task accelerates but review, queues, coordination, or incentives absorb the gain.
- Stop
- Rework, failure, or weak outcomes consume the apparent benefit.
The strongest objection is that leaders cannot wait for lagging business outcomes before managing a new technology. Correct. Leading indicators are necessary. But a leading indicator earns decision weight by predicting the next boundary. Until that relationship is demonstrated, it is a hypothesis to monitor—not a return to report.
What would change our view
This position should be falsifiable.
If organisations repeatedly show that a particular activity metric reliably predicts later accepted throughput or captured value across comparable workflows—and that the relationship survives changes in task mix, worker experience, model generation, and review burden—then that metric should receive more weight in investment decisions.
The standard is not perfection. It is an observed bridge between the number and the decision.
Until that bridge exists, the honest conclusion is often narrower: people used the tool; a task may have improved; the workflow has not yet proved it; value remains uncaptured.
That answer will produce fewer celebratory slides. It will prevent more expensive mistakes.
A metric deserves the name productivity only after the evidence crosses the same boundaries as the work.
Sources and further reading
- Joel Becker, Nate Rush, Beth Barnes, and David Rein, Measuring the Impact of Early-2025 AI on Experienced Open-Source Developer Productivity, METR, July 2025.
- METR, We are Changing our Developer Productivity Experiment Design, February 2026.
- Erik Brynjolfsson, Danielle Li, and Lindsey R. Raymond, Generative AI at Work, NBER Working Paper 31161, revised November 2023 and published in The Quarterly Journal of Economics in 2025.
- Fabrizio Dell'Acqua and colleagues, Navigating the Jagged Technological Frontier, published in Organization Science in 2026.
- DORA, State of AI-assisted Software Development 2025.