DIGNALEGI

Digna Legi guide / AI experimentation

How to Run an AI Productivity Experiment That Can Say No

A practical protocol for testing whether AI improves real work: define the decision, build a credible comparison, measure rework, and set a kill condition.

Two parallel paper-card test lanes pass through measuring instruments toward an open route and a blocked route marked by a yellow threshold.

Most AI pilots are demonstrations wearing laboratory coats.

The team gives a model to enthusiastic volunteers. Usage rises. A few people report dramatic time savings. The best examples become slides. The pilot ends with a recommendation to buy more licences.

Nothing in that sequence establishes what the tool changed. The volunteers may have been unusually capable. Their work may have become easier for unrelated reasons. The model may have accelerated drafting while pushing more work into review. People who did not use it may have improved just as quickly. A before-and-after chart cannot separate those explanations.

A larger pilot does not fix this. The work needs a different object: a decision-shaped experiment.

A decision-shaped experiment starts with the action its result must change: scale, narrow, redesign, or stop. It then works backwards to the evidence that would justify that action. It defines the intervention, comparison, outcome, guardrails, segments, and failure condition before the result is known.

An AI experiment that cannot produce a decision to stop is rollout theatre.

This guide offers a practical protocol for testing AI in real work. High-stakes studies still require specialist statistical design. The protocol stops weak pilots from manufacturing strong claims.

Start with the decision, not the tool

Write one sentence before choosing a metric:

We are testing whether to [scale, narrow, redesign, or stop] this intervention for [defined people and work] because we expect it to change [primary outcome] through [mechanism], without violating [guardrails].

“Test Copilot” is not a decision. “Determine whether to extend AI-assisted refund drafting from one support team to all low-risk refund cases” is.

The distinction forces five choices into the open:

  1. Decision owner. Name the person who will act on the result. A pilot without an owner can generate evidence indefinitely.
  2. Population. Define the people, customers, cases, repositories, or workflows to which the result may apply.
  3. Mechanism. State why the intervention should work. Faster first drafts, better retrieval, fewer hand-offs, or more consistent decisions are different theories.
  4. Primary outcome. Choose the one result that determines the decision. Other measures can explain or constrain it; they should not rescue it after failure.
  5. Decision date. Fix when evidence will be reviewed and who can extend the test.

This is also where the team names the minimum improvement worth buying. Statistical detectability is not commercial importance. A tiny, confidently measured gain can still be too small to justify licences, integration, review, governance, and switching costs. Set the minimum worthwhile effect first; then ask an analyst what sample and duration are needed to detect it with defensible uncertainty.

Stop Calling AI Activity Productivity explains why prompts, users, accepted suggestions, and self-reported hours are weak executive outcomes. They can describe exposure and help diagnose adoption. They should not become the primary outcome merely because they are easy to collect.

Separate the three units

AI experiments fail quietly when they treat three different units as if they were the same.

The unit of assignment is what receives the intervention: a task, person, team, site, customer, or time period.

The unit of work is the repeatable object that moves through the process: a ticket, pull request, claim, campaign, analysis, or decision.

The unit of value is where the organisation expects the benefit or harm to appear: a resolved case, released feature, retained customer, avoided loss, or unit of reliable capacity.

Suppose support agents receive an AI drafting tool. Agents are the unit of assignment, cases are the unit of work, and resolved cases that do not reopen may be the unit of value. Measuring message-generation time observes only part of the work. Measuring individual output while the same specialist review queue serves treatment and comparison groups can hide the bottleneck. Randomising cases within one agent may contaminate the comparison because the agent learns techniques from AI-assisted cases.

Randomise where contamination can be controlled; measure where value is supposed to appear.

There is no universally correct level.

Task-level assignment can create many observations quickly, but switching, learning, and reuse of generated material can leak treatment into control work.

Person-level assignment is easier to operate, but collaboration may spread templates and techniques between groups. It can also miss effects that appear only when a whole team changes its workflow.

Team- or site-level assignment better captures coordination and queue effects, but usually provides fewer independent groups and makes pre-existing differences more dangerous.

Time-based or phased assignment fits a staged rollout, but seasonality, concurrent projects, and model changes can become alternative explanations.

Choose the assignment level by mapping how knowledge, work, and constraints travel. Then make the limits explicit. A clean estimate of individual drafting speed does not automatically answer whether the department should scale the tool.

Specify the intervention before it changes

“Using AI” is not a reproducible treatment.

Record the model and version, system instructions, tools, retrieved data, permissions, interface, training, human responsibilities, review rules, escalation path, and excluded tasks. Capture meaningful changes during the experiment. If the vendor silently changes the model or the team improves the prompt library, the intervention has changed even if the product name has not.

The human process belongs inside the specification. Access alone is one intervention. Access plus training is another. A model that drafts and a person who verifies is not the same system as an agent that acts and asks for approval only on exceptions.

This matters because AI results belong to a configuration, not to a brand. The Harvard–BCG field experiment found gains on tasks inside the tested model's capability frontier and worse correctness on a task outside it. The result makes eligibility rules part of the treatment.

Write them down:

  • which work is eligible;
  • which work is prohibited;
  • what context the system may use;
  • what a human must verify;
  • when the system must escalate;
  • what counts as a completed case.

How to Evaluate LLMs Before Production is useful upstream: it helps establish whether a model is capable and safe enough to enter a live experiment. The experiment asks the next question: whether that capability improves the operating system in which it is placed.

Establish the baseline before behaviour changes

Baseline collection should begin before access, training, incentives, or public enthusiasm changes how people work.

Document business as usual as an actual process, not a label. Record the full path from intake to accepted completion: hands-on time, waiting, review, rework, exceptions, abandonment, and downstream correction. Preserve task mix and difficulty where possible. Note other changes that could affect the same outcome, such as staffing, demand, policy, seasonality, or a concurrent software release.

HM Treasury's 2026 guidance on evaluating AI interventions makes the same point: if the comparison is business as usual, the evaluator must define what business as usual consists of before implementation. This is especially important when AI supports subjective work whose quality is not already measured consistently.

A baseline is not automatically a counterfactual. “The team was slower last month” does not establish what would have happened this month without AI. Demand, experience, staffing, and task difficulty move. Baseline data helps describe the starting point, estimate variance, and test whether groups were comparable. The comparison design supplies the causal claim.

Build the strongest comparison the operation permits

Begin with random assignment, then step down only when the operation or risk makes it unsuitable.

Randomised access
Assign eligible tasks, people, teams, or sites to AI access or business as usual. Randomisation makes the groups comparable on average and gives the cleanest basis for attribution.
Randomised timing
If everyone will eventually receive access, randomise the order of a phased rollout. Early and later groups can provide a comparison while preserving the deployment plan.
Randomised encouragement
If access cannot be withheld, randomly assign training, prompts, office hours, or workflow nudges. This estimates the effect of the encouragement strategy, not automatically the effect of tool use itself.
Quasi-experimental comparison
If randomisation is impossible, use a credible unaffected group and a design such as matching or difference-in-differences with qualified analytical support. The burden is to show why the comparison approximates what would have happened without the intervention.
Structured process evaluation
Rare, high-stakes, or highly heterogeneous work may not produce enough comparable cases for a useful controlled estimate. Use expert review, scenario tests, trace analysis, incident evidence, and a clearly stated causal theory. Do not manufacture a precise productivity percentage from a handful of incomparable cases.

The strongest counterargument is speed: models evolve faster than conventional evaluations, so a rigorous experiment may measure a system that no longer exists. That is a design constraint, not permission to make an unidentifiable claim. Use shorter stages, stable intervention windows, phased access, and explicit version records. Make a narrower decision sooner. Repeat the evaluation when the model, workflow, or population materially changes.

The Magenta Book guidance for AI interventions supports this proportionate hierarchy: consider experimental methods first, use quasi-experimental or theory-based methods when appropriate, and align evaluation stages with an evolving intervention.

Pre-commit the scorecard

Before looking at results, write a one-page experiment contract.

Primary outcome
One measure tied to the decision, such as accepted cases per paid hour, end-to-end lead time, first-pass resolution, or loss avoided. Define its numerator, denominator, time window, and exclusions.
Guardrails
Measures that must not deteriorate beyond an agreed boundary: escaped defects, reopen rate, customer harm, security incidents, escalation load, reviewer burden, or outcome disparity between relevant groups.
Minimum worthwhile effect
The smallest improvement that would justify the total cost and operational change. This is a business threshold, not a p-value.
Segments
The few groups for which a different result would change deployment: task type, difficulty, worker experience, input quality, risk band, language, or customer group. Choose them before analysis. Searching dozens of segments after the result invites a flattering story.
Kill condition
The evidence that pauses or ends the experiment before its scheduled close. Examples include a serious safety event, a predefined quality breach, reviewer overload, or inability to maintain the comparison.
Decision rule
State how the scorecard maps to scale, narrow, redesign, or stop. Do not reduce judgment to a mechanical formula; do prevent the team from changing the standard after seeing an inconvenient result.

Consider this hypothetical support experiment. The decision is whether to scale AI-assisted drafting for low-risk refund exceptions. The primary outcome is accepted resolutions per paid hour. The team sets an illustrative minimum worthwhile improvement of 10 per cent, with no more than a two-percentage-point increase in reopen rate and no rise in policy errors. Cases are segmented by complexity and agent experience. A material policy error triggers immediate review. The thresholds are examples, not universal recommendations; the organisation must derive its own from economics, baseline variation, and risk.

A primary outcome decides whether the intervention worked. Guardrails decide whether the win is acceptable.

Measure the complete workflow

Instrument the causal path from access to captured value.

Exposure
Who received access, what they used, for which eligible work, and whether the system was available.
Task effect
Total completion time, hands-on time, first-pass quality, corrections, verification effort, and failure, rather than merely generation speed.
Workflow effect
Accepted throughput, queues, hand-offs, total lead time, rework, exceptions, reliability, and burden transferred to reviewers or downstream teams.
Captured value
Cost removed or avoided, revenue changed, loss prevented, service improved, or capacity deliberately reassigned to named work.

The AI productivity metrics guide develops this evidence chain; the experiment makes it testable. The AI ROI guide then converts a credible operational effect into an economic decision without pretending that every saved minute became cash.

Two measurement errors deserve special attention.

First, do not measure only successful outputs. Capture abandoned attempts, model outages, fallback work, corrections, and cases routed away from AI. Excluding failures makes the intervention look more reliable precisely because it failed.

Second, do not make invasive employee telemetry the default. Collect the minimum data needed for the decision, define its use, restrict access, and separate workflow evaluation from individual performance management. Surveillance changes behaviour and can destroy participation; it also produces an experiment about surveillance rather than ordinary work.

Treat adoption as part of the result

Analyse outcomes by assignment first: what happened to the group offered the intervention compared with the group assigned to business as usual? This captures the policy effect of deploying the system, including non-use, friction, and training failure.

Then examine actual use to understand the mechanism. But do not naively compare users with non-users. People who choose to use AI may have different tasks, confidence, skill, or time pressure. Their better result may explain adoption rather than be caused by it.

METR's 2026 update to its developer-productivity research is instructive. As AI tools improved and became normal, participation and compliance changed: developers were less willing to accept tasks on which AI was prohibited, making the original experiment harder to sustain and the resulting estimate less reliable. That is not a footnote. It demonstrates that experimental design is part of the evidence, and that changing selection can invalidate a familiar protocol.

Likewise, an NBER field experiment across 66 firms found that access to an integrated generative-AI tool reduced time spent on email among users and reduced after-hours work, but did not detect broader changes in the quantity or composition of tasks over the study period. Individual time use changed more clearly than organisational work. A deployment can improve one layer without moving the next.

Non-use, workarounds, and reviewer resistance are not noise to delete. They are properties of the intervention being tested.

Explain variation before averaging it away

Report the overall effect with uncertainty, then show the pre-specified segments that matter for deployment.

An average gain can conceal a strong benefit for routine cases and a loss for ambiguous ones. A quality average can hide a rare but unacceptable failure. A team average can combine novices who gain with experts who slow down. Those patterns are not reasons to discard the experiment. They are often the most actionable result.

Variation should change the product decision:

Scale when the complete workflow improves, guardrails hold, the effect survives relevant segments, and the organisation can capture the value.

Narrow when benefit belongs to defined tasks, people, or risk bands.

Redesign when local speed is real but verification, queues, hand-offs, poor tools, or incentives absorb it.

Stop when harm, rework, weak adoption, or trivial economics defeats the hypothesis.

Do not hide a failed primary outcome behind a successful secondary metric. A pilot that misses accepted throughput but raises weekly use has taught the team something about adoption, not proved productivity.

You Cannot Mandate an AI Transformation describes the organisational version of this mistake: access is distributed while the surrounding system remains unchanged. A well-designed experiment reveals that mismatch before it becomes an enterprise commitment.

End with a decision record

The final deliverable should fit on one page before the appendix begins.

  1. Decision. Scale, narrow, redesign, stop, or run a named next test.
  2. Population. Where the conclusion applies and where it does not.
  3. Intervention. The model, workflow, training, review, and eligibility rules tested.
  4. Comparison. What represented the no-AI counterfactual and why it was credible.
  5. Result. Primary outcome, uncertainty, guardrails, and relevant segments.
  6. Mechanism. Why the observed change probably occurred, including evidence against alternative explanations.
  7. Economics. The value captured or the explicit plan for capturing capacity.
  8. Next trigger. The model, workflow, risk, or volume change that requires re-evaluation.

NIST's AI Risk Management Framework calls for testing that is objective, repeatable, documented, representative of deployment conditions, and monitored after deployment. The last condition is easy to neglect. An experiment is evidence about a bounded configuration at a point in time. It is not a permanent certificate for every future model, prompt, population, and workflow.

This is why the experiment should end with both a decision and an expiry condition.

Our view: the result belongs to the operating system

Companies keep asking for the productivity of a model. The question is malformed.

A model does not enter the business alone. It arrives with task selection, context, tools, permissions, training, review, escalation, incentives, queues, and some mechanism for capturing released capacity. That mechanism may be absent. The experiment must therefore test the intervention as an operating system, not the model as an isolated object.

That standard is demanding but proportionate. A reversible, low-risk workflow can use a short randomised or phased test with a small scorecard. A consequential deployment deserves independent review, stronger quality measures, subgroup analysis, and longer monitoring. A rare workflow may require structured case evidence rather than a false claim of statistical certainty.

What must not change is the logic: a defined decision, a credible counterfactual, the complete workflow, explicit guardrails, and a result allowed to disappoint its sponsor.

An AI experiment should discover where the system earns the right to change, rather than prove that the tool works.

What would change our view

This protocol carries a cost. It would be excessive if lighter evidence proved equally reliable.

We would relax it if repeated deployments showed that simple before-and-after or activity measures predicted later workflow and business outcomes as accurately as decision-shaped counterfactual experiments, across changes in task mix, worker experience, model version, and verification burden.

We would also give a validated leading indicator more decision weight if it repeatedly predicted the next boundary in the same class of workflow. That would not abolish evaluation. It would shorten the evidence chain.

Until then, the cheap pilot is often the expensive choice. It converts uncertainty into confidence without producing knowledge, then sends the organisation into a larger rollout with the causal question still unanswered.

Sources and further reading