The core idea
A pilot explores feasibility; an impact evaluation asks what changed because of an intervention. An A/B test may use random allocation to compare approaches, but not every pilot is randomised or suitable for estimating impact. Define the outcome and analysis before seeing results where possible. The comparison should approximate what would have happened without the intervention, and practical or ethical constraints must shape the design.
Source and attribution [1]Using it in practice
Specify eligibility, allocation, outcome window and how people will be analysed. Consider whether assignment should occur by team or site to reduce spillovers. Estimate the sample needed for a meaningful effect and plan how missing outcomes will be handled. Do not withhold legal entitlements or essential support. Where randomisation is unsuitable, select a defensible alternative and report its assumptions.
An example, not a reported case
Worked example · illustrative
In an invented pilot, 90-day retention is 84% among 100 starters offered a new onboarding approach and 80% among 100 under the comparison approach. The observed difference is four percentage points. Under a simple independent-binomial approximation, its standard error is about 5.4 percentage points, so an approximate 95% interval is −6.6 to +14.6 points. The pilot is too imprecise to claim a clear retention improvement from these figures alone.
What to watch for
Before-and-after improvement can reflect seasonality or other changes. Repeatedly checking results and stopping at a favourable moment can distort inference. Small pilots may answer feasibility questions without resolving impact. Clustered allocation, non-compliance and spillovers require appropriate analysis beyond the simple example.
Define success in advance
Choose a primary outcome, meaningful effect size and review point. Include implementation measures and possible harms. A pilot can be useful if it reveals that delivery is impractical, even when it cannot estimate a precise effect.
Report the whole result
Show group counts, outcomes, uncertainty, deviations from the plan and contextual changes. Distinguish lack of evidence from evidence of no effect. Explain what the result supports doing next rather than reducing the decision to a p-value.
Take it into your next conversation
Three useful questions.
- What question can this method answer, and what can it not establish?
- Are the comparison, observation period and assumptions defensible?
- What decision follows, and how will its consequences be reviewed?
Related terms
Go to the evidence
Sources & attribution
[1] HM Treasury: Magenta Book Annex A analytical methods ↗
The core idea is an original summary of the cited work. Application notes, examples and sketches are our interpretations, not quotations or reproductions of the authors’ figures. Publisher records may require access to read the full original work.
Published 2026-09-20 · Reviewed 2026-09-20. Editorial approach