Valorice · Phase 4 · Deliver · Impact Measurement
Deliverable 6

From claimed value to proven value.

A working method to measure impact per initiative — using full causal techniques so every euro on the Journey-to-ROI scorecard is defended by evidence, not by a slide. Numbers in the worked examples are illustrative.

Why now

Most CX claims fail the audit — measurement is what survives it.

Four numbers that justify the discipline.
60%
of CX programmes
cannot demonstrate measurable financial impact to the CFO.
Forrester
gap between claimed and realised value
across digital transformation initiatives.
McKinsey
70%
of A/B tests run flat or negative
when measured rigorously — most “wins” are noise.
Harvard Business Review
85%
of executives demand evidence
before approving incremental CX investment.
Gartner
Continuity

Measurement closes the loop — without it, ROI is a forecast forever.

Where impact measurement sits in the Valorice sequence.
TOM
Target Operating Model
POLISM design
TJM
Target Journey Maps
Future-state journeys
ROI
Journey-to-ROI
Sized opportunity portfolio
CHG
Change & Adoption
Behaviours that stick
MEAS
Impact Measurement
Proven euros — and what to do next
↺ loop-back   Measurement feeds actuals back into the Journey-to-ROI scorecard. Confidence is refreshed, sequencing is re-evaluated, sunset criteria fire. A claimed benefit is a hypothesis; only a counterfactual turns it into evidence.
Six principles

Discipline before data — these six rules earn the right to a number.

Skip any one and the read is no longer trustworthy.
01Pre-register
Hypothesis, metric, threshold and analysis plan locked before treatment. No moving the goalposts.
02Baseline
Stable pre-intervention window measured the same way as post. No baselines invented after the fact.
03Counterfactual
A credible answer to “what would have happened anyway?” RCT, control group, synthetic, or difference-in-differences.
04Dose-response
Stronger treatment should produce stronger effect. If it doesn’t, the mechanism is wrong.
05Sustainment
Measure again at T+6 and T+12 months. Many “wins” decay; some only emerge later.
06Sunset
Pre-commit to when measurement stops and the metric is retired. Avoid forever-pilot syndrome.
Design canvas

Lock these seven cells per initiative — before treatment starts.

Print one canvas per opportunity. Sign it. Log it. Then run.
Outcome
The business outcome in plain language. e.g. “Reduce onboarding drop-off and lift first-90-day revenue.”
KPI
The single primary metric. One. Pre-named guardrail metrics allowed; no metric hopping.
Baseline
Window, definition, value. Same calculation pre and post. Document the SQL or query.
Treatment
Population, intensity, start date, end date. Who actually gets it.
Control
How counterfactual is constructed: RCT, hold-out, matched, synthetic, DiD pair.
Horizon
T+0 read · T+90 confirm · T+180 sustain · T+365 sunset review. Cadence fixed.
Confidence
Effect size needed to call a win, statistical threshold (p<0.05 or Bayesian posterior), and minimum sample size pre-computed from minimum detectable effect.
Method decision tree

Pick the method that fits the constraints — not the constraint that fits the method.

Three questions decide it.
Can you randomise the treatment?
Volume · ethics · ops
YES
RCT
A/B test · hold-out · randomised pilot
↓ NO
Is there a comparable untreated group?
Geo, cohort, segment
YES
Quasi-experiment
DiD · pre-post + control · stepped wedge
↓ NO
Aggregate trend + donor pool?
Multiple similar units
YES
Synthetic control
Weighted donor pool builds the counterfactual
↓ NO
Uplift modelling on observational data
Propensity-matched · instrumental variable · heterogeneous treatment effects
Four methods

From randomised to observational — four ways to defend a number.

In order of evidential strength.
Method 01
RCT — randomised controlled trial

How it works

  • Randomly assign eligible units to treatment or control
  • Treatment receives the redesigned journey; control keeps the old
  • Measure the primary KPI for both groups over the horizon
  • Effect = mean(treatment) − mean(control), with confidence interval
  • Pre-specify sample size from minimum detectable effect

Pitfalls

  • Contamination between arms
  • Underpowered sample
  • Novelty effect (decays)
When to use — high volume, low ethical risk, treatment can be withheld from control, randomisation operationally feasible (cohorts, geos, app flags).
Method 02
Quasi-experiment

Three flavours

  • DiD — compare change in treated vs change in control across the same pre/post window. Removes time-invariant differences.
  • Matched control — build a comparison cohort on propensity score, then compare deltas. Works at customer level.
  • Stepped-wedge — sequential rollout across waves; early units treated, later units form rolling control until their turn.

Pitfalls

  • Non-parallel pre-trends invalidate DiD
  • Selection bias if treatment was self-chosen
  • Time-varying confounders (competitor moves, macro shocks)
When to use — can’t randomise but a credible comparison group exists.
Method 03
Synthetic control

How it works

  • Identify the treated unit (e.g. one market, branch, segment)
  • Identify a donor pool of similar untreated units
  • Algorithmically weight donors so they reproduce the pre-treatment trend of the treated unit
  • The weighted combination becomes the synthetic counterfactual
  • Effect = treated − synthetic, tested with placebo permutations

Pitfalls

  • Donor pool contamination (donors partially treated)
  • Poor pre-period fit invalidates the synthetic
  • Over-fitting on short series
When to use — treatment rolled out to a single unit, multiple plausible donors exist, enough pre-treatment history to fit.
Method 04
Uplift modelling

How it works

  • Model the heterogeneous treatment effect — different customers respond differently
  • Predict τ(x) = P(outcome | treated, x) − P(outcome | control, x) per customer
  • Segment by predicted uplift; target only Persuadables
  • Validate with held-out RCT slice or propensity-matched test

Persuadables matrix

Persuadables
respond positively to treatment
Sure-things
convert anyway — don’t spend
Lost causes
don’t convert either way
Sleeping dogs
treatment makes them churn
When to use — heterogeneous response, scale targeting, churn-aware comms.
ROI loop-back

Each measurement refreshes the Journey-to-ROI scorecard — automatically.

Three steps, then a pre-committed rule.
Read
Measurement read
T+90 / T+180 / T+365 read against pre-registered design canvas.
Update
Confidence update
Effect size + CI feed back into the OPP-XX scorecard: confidence goes up or down.
Decide
Portfolio decision
Re-sequence, double-down, or sunset. Decision rule pre-committed at design.

Pre-committed decision rules

Effect ≥ pre-registered threshold and CI excludes zero
→ Confidence high · scale to full population
Effect positive but CI crosses zero
→ Confidence medium · extend measurement window
Effect null or negative
→ Confidence low · sunset and free capacity for next OPP
Governance

Six rules that stop measurement from drifting into theatre.

Without these, every read becomes optional.
Single source of truth
One owner per metric. One query repository. One scorecard system. No parallel claims.
Pre-registration registry
Every initiative’s design canvas is logged before treatment. Auditable, time-stamped.
Sunset clock
Default sunset at T+365. Renewal requires positive evidence — not silence.
Independent read
Analytics function (not the project owner) runs and reports the read. Removes optimism bias.
Negative results published
Null and negative reads land in the same place as wins. Penalises hiding, not failing.
Annual portfolio audit
Sample 10% of past wins, re-measure with current data. Catches regression and gaming.
Next steps

From method to running system in four moves.

Four weeks, then it runs itself.
References
Forrester CX ROI research · McKinsey on digital transformations · HBR — the surprising power of online experiments · Imbens & Rubin — Causal Inference for Statistics · Abadie — synthetic control review · Athey & Imbens — heterogeneous treatment effects