Valorice · Phase 4 · Deliver · Impact Measurement
From claimed value to proven value.
A working method to measure impact per initiative — using full causal techniques so every euro on the Journey-to-ROI scorecard is defended by evidence, not by a slide. Numbers in the worked examples are illustrative.
Why now
Most CX claims fail the audit — measurement is what survives it.
Four numbers that justify the discipline.
60%
of CX programmes
cannot demonstrate measurable financial impact to the CFO.
Forrester
3×
gap between claimed and realised value
across digital transformation initiatives.
McKinsey
70%
of A/B tests run flat or negative
when measured rigorously — most “wins” are noise.
Harvard Business Review
85%
of executives demand evidence
before approving incremental CX investment.
Gartner
Continuity
Measurement closes the loop — without it, ROI is a forecast forever.
Where impact measurement sits in the Valorice sequence.
TOM
Target Operating Model
POLISM design
→
TJM
Target Journey Maps
Future-state journeys
→
ROI
Journey-to-ROI
Sized opportunity portfolio
→
CHG
Change & Adoption
Behaviours that stick
→
MEAS
Impact Measurement
Proven euros — and what to do next
↺ loop-back Measurement feeds actuals back into the Journey-to-ROI scorecard. Confidence is refreshed, sequencing is re-evaluated, sunset criteria fire. A claimed benefit is a hypothesis; only a counterfactual turns it into evidence.
Six principles
Discipline before data — these six rules earn the right to a number.
Skip any one and the read is no longer trustworthy.
01Pre-registerHypothesis, metric, threshold and analysis plan locked before treatment. No moving the goalposts.
02BaselineStable pre-intervention window measured the same way as post. No baselines invented after the fact.
03CounterfactualA credible answer to “what would have happened anyway?” RCT, control group, synthetic, or difference-in-differences.
04Dose-responseStronger treatment should produce stronger effect. If it doesn’t, the mechanism is wrong.
05SustainmentMeasure again at T+6 and T+12 months. Many “wins” decay; some only emerge later.
06SunsetPre-commit to when measurement stops and the metric is retired. Avoid forever-pilot syndrome.
Design canvas
Lock these seven cells per initiative — before treatment starts.
Print one canvas per opportunity. Sign it. Log it. Then run.
Outcome
The business outcome in plain language. e.g. “Reduce onboarding drop-off and lift first-90-day revenue.”
KPI
The single primary metric. One. Pre-named guardrail metrics allowed; no metric hopping.
Baseline
Window, definition, value. Same calculation pre and post. Document the SQL or query.
Treatment
Population, intensity, start date, end date. Who actually gets it.
Control
How counterfactual is constructed: RCT, hold-out, matched, synthetic, DiD pair.
Horizon
T+0 read · T+90 confirm · T+180 sustain · T+365 sunset review. Cadence fixed.
Confidence
Effect size needed to call a win, statistical threshold (p<0.05 or Bayesian posterior), and minimum sample size pre-computed from minimum detectable effect.
Method decision tree
Pick the method that fits the constraints — not the constraint that fits the method.
Three questions decide it.
Can you randomise the treatment?
Volume · ethics · ops
RCT
A/B test · hold-out · randomised pilot
↓ NO
Is there a comparable untreated group?
Geo, cohort, segment
Quasi-experiment
DiD · pre-post + control · stepped wedge
↓ NO
Aggregate trend + donor pool?
Multiple similar units
Synthetic control
Weighted donor pool builds the counterfactual
↓ NO
Uplift modelling on observational data
Propensity-matched · instrumental variable · heterogeneous treatment effects
Four methods
From randomised to observational — four ways to defend a number.
In order of evidential strength.
Method 01
RCT — randomised controlled trial
How it works
- Randomly assign eligible units to treatment or control
- Treatment receives the redesigned journey; control keeps the old
- Measure the primary KPI for both groups over the horizon
- Effect = mean(treatment) − mean(control), with confidence interval
- Pre-specify sample size from minimum detectable effect
Pitfalls
- Contamination between arms
- Underpowered sample
- Novelty effect (decays)
When to use — high volume, low ethical risk, treatment can be withheld from control, randomisation operationally feasible (cohorts, geos, app flags).
Method 02
Quasi-experiment
Three flavours
- DiD — compare change in treated vs change in control across the same pre/post window. Removes time-invariant differences.
- Matched control — build a comparison cohort on propensity score, then compare deltas. Works at customer level.
- Stepped-wedge — sequential rollout across waves; early units treated, later units form rolling control until their turn.
Pitfalls
- Non-parallel pre-trends invalidate DiD
- Selection bias if treatment was self-chosen
- Time-varying confounders (competitor moves, macro shocks)
When to use — can’t randomise but a credible comparison group exists.
Method 03
Synthetic control
How it works
- Identify the treated unit (e.g. one market, branch, segment)
- Identify a donor pool of similar untreated units
- Algorithmically weight donors so they reproduce the pre-treatment trend of the treated unit
- The weighted combination becomes the synthetic counterfactual
- Effect = treated − synthetic, tested with placebo permutations
Pitfalls
- Donor pool contamination (donors partially treated)
- Poor pre-period fit invalidates the synthetic
- Over-fitting on short series
When to use — treatment rolled out to a single unit, multiple plausible donors exist, enough pre-treatment history to fit.
Method 04
Uplift modelling
How it works
- Model the heterogeneous treatment effect — different customers respond differently
- Predict τ(x) = P(outcome | treated, x) − P(outcome | control, x) per customer
- Segment by predicted uplift; target only Persuadables
- Validate with held-out RCT slice or propensity-matched test
Persuadables matrix
Persuadables
respond positively to treatment
Sure-things
convert anyway — don’t spend
Lost causes
don’t convert either way
Sleeping dogs
treatment makes them churn
When to use — heterogeneous response, scale targeting, churn-aware comms.
ROI loop-back
Each measurement refreshes the Journey-to-ROI scorecard — automatically.
Three steps, then a pre-committed rule.
Read
Measurement read
T+90 / T+180 / T+365 read against pre-registered design canvas.
Update
Confidence update
Effect size + CI feed back into the OPP-XX scorecard: confidence goes up or down.
Decide
Portfolio decision
Re-sequence, double-down, or sunset. Decision rule pre-committed at design.
Pre-committed decision rules
Effect ≥ pre-registered threshold and CI excludes zero
→ Confidence high · scale to full population
Effect positive but CI crosses zero
→ Confidence medium · extend measurement window
Effect null or negative
→ Confidence low · sunset and free capacity for next OPP
Governance
Six rules that stop measurement from drifting into theatre.
Without these, every read becomes optional.
Single source of truth
One owner per metric. One query repository. One scorecard system. No parallel claims.
Pre-registration registry
Every initiative’s design canvas is logged before treatment. Auditable, time-stamped.
Sunset clock
Default sunset at T+365. Renewal requires positive evidence — not silence.
Independent read
Analytics function (not the project owner) runs and reports the read. Removes optimism bias.
Negative results published
Null and negative reads land in the same place as wins. Penalises hiding, not failing.
Annual portfolio audit
Sample 10% of past wins, re-measure with current data. Catches regression and gaming.
Next steps
From method to running system in four moves.
Four weeks, then it runs itself.
1
Pick three OPPs
Take the top three from the Journey-to-ROI scorecard and apply the design canvas.
2
Pre-register
Lock canvas + method per OPP in the registry before any treatment goes live.
3
Run reads
T+90 confirm · T+180 sustain · T+365 sunset — feed every read into the scorecard.
4
Quarterly review
Refresh confidence, re-sequence the portfolio, retire what didn’t work.