A single live-event win-back campaign is aimed at a lapsed/at-risk
cohort of an alliance MMO. The data is generated with a real causal
structure so it can showcase four failure modes from
docs/reengagement_experiment_design.md:
over-simplification, contamination,
always-positive attribution, and small-sample
& whale fragility. Each mistake pushes toward the same
confident, wrong conclusion.
| players | reachable | overall_return | alliances |
|---|---|---|---|
| 14000 | NA | 54% | 460 |
“Not seen in X days” and install-anchored D14/D30 conflate slow-cadence committed players with real churners. A cadence-aware churn score separates actual returners far better than recency, and most “returns” are shallow event bounces.
## Of 'returners', only 22% are durable (>=5 active days) — the rest are shallow event bounces.
## days_since_install ranges 45-900 days, so install-anchored D14/D30 is undefined for this cohort.
The nudge lifts returns, and a returning player rallies their alliance — pulling alliance-mates (control included) back too. Under per-player randomization treated and control share alliances, so spillover lifts control and the gap collapses. Under alliance-cluster randomization a pure-control alliance never gets the seed, so the true (total) effect reappears.
## per-player gap +0.027 vs alliance gap +0.149 -> per-player understates the true effect 5.6x
Players were selected on a transient low (they lapsed), so their spend reverts upward regardless of treatment. Counting all post-campaign revenue as “incremental,” or a single-arm pre/post, shows a big win for the treated and an equally big one for control — the method can only ever say positive. Only the randomized control reveals the true (small, non-significant) effect.
## control pre->post: $0.18 -> $1.81 (regression to the mean); valid T-C: $+0.198 (p=0.26)
The per-player effect is small (§2). At the full 14k it is technically significant, but a real campaign — after the reachability and eligibility funnel — is small, and at those sizes the per-player gap is not distinguishable from zero: it reads as a null.
## revenue ARPU lift @n=1500: $+0.55 (95% CI [$-0.3, $+1.4]); drop top whale ($180) -> $+0.30 — one player swings it.
On one dataset, four independent mistakes each push toward a confident wrong call: over-simplification picks the wrong at-risk players and counts shallow bounces as wins; contamination makes the per-player test understate the real effect ~6× (a returning player rallies the whole alliance, control included); always-positive attribution manufactures a win the control group shares; and small samples turn the real (attenuated) effect into an unfalsifiable null while whale-skewed revenue makes any spend read unstable.
The fixes are in docs/reengagement_experiment_design.md:
campaign-anchored rolling durable metrics, uplift targeting, an
alliance/cluster unit matched to the interference
graph, a frozen reachable frame with
ITT/CACE/population-incremental reporting, a randomized
hold-back for incrementality, and power sized for
clusters, not players.
Generated by the Savepoint Analytics video-game A/B testing case
study. Companion to data/mocks/ and
docs/reengagement_experiment_design.md.