
Game analytics · Part 05 of 10 · August 23, 2026
Re-engagement, Honestly Measured
Simulated data with recorded ground truth: the mean true direct uplift is 12.7 percentage points, and the fixture also records each player’s organic return probability, so every estimate below can be scored against the answer.
The question
Should the studio keep funding a live-event win-back campaign aimed at a lapsed cohort of an alliance-based MMO?
The default answer, absent an experiment, is always yes, because the reported numbers are always good. That is what interested me. Win-back readouts have a way of never coming back negative, and a measurement that cannot produce a negative result is not a measurement. So the real question is how a campaign gets funded on nothing, and whether the same data, read properly, would have funded it anyway.
Why the effect is harder to see than it looks
Re-engagement has four separate ways to fool you, and they are independent. Fix any three and the fourth still gets the decision wrong.
The population is defined by absence, so you cannot reach most of it, and the ones you can reach are not a random sample of the ones you want. The outcome, “came back”, is trivially satisfied by logging in once for an event and leaving again. The players you are nudging live in alliances, and a returning player drags their alliance-mates back with them, including the ones in the control arm. And the whole cohort was selected at a moment of unusually low activity, which means it will regress upward whether you do anything or not.
Any one of these produces a positive number. Together they produce a very positive number, delivered with a straight face.
My approach
Fourteen thousand at-risk players across 460 alliances, mean alliance size 30.4. Tenure runs from 45 to 900 days since install. Only 37.6% of the cohort is reachable at all, by push, email or paid social. The rest cannot be contacted, for the simple reason that a player who is not logging in cannot be reached by in-game mail.
The same players are simulated under two randomization designs, per-player and alliance-cluster. Running both on identical players is what makes the comparison clean: any difference between the two estimates is the design, not the sample.
Every estimate is intention-to-treat on the assigned population unless labelled otherwise, and every conditioning variable is checked for being fixed before the campaign went out.
The assumptions doing the heavy lifting
Walking through the four pitfalls is walking through the assumptions, because each pitfall is an assumption that felt too obvious to state.
Pitfall one: the metric and the target
The assumption is that “retention” means something for this population. It does not. Install-anchored D14/D30 retention is not a hard metric here but an undefined one: tenure spans 45 to 900 days, and a player who installed eight months ago is long past D30. The anchor has to be the campaign date, not the install date.
The targeting assumption is worse. “Not seen in 30+ days” flags 20% of the cohort and correlates with the true organic return probability at 0.21. A cadence-aware churn score correlates at 0.87. Return rate by churn-score decile runs monotonically from 69.9% down to 46.4%, a 24-point spread. By recency decile it runs 54.9%, 56.5%, 53.9%, 53.5%, 54.6%, 55.2%, 53.0%, 55.1%, 53.4%, 48.6%. Essentially flat, with no ordering.
The reason is cadence. Every player has a natural rhythm, and a single recency threshold sweeps in the slow ones first:
| Player’s normal cadence | Players | Mean days since seen | Return rate |
|---|---|---|---|
| 1–5 days | 1,375 | 10.2 | 58.1% |
| 6–10 days | 4,831 | 15.2 | 56.8% |
| 11–30 days | 7,152 | 26.7 | 51.8% |
| 30+ days | 642 | 55.0 | 45.5% |
A committed fortnightly player looks identical to a churner at day 12.
Then the outcome definition. Only 22% of “returners” are durable, meaning five or more active days. In Altered Carbon, being re-sleeved into a body for an afternoon’s testimony does not make you alive again in any sense that matters; the campaign paid for the other 78% on exactly those terms. They were spun up for an event and stacked again afterwards.
The deeper point is that risk is not the same as treatability. The players most likely to churn are often the least persuadable, because they are genuinely gone. The at-risk pool is diluted by sure-things who return anyway and lost-causes who never will, and both have roughly zero uplift. Survival models answer risk; uplift models answer treatability, and treatability is the targeting objective most teams skip.
Pitfall two: interference
The assumption is that one player’s treatment does not affect another’s outcome. In an alliance game that is false by design. The nudge lifts returns for reachable persuadables; a returning player rallies their alliance; the cascade pulls alliance-mates back, including control-arm players, because under per-player randomization treatment and control share alliances.
Same players, two designs:
| Design | Control | Treatment | Gap | p |
|---|---|---|---|---|
| Per-player | 52.52% | 55.18% | +2.65pp | 0.00164 |
| Alliance cluster | 43.24% | 58.12% | +14.88pp | 1.23e-05 |
The per-player design understates the effect 5.6×. Against the recorded truth of 12.67pp, the per-player estimate recovers 21% of it. The alliance estimate lands slightly above the truth, because it measures the total effect, direct plus spillover, rather than the direct one.
Neither is wrong. They are not two estimates of the same thing. The choice of unit is a choice of estimand: per-player gives you the direct effect, cluster gives you the total effect, and the studio is paying for the total. Match the unit to the interference graph, not to convenience. And note the asymmetry that follows: a per-player null under social interference is ambiguous, not proof the nudge failed.
Pitfall three: attribution without a counterfactual
The assumption is that the post-campaign behaviour of treated players tells you what the campaign did. Here is the chain of individually reasonable steps that makes a negative result impossible:
- Define the population as lapsed players.
- Send the push to all of them.
- Measure who returned within seven days.
- Attribute the returners’ subsequent spend to the campaign.
Every step is defensible alone. Together they measure the natural return rate of lapsed players and then value it at the spend rate of the people who chose to come back. It is Ozymandias at the end of Watchmen, explaining that he did it thirty-five minutes ago: the conclusion was reached before the analysis began, and the analysis was arranged to arrive there.
“Incremental” ARPU per player, by method:
| Method | Value |
|---|---|
| All post-revenue counted as incremental (treated) | $2.010 |
| The same method applied to control | $1.812 |
| Single-arm pre → post (treated) | $1.828 |
| Randomized control (T − C) | $+0.198 |
The naive method credits the untreated control arm, an arm that received nothing, with $1.81 per player. Control’s own revenue moved from $0.178 to $1.812 with no treatment at all. That is regression to the mean, on a population selected at a transient low, and it is the engine of every win-back deck that has ever been funded without a control.
The valid estimate is p = 0.26. Not significant.
Pitfall four: small samples and whale fragility
The assumption is that the sample is large. It is large on paper. A live campaign is far smaller than 14,000 after the reachability funnel, and the per-player effect is real but small:
| n | Gap | 95% CI | Verdict |
|---|---|---|---|
| 14,000 | +0.027 | [+0.010, +0.043] | significant |
| 6,000 | +0.035 | [+0.010, +0.061] | significant |
| 3,000 | +0.036 | [−0.000, +0.071] | reads as null |
| 1,500 | +0.025 | [−0.025, +0.076] | reads as null |
| 800 | +0.060 | [−0.009, +0.129] | reads as null |
Revenue is worse. At n = 1,500 the ARPU lift reads $+0.55 with a 95% interval of [−$0.3, +$1.4]; drop the single largest spender, worth about $180, and it falls to $+0.30.
An underpowered test does not return “no effect”. It returns “no information”, and the two get written up identically.
Where the data had its own opinion
Two things the fixture showed me that I had not fully priced in.
What the cluster structure actually costs. Estimated on the alliance design: mean cluster size 30.4, ICC 0.52, design effect 16.3×. Fourteen thousand enrolled players carry about 858 players’ worth of independent information. A player-level power calculation would overstate this experiment’s precision by more than an order of magnitude, which is what happens whenever the alliance is treated as a segment rather than as the unit.
The reachability decomposition. Reachability is a pre-treatment property, so conditioning on it is legitimate, unlike conditioning on whether someone opened the message, which is post-treatment and a collider.
| Frame | Gap | p |
|---|---|---|
| Everyone targeted (ITT) | +2.65pp | 0.00164 |
| Unreachable subgroup (placebo) | −1.14pp | 0.287 |
| Within the reachable frame | +8.92pp | 5.59e-11 |
| Complier effect (CACE) | +7.06pp | — |
The unreachable subgroup moved by nothing, which is exactly what a clean placebo looks like, and it is the check I would want on any campaign readout before believing the rest. The reachable-frame effect then has to be discounted back to the population: a lift on the reachable share moves the whole base by roughly that share times the lift. Decide on the population number or overstate impact several-fold.
Per channel, within the reachable frame: push +11.8pp, attribution-linked +9.9pp, email +8.1pp, paid social +5.9pp.
What the estimates actually say
Cannot say: that the campaign generated the post-campaign revenue of the treated cohort. That method credits the control arm just as generously.
Can say: with alliance-cluster randomization the campaign lifts durable returns by +14.9pp (p = 1.2e-05), consistent with the recorded truth once spillover is counted.
Cannot yet say anything about revenue: the spend read is whale-dominated and underpowered at realistic campaign sizes.
Which is a more interesting outcome than either “it works” or “it doesn’t”. The campaign probably does work, on the one metric that was measured honestly, and the deck that got it funded was wrong about everything except the conclusion.
What I would change next time
In the order it changes the answer:
- Anchor metrics on the campaign rather than the install, and score durable return, not any return.
- Target with an uplift model against a cadence-aware churn score, not a recency threshold. Recency is nearly uninformative here; cadence is nearly everything.
- Randomize on the interference graph and power for clusters. The alliance is the unit. The 858-player effective sample is the budget, not the 14,000.
- Freeze the reachable frame before launch, report ITT alongside the complier effect, and check the unreachable placebo every time.
- Keep a permanent never-messaged hold-back as the standing baseline, so regression to the mean has somewhere to show itself other than in the treatment arm’s favour.
The broader lesson
Four independent mistakes on one dataset, each sufficient on its own to get the decision wrong, and all four pushing the same way. That last part is not a coincidence. Every one of them is a shortcut that happens to flatter the programme, which is why they survive: nobody audits a win. Together they are how win-back programmes get funded on nothing, and, occasionally, how a programme that actually works gets funded for the wrong reasons and cancelled the first time someone measures it properly.
Rules this demonstrates: 1 — define the decision first · 3 — verify assignment · 4 — count whom you assigned, not whom the treatment selected · 11 — calculate power before launching · 17 — do not mistake conviction for evidence
Worked examples
Reengagement pitfalls — four ways to fool yourself in a win-back campaign
All four pitfalls with their code, plus the cluster-variance and reachability decompositions.