
Game analytics · Part 04 of 10 · August 23, 2026
Whale KPI Distortion
Simulated data with a known answer. The true effect is exactly zero, set that way on purpose, and the analysis is scored against it at the end.
The question
A treatment shows a +50% ARPU lift. Ship it?
That is the question as it arrives. The one that actually decides it, and the one I built this fixture to answer, is narrower: on zero-inflated, whale-skewed revenue, which statistic is allowed to determine whether an experiment won? Because if the answer is “the mean”, then the answer to the first question is being decided by whichever arm the biggest spender happened to land in, and the experiment is a coin toss with a spreadsheet attached.
Why the effect is harder to see than it looks
Free-to-play revenue is not a bell curve with a long tail. It is a spike at zero and a tail so heavy that the tail is the distribution. Most players never pay. A few pay a little. A handful pay amounts that would be a rounding error on the studio’s annual books and are a decisive fraction of any one experiment arm.
The consequence for the arithmetic is that the sample mean stops behaving like an estimate of the typical player and starts behaving like an estimate of the top few rows. Vought International, in The Boys, has an average employee who is extraordinarily strong, provided you include Homelander in the average. The median Vought employee works in marketing. Both statements are true, and only one of them tells you anything about who you are going to meet in the lift.
The statistical version is that the mean of a heavy-tailed variable converges slowly, its standard error is dominated by the same tail, and the t-test’s comforting normality applies to the sampling distribution of that mean at sample sizes a live game will not reach. So I wanted a fixture where the mechanism was undeniable: a treatment with a known effect of zero and a headline of +50%, produced by one row.
My approach
Player-level randomization, two arms, 8,000 players each. Both arms are drawn from the same data-generating process, so retention, conversion, session activity and the entire spend distribution are identical by construction. The seed was chosen by a balance search precisely so that every non-spend KPI comes out statistically equal. The fixture is built to make the failure vivid rather than to hide it.
Then one thing changes. The treatment arm’s largest existing whale is promoted
to megawhale, with a spend value solved algebraically to hit a +50% ARPU lift
exactly:
megawhale = 1.50 × control_total − treatment_total + victim_original
That single row is the entire difference between the arms. The whale’s spend was not sampled. It was solved for. This is the one place in the series where I put my thumb on the scale, and I am telling you where.
The realized distribution is ordinary for a free-to-play game:
| Tier | Players | % of players | % of revenue |
|---|---|---|---|
| non-spender | 14,036 | 88.0 | 0.0 |
| minnow | 1,581 | 10.0 | 13.0 |
| dolphin | 259 | 2.0 | 16.0 |
| whale | 123 | 1.0 | 51.0 |
| megawhale | 1 | 0.0 | 21.0 |
The assumptions doing the heavy lifting
There is really only one, and it is the one nobody states: that the mean is an appropriate summary of this variable. Everything downstream inherits it. The Welch test assumes it. The “+50%” assumes it. The deck assumes it. And it is false, not as a matter of taste but as a matter of arithmetic: 88% of players spend nothing and 0.775% of players generate 71.3% of all revenue, so the sample mean is a weighted readout of a handful of observations wearing the costume of a population statistic.
A second, quieter assumption is that the guardrails would catch anything wrong. They would not, and that is the more dangerous of the two, because passing guardrails make a bad headline more convincing rather than less.
What the readout says
ARPU moves $11.598 → $17.397, a +50.0% lift. And every guardrail is clean:
| KPI | Control | Treatment | p |
|---|---|---|---|
| D1 retention | 0.4618 | 0.4632 | 0.849 |
| D7 retention | 0.2535 | 0.2504 | 0.649 |
| D30 retention | 0.1235 | 0.1234 | 0.981 |
| conversion | 0.1221 | 0.1234 | 0.810 |
| session days | 4.5625 | 4.5551 | 0.887 |
Engagement untouched, monetization up half. That is the deck that gets presented, and I understand why. It is a beautiful deck.
Where the data had its own opinion
One planted row worth $48,164.74, spread across an arm of 8,000, adds $6.02 per player to treatment ARPU, enough to move an $11.60 base by exactly the 50% it was engineered to move. That player is 34.6% of the treatment arm’s revenue and 20.8% of all revenue in the experiment, and spends 17.7× the next-biggest spender in either arm.
The tell is already in the table that produces the win. The variance moves with the point estimate. The same tail that inflates the mean inflates its standard error, so the +50% arrives with an interval that spans zero. ARPU’s point estimate is roughly fifty times the guardrail KPIs’ and its confidence interval is roughly forty times wider, while being no more significant. A large point estimate paired with a huge, non-significant interval is a tell, not a win. It is the statistical equivalent of a witness who is very confident and cannot remember what day it was.
Note also what the guardrails did here, which is nothing. Retention, conversion and engagement all passed. The failure was not a missing metric. It was the choice of statistic.
The robust battery
Run the same comparison four ways:
| Statistic | Control | Treatment | Relative lift | p |
|---|---|---|---|---|
| Mean / ARPU | 11.598 | 17.397 | +50.0% | 0.350 |
| Trimmed mean (1%) | 3.564 | 3.527 | −1.0% | — |
| Winsorized mean (p99) | 5.686 | 5.650 | −0.6% | — |
| Median (spenders) | 18.990 | 18.700 | −1.5% | 0.862 |
- Welch t-test: t = +0.93, p = 0.35, 95% CI on the ARPU difference [−$6.37, +$17.97].
- Bootstrap, 5,000 resamples: 95% CI [−$2.41, +$20.32].
- Mann-Whitney U on the full spend distribution: p = 0.862. No distributional shift at all.
Welch’s t of 0.93 is one standard error of nothing.
Two reads that localize the damage
Decompose ARPU. Revenue per user is conversion × purchases per spender × mean ticket, and splitting it shows exactly where the distortion enters:
| Factor | Control | Treatment | Relative | p |
|---|---|---|---|---|
| Conversion | 0.1221 | 0.1234 | +1.02% | 0.810 |
| Purchases / spender | 5.704 | 6.871 | +20.46% | 0.143 |
| Mean ticket | $16.65 | $20.52 | +23.26% | — |
| Median ticket | $8.09 | $7.86 | −2.81% | 0.454 |
The typical spender changed on no factor. The two mean-based factors each moved more than 20% and multiply into the headline. This is worth dwelling on, because the decomposition can look like corroboration, two independent signals pointing the same way, when in fact one player inflated frequency and ticket simultaneously, with 640 purchases at roughly $75 each. Two symptoms of one cause are not two pieces of evidence.
Leave out the top N. The cheapest robustness check there is:
| Top-N excluded per arm | Lift |
|---|---|
| 0 | +50.00% |
| 1 | +1.05% |
| 2 | +1.12% |
| 3 | +1.72% |
| 5 | +2.90% |
| 8 | +3.61% |
The curve collapses at N = 1 and then stays flat. Remove one player from sixteen thousand, 0.00625% of the sample, and +50% becomes +1.05%.
What the estimates actually say
The true effect is zero by construction, and every outlier-resistant read agrees with the truth. The trimmed mean, the winsorized mean, the median, the rank test and the leave-one-out curve all say the same thing, which is nothing. Only the raw mean disagrees, and it disagrees by exactly the amount one row was solved to produce.
So there is nothing to ship. The naive ARPU read would have shipped a change worth exactly nothing on the strength of a number that was, in the most literal sense, one person.
What I would change next time
- Pre-register the decision statistic. A trimmed or winsorized mean, or a capped ARPU with the cap fixed before launch. Choosing the statistic after seeing the data is choosing the answer.
- Read the interval before the point estimate. Width first.
- Pair the mean with a rank test and a bootstrap. Disagreement between them is the signal, not a nuisance to be explained away.
- Decompose into conversion × frequency × ticket, and report the median alongside every mean.
- Show the leave-out-top-N curve. A result that depends on one row is not a result.
The broader lesson
On zero-inflated, whale-skewed revenue, reporting the mean is a choice, not a measurement. It is a defensible choice for some questions (total revenue is, after all, a sum), but it has to be made before the data arrives and defended with a robust statistic alongside it, because otherwise the experiment is not measuring the treatment. It is measuring where Homelander sat.
Rules this demonstrates: 5 — one primary success metric · 8 — significance is not value · 14 — respect the shape of the distribution
Worked examples
Whale KPI Distortion — when the mean lies
The full battery: robust statistics, the ARPU decomposition, and the leave-out-top-N curve, computed on the seeded fixture.