Game analytics · Part 04 of 10 · August 23, 2026

Whale KPI Distortion

Simulated data with a known answer. The true effect is exactly zero, set that way on purpose, and the analysis is scored against it at the end.


The question

A treatment shows a +50% ARPU lift. Ship it?

That is the question as it arrives. The one that actually decides it, and the one I built this fixture to answer, is narrower: on zero-inflated, whale-skewed revenue, which statistic is allowed to determine whether an experiment won? Because if the answer is “the mean”, then the answer to the first question is being decided by whichever arm the biggest spender happened to land in, and the experiment is a coin toss with a spreadsheet attached.

Why the effect is harder to see than it looks

Free-to-play revenue is not a bell curve with a long tail. It is a spike at zero and a tail so heavy that the tail is the distribution. Most players never pay. A few pay a little. A handful pay amounts that would be a rounding error on the studio’s annual books and are a decisive fraction of any one experiment arm.

The consequence for the arithmetic is that the sample mean stops behaving like an estimate of the typical player and starts behaving like an estimate of the top few rows. Vought International, in The Boys, has an average employee who is extraordinarily strong, provided you include Homelander in the average. The median Vought employee works in marketing. Both statements are true, and only one of them tells you anything about who you are going to meet in the lift.

The statistical version is that the mean of a heavy-tailed variable converges slowly, its standard error is dominated by the same tail, and the t-test’s comforting normality applies to the sampling distribution of that mean at sample sizes a live game will not reach. So I wanted a fixture where the mechanism was undeniable: a treatment with a known effect of zero and a headline of +50%, produced by one row.

My approach

Player-level randomization, two arms, 8,000 players each. Both arms are drawn from the same data-generating process, so retention, conversion, session activity and the entire spend distribution are identical by construction. The seed was chosen by a balance search precisely so that every non-spend KPI comes out statistically equal. The fixture is built to make the failure vivid rather than to hide it.

Then one thing changes. The treatment arm’s largest existing whale is promoted to megawhale, with a spend value solved algebraically to hit a +50% ARPU lift exactly:

megawhale = 1.50 × control_total − treatment_total + victim_original

That single row is the entire difference between the arms. The whale’s spend was not sampled. It was solved for. This is the one place in the series where I put my thumb on the scale, and I am telling you where.

The realized distribution is ordinary for a free-to-play game:

Tier Players % of players % of revenue
non-spender 14,036 88.0 0.0
minnow 1,581 10.0 13.0
dolphin 259 2.0 16.0
whale 123 1.0 51.0
megawhale 1 0.0 21.0

The assumptions doing the heavy lifting

There is really only one, and it is the one nobody states: that the mean is an appropriate summary of this variable. Everything downstream inherits it. The Welch test assumes it. The “+50%” assumes it. The deck assumes it. And it is false, not as a matter of taste but as a matter of arithmetic: 88% of players spend nothing and 0.775% of players generate 71.3% of all revenue, so the sample mean is a weighted readout of a handful of observations wearing the costume of a population statistic.

A second, quieter assumption is that the guardrails would catch anything wrong. They would not, and that is the more dangerous of the two, because passing guardrails make a bad headline more convincing rather than less.

What the readout says

ARPU moves $11.598 → $17.397, a +50.0% lift. And every guardrail is clean:

KPI Control Treatment p
D1 retention 0.4618 0.4632 0.849
D7 retention 0.2535 0.2504 0.649
D30 retention 0.1235 0.1234 0.981
conversion 0.1221 0.1234 0.810
session days 4.5625 4.5551 0.887

Engagement untouched, monetization up half. That is the deck that gets presented, and I understand why. It is a beautiful deck.

Where the data had its own opinion

One planted row worth $48,164.74, spread across an arm of 8,000, adds $6.02 per player to treatment ARPU, enough to move an $11.60 base by exactly the 50% it was engineered to move. That player is 34.6% of the treatment arm’s revenue and 20.8% of all revenue in the experiment, and spends 17.7× the next-biggest spender in either arm.

The tell is already in the table that produces the win. The variance moves with the point estimate. The same tail that inflates the mean inflates its standard error, so the +50% arrives with an interval that spans zero. ARPU’s point estimate is roughly fifty times the guardrail KPIs’ and its confidence interval is roughly forty times wider, while being no more significant. A large point estimate paired with a huge, non-significant interval is a tell, not a win. It is the statistical equivalent of a witness who is very confident and cannot remember what day it was.

Note also what the guardrails did here, which is nothing. Retention, conversion and engagement all passed. The failure was not a missing metric. It was the choice of statistic.

The robust battery

Run the same comparison four ways:

Statistic Control Treatment Relative lift p
Mean / ARPU 11.598 17.397 +50.0% 0.350
Trimmed mean (1%) 3.564 3.527 −1.0% —
Winsorized mean (p99) 5.686 5.650 −0.6% —
Median (spenders) 18.990 18.700 −1.5% 0.862
  • Welch t-test: t = +0.93, p = 0.35, 95% CI on the ARPU difference [−$6.37, +$17.97].
  • Bootstrap, 5,000 resamples: 95% CI [−$2.41, +$20.32].
  • Mann-Whitney U on the full spend distribution: p = 0.862. No distributional shift at all.

Welch’s t of 0.93 is one standard error of nothing.

Two reads that localize the damage

Decompose ARPU. Revenue per user is conversion × purchases per spender × mean ticket, and splitting it shows exactly where the distortion enters:

Factor Control Treatment Relative p
Conversion 0.1221 0.1234 +1.02% 0.810
Purchases / spender 5.704 6.871 +20.46% 0.143
Mean ticket $16.65 $20.52 +23.26% —
Median ticket $8.09 $7.86 −2.81% 0.454

The typical spender changed on no factor. The two mean-based factors each moved more than 20% and multiply into the headline. This is worth dwelling on, because the decomposition can look like corroboration, two independent signals pointing the same way, when in fact one player inflated frequency and ticket simultaneously, with 640 purchases at roughly $75 each. Two symptoms of one cause are not two pieces of evidence.

Leave out the top N. The cheapest robustness check there is:

Top-N excluded per arm Lift
0 +50.00%
1 +1.05%
2 +1.12%
3 +1.72%
5 +2.90%
8 +3.61%

The curve collapses at N = 1 and then stays flat. Remove one player from sixteen thousand, 0.00625% of the sample, and +50% becomes +1.05%.

What the estimates actually say

The true effect is zero by construction, and every outlier-resistant read agrees with the truth. The trimmed mean, the winsorized mean, the median, the rank test and the leave-one-out curve all say the same thing, which is nothing. Only the raw mean disagrees, and it disagrees by exactly the amount one row was solved to produce.

So there is nothing to ship. The naive ARPU read would have shipped a change worth exactly nothing on the strength of a number that was, in the most literal sense, one person.

What I would change next time

  • Pre-register the decision statistic. A trimmed or winsorized mean, or a capped ARPU with the cap fixed before launch. Choosing the statistic after seeing the data is choosing the answer.
  • Read the interval before the point estimate. Width first.
  • Pair the mean with a rank test and a bootstrap. Disagreement between them is the signal, not a nuisance to be explained away.
  • Decompose into conversion × frequency × ticket, and report the median alongside every mean.
  • Show the leave-out-top-N curve. A result that depends on one row is not a result.

The broader lesson

On zero-inflated, whale-skewed revenue, reporting the mean is a choice, not a measurement. It is a defensible choice for some questions (total revenue is, after all, a sum), but it has to be made before the data arrives and defended with a robust statistic alongside it, because otherwise the experiment is not measuring the treatment. It is measuring where Homelander sat.


Rules this demonstrates: 5 — one primary success metric · 8 — significance is not value · 14 — respect the shape of the distribution

Worked examples

ReportR Markdown · interactive

Whale KPI Distortion — when the mean lies

The full battery: robust statistics, the ARPU decomposition, and the leave-out-top-N curve, computed on the seeded fixture.

Open full report ↗