
Series · Game analytics · 10 parts
A/B Testing in Games
A decade of running experiments on live games, reduced to seventeen rules anyone can apply, a worked case study for each way an experiment lies to you, and the module I built so the pipeline refuses to make the same mistakes twice.
Parts
A/B Testing in Games
A decade of running experiments on live games, reduced to seventeen rules anyone can apply, a worked case study for each way an experiment lies to you, and the module I built so the pipeline refuses to make the same mistakes twice.
Rules for Reading an Experiment
Seventeen standing rules for running and reading A/B tests on a live game, written so a designer, producer, or monetization lead can interrogate a result without reproducing the analysis. Each rule links to the case study where ignoring it went wrong.
The Experiment Module
A self-contained experimentation layer: assignment and exposure logging, guardrails built into the modelling layer, primary metrics on a separate compute path, and the post-experiment deep dives. What it enforces, what it refuses to do, and where it still says prototype.
Promotional Offers and Borrowed Revenue
Raising the per-event purchase cap on a discounted pack lifts sale-weekend revenue 23%, significantly. Over the full window it nets to −1.2%. The same money, spent earlier, and a case study in choosing the horizon before the result chooses it for you.
Includes 1 interactive reportWhale KPI Distortion
A +50% ARPU win, produced entirely by one planted player out of sixteen thousand. Every guardrail passes, the t-test is not significant, and every outlier-resistant statistic says zero. On whale-skewed revenue, the mean is a choice, not a measurement.
Includes 1 interactive reportRe-engagement, Honestly Measured
Four independent mistakes on one win-back campaign, each sufficient on its own to get the decision wrong, and all four pushing toward the same confident wrong call. Same players, two randomization designs, +2.7pp versus +14.9pp, and both of them right.
Includes 1 interactive reportHoldouts and Interference
A feature with no direct effect on spend wins its test by +38%, because the edge it grants in a shared economy depresses the control players on the same server. One dataset, three answers, and only one of them visible from inside the experiment.
Includes 1 interactive reportThe Sink That Booked Revenue It Never Kept
A premium-currency sink, randomized by server. Gross bookings up 5%, net revenue down 3%, refund rate up from 2.8% to 9.8%, and the modelled fact table cannot see any of it, because it drops the gross amount and the refund flag on the way in.
Content Demand Substitution
Cut one unit's cost by 25% and its share of player investment rises 8 percentage points. Three quarters of that comes from units doing the same job. The lift is not the decision; the diversion is, and the headline cannot tell the two apart.
Includes 1 interactive reportPrice and Income Elasticity
Three demand parameters recovered from a game economy where no dollar price was ever randomized, by randomizing the prices the studio sets and converting through the chain. The one experiment in the series that leaves behind a number you can reuse.
Includes 1 interactive reportMatchmaking, Rewards, and Dose
Designed opponents in a thin matchmaking tier: whether players prefer the harder content, how much reward premium a unit of difficulty costs, and three ways to estimate the same thing (a dose design, a difference-in-differences, and a server-randomized holdout) that disagree for instructive reasons.