
Game analytics · Part 07 of 10 · August 23, 2026
The Sink That Booked Revenue It Never Kept
Simulated data. The registered hypothesis says so out loud: “simulated with a large refund-rate side effect, so the refund guardrail should breach.” Ground truth is opened once, at the end, to score the method.
The question
A live-service game wants to add a sink for premium currency: a compelling place to spend gems, on the theory that it will make players buy more gems. Does it raise revenue enough to ship, and does it do damage on the way?
The question interested me for two reasons that turned out to be one. First, the honest randomization unit is the server, and I wanted to see exactly what that costs when the game has 32 of them. Second, the harm in this fixture is a refund wave, and refunds are the kind of harm that a studio’s daily surfaces are built not to see. Those two facts collide in a specific way: the one metric with enough precision to decide the test is the one the fact table drops.
Why the effect is harder to see than it looks
A live economy is faucets and sinks. Weak sinks inflate: currency accumulates, the marginal utility of the next unit falls, a veteran eventually owns everything worth owning, and the store stops working because nothing in it is scarce. A soft-currency sink drains a stock the player already has. A premium-currency sink is a monetization feature dressed as an economy fix, because it converts a gameplay desire into a purchase.
Two things make the effect hard to see. The first is interference: a shard is one shared world, players trade and compete for the same alliance resources and face one aggregate money supply, so removing currency from circulation for half a shard is felt by the other half through prices. A within-shard comparison measures a contaminated difference. The shard is the smallest self-contained unit, and there are 32 of them.
The second is that the outcome that matters is not the one the dashboards show. Bookings are what a store reports. Net revenue is what the studio keeps. The gap between them is refunds, and refunds arrive later, through a different system, and get reconciled by finance rather than watched by live ops. A feature that converts through urgency and then gives the money back will look like a win on every surface the studio checks daily.
My approach
Fifteen control servers (5,623 players) against seventeen treatment servers (6,369 players), over 56 days. 32 randomization units, not 12,000 players. Almost everything below is a consequence of that one line.
Primary metric: net revenue per assigned player, standard errors clustered on server. Guardrails registered before launch, each with a threshold and a two-part breach rule: materially over threshold and distinguishable from noise. Refund rate is one of them, with a +0.50pp threshold.
SRM, at the grain the coin was flipped
| Grain | Control | Treatment | χ² p | Verdict |
|---|---|---|---|---|
| Servers (the randomization unit) | 15 | 17 | 0.724 | balanced |
| Players (the wrong grain) | 5,623 | 6,369 | 9.6e-12 | “catastrophic” |
The player-grain gap is 13.3%, which is exactly what 17 versus 15 servers of ~375 players produces. A split of 32 fair coin flips at least this uneven happens 86% of the time. The chi-square at the player grain treats 12,000 correlated observations as independent, so it detects the design, not a defect.
Wired to the player grain, this check fires a stop-the-test alarm on every single run. A check that always fires is one everyone learns to dismiss, and then it is worse than no check, because it has trained the room to ignore alarms.
The assumptions doing the heavy lifting
That the arms are comparable on things the shard determines. Player-level covariates balance beautifully: max standardised difference across tenure, pre-window spend, purchases and lifetime revenue is 0.018, an order of magnitude inside threshold. Country tier does not balance. Shards are named by region, and each region runs exactly two, so region is not a player attribute that happens to correlate with the cluster. It is a cluster attribute. Of sixteen regions, seven split one server to each arm, and nine sent both shards to the same arm.
| Country tier | Control share | Treatment share | Gap | Control ARPU |
|---|---|---|---|---|
| APAC | 13.2% | 23.5% | +10.3pp | $2.87 |
| Tier 1 | 53.4% | 59.5% | +6.1pp | $2.23 |
| ROW | 13.6% | 11.4% | −2.1pp | $1.99 |
| US | 19.9% | 5.6% | −14.3pp | $2.82 |
The arms hold different market mixes, and those markets are not worth the same. No amount of sample size fixes this, because the effective sample for a region-level covariate is 32, not 12,000.
That the fact table contains the outcome. This one I did not know I was making until it failed. The modelled fact table defines daily spend as revenue net of refunds. The gross amount and the refund flag exist only on the raw revenue rows and do not survive into the fact table. The evaluation pipeline reads the fact table. So an analyst working from the marts, which is to say the normal way of working, literally cannot see the two bases diverge. They see one revenue number and no refund flag. Getting at the finding meant querying the staging views directly.
In Chernobyl, the first dosimeter reads 3.6 roentgen, “not great, not terrible”, and the number goes up the chain and into the record, because 3.6 was the highest value that instrument could display. The fact table here is that dosimeter. It reports what it is built to report, faithfully, and the thing that matters is above its ceiling.
What the readout says
Per assigned player, standard errors clustered on server:
| Basis | Control | Treatment | Absolute | Relative | p |
|---|---|---|---|---|---|
| Gross bookings | $2.4733 | $2.5986 | +$0.1253 | +5.1% | 0.393 |
| Net revenue | $2.3969 | $2.3175 | −$0.0794 | −3.3% | 0.575 |
| Refunded USD | $0.0764 | $0.2811 | +$0.2047 | +268% | 0.0000 |
| Purchases / player | 0.2518 | 0.2591 | +0.0072 | +2.9% | 0.490 |
| Conversion rate | 0.1595 | 0.1547 | −0.0049 | −3.1% | 0.380 |
Booked $0.125 per player. Returned $0.205 per player. 1.63× the gain came back.
A store dashboard shows bookings. A live-ops readout shows bookings. This feature would have looked like a win in every surface a studio checks daily, for eight weeks, and the refunds would have arrived quietly in a finance reconciliation later.
Adjusting for the market-mix imbalance moves gross to +6.6% and net to −1.3%, shrinks the standard error by about 11%, and changes no conclusion, which is the useful thing to know about it.
Where the data had its own opinion
The guardrail that decides it
A refund rate is refunds per purchase, so it is estimated on the purchase rows, clustered on server: not on players, and not on the arm total, which would weight a whale’s twenty transactions the same as a minnow’s one.
Across 3,066 purchases: 2.82% → 9.76%, absolute +6.93pp, 95% CI [+4.91pp, +8.96pp], p = 1.9e-11, against a +0.50pp threshold. That is 13.9× the threshold, and the CI floor alone is 9.8× it.
Everything else passes. Crash rate +0.24pp (p = 0.14). D7 retention −0.63pp (0.46). D1 and D14 retention drift just past the stated band at −1.3pp and −1.4pp but neither is resolvable, so both log as watch rather than breach. A one-part rule keyed on the threshold alone would stop the test on two metrics it cannot measure; a one-part rule keyed on significance alone would let a large, imprecise harm through. Both parts, or the rule is decorative.
This is not a feature that broke the game or drove players away. Players kept playing. They just kept asking for their money back, and nothing except the refund guardrail was looking.
Achieved precision, which belongs on every readout
Replace the asserted power claim with the minimum effect the test could actually detect, 2.80 × SE:
| KPI | Relative effect | Achieved MDE | Observed / MDE |
|---|---|---|---|
| Gross bookings / player | +5.1% | 16.6% | 0.31× |
| Net revenue / player | −3.3% | 16.5% | 0.20× |
| Purchases / player | +2.9% | 11.7% | 0.25× |
| Conversion rate | −3.1% | 9.7% | 0.31× |
| Refund rate / purchase | +245% | 102% | 2.40× |
The test registered a +9% effect and achieved a 16.5% MDE. It could only resolve an effect 1.8× larger than the one it was built to find. The observed net effect sits at 0.20× its MDE, which is not a null result. It is no result.
Servers per arm needed at 80% power: about 54 for a 9% effect, 174 for 5%, and 482 for 3%. The game has 32 shards in total. The honest power calculation for the primary metric was never “how long should we run this” but “this game cannot run this test”, and that is a useful thing to know on day zero.
So why did the guardrail escape the same fate? Refunds are a property of purchases, and there are 3,066 of them. Rates over events survive the cluster penalty far better than per-player means. The generalizable question for any coarse-unit test: which of your metrics have a within-cluster denominator? Those are the ones the test can actually resolve.
What clustering did, in both directions
| Outcome | Naive SE | Cluster SE | Ratio | ICC | Design effect | Effective n |
|---|---|---|---|---|---|---|
| Net revenue | 0.1547 | 0.1417 | 0.92× | 0.000 | 1.00 | 11,992 |
| Gross | 0.1622 | 0.1465 | 0.90× | 0.000 | 1.00 | 11,992 |
| Refund rate | 0.0085 | 0.0103 | 1.21× | 0.024 | 3.31 | 925 |
Same 32 clusters, opposite directions. Player spend is dominated by individual heterogeneity (one whale swamps a shard’s average) so almost none of its variance is between-server and the ICC is effectively zero. Refunds are the opposite: treatment is assigned at server level and moves refunds hard, so refund propensity genuinely clusters, and 3,066 purchase rows collapse to an effective 925.
The diagnostic worth internalising: an outcome the treatment moves strongly will tend to show a real ICC precisely because the treatment is clustered. The metric the test is about is usually the one that most needs cluster-robust inference, and the one where skipping it overstates significance.
A note on the estimator, because it bit me. The ICC here is a closed-form one-way random-effects moment estimator, deliberately. A mixed-model fit is unstable on an outcome this heavy-tailed: the likelihood is near-flat in the between-cluster variance, the optimizer lands on the boundary, and it reports an ICC of 0.70 on an outcome whose cluster means barely differ. The moment estimator is closed-form and cannot do that. Sophisticated is not the same as stable.
Where the refunds came from
Three candidate stories, each with a signature. A billing or entitlement defect concentrates in one SKU. Big-ticket regret concentrates in the top price band. Impulse conversion (the sink creates urgency, urgency converts marginal buyers, marginal buyers refund) is roughly multiplicative across the catalogue.
| Ticket | Control | Treatment | Ratio | n |
|---|---|---|---|---|
| < $5 | 2.17% | 6.19% | 2.86× | 1,200 |
| $5–15 | 2.81% | 12.68% | 4.51× | 1,342 |
| $15–40 | 4.62% | 10.33% | 2.24× | 466 |
| $40+ | 3.70% | 12.90% | 3.48× | 58 |
Elevated in every band, 2.2× to 4.5×, concentrated nowhere. Impulse conversion.
One SKU refunds at exactly 0.0% in both arms across 307 purchases. That is not a finding about the SKU. Those rows come from a layer the simulator writes as unrefundable by construction, which makes them an accidental placebo: transactions in the same window, for the same players, that cannot respond, and they read flat as they must. They also quietly dilute the headline, which is why the control refund rate reads 2.8% against a 4.0% base rate in the generator. Check what is mechanically incapable of moving before quoting a rate.
The targeting escape hatch does not exist
Prior spenders are 55% of the population and 77% of gross bookings, and they carry the booked gain: +11.6% gross, against −12.8% for players with no purchase history.
The instinct on a guardrail breach is to ship to the segment where it works. But the refund rate rises in both segments, 2.2% → 9.8% among prior spenders and 4.3% → 9.5% among the rest. Restricting the launch would keep the harm and shrink the population it is spread over. Worth checking before proposing it, and worth reporting when it fails.
What the estimates actually say
Do not ship.
The verdict does not rest on the primary KPI, and it cannot. Net revenue came in at −3.3% against a 16.5% MDE, so on the registered metric the honest answer is inconclusive. The refund guardrail decides it, unambiguously.
An underpowered test is not this feature’s main problem. A 3.5× refund rate is, and that is a product problem, not a measurement one. A sink that converts through urgency and returns the money is not a monetization feature. It is a loan, extended on terms Tony Soprano would have considered soft, and unlike his it comes with a chargeback rate the storefront can see. Sustained refund volume puts merchant standing at risk, a cost that appears in no revenue KPI.
The economy problem, currency accumulation, is solved by any sink. Only the monetization framing needed it priced in premium currency, and the monetization framing is the part that failed.
What I would change next time
Block on region. Every region runs exactly two shards, a natural matched pair. Assign one of each pair to each arm and region balance becomes exact by construction, with the comparison happening within region, where the between-market variance that dominates the noise cancels.
The seven regions that happened to split 1–1 put a number on it:
| Sample | Estimator | Relative | SE | p | Achieved MDE |
|---|---|---|---|---|---|
| All 32 servers | unblocked | −3.3% | 0.1417 | 0.575 | 16.6% |
| 14 paired servers | unblocked | +8.1% | 0.1631 | 0.289 | 21.4% |
| 14 paired servers | region fixed effects | +8.6% | 0.0825 | 0.026 | 10.8% |
Fourteen blocked servers beat thirty-two unblocked ones. The standard error falls 42%, on fewer than half the clusters, and the interval finally excludes zero.
The point estimate is not the lesson. The blocked subset reads +8.6% against a simulated truth of +5.5%; it overshoots, because fourteen clusters is still small and those seven regions are a particular seven. The transferable claim is about precision. Post-hoc regression adjustment recovered about 11% of the standard error; blocking recovers 42%, and blocking is a decision made at randomization time, not at analysis time. In a shard-based game the pairing is already in the world layout, so it is close to free, and I did not take it.
And, beyond the design: carry gross and the refund flag into the fact table. The finding here was only reachable by going around the pipeline. That is a pipeline defect, and the experiment module exists partly because of it.
Checking the answer
Ground truth is opened once, at the end.
| Quantity | Truth | Estimated | Covered by 95% CI? |
|---|---|---|---|
| Gross revenue effect | +5.48% | +5.06% | yes |
| Refund rate multiplier | 2.83× | 3.45× | — |
The gross estimate lands 0.4pp from the truth and its interval covers it. The point estimate was fine all along. What the design lacked was the precision to say so: the interval is 23 percentage points wide. The correctness here is luck rather than evidence, which is exactly why an achieved-precision statement belongs next to every effect. The number and its usefulness are separate facts.
The broader lesson
When the randomization unit is coarse, the test can only resolve the metrics that have a denominator inside the cluster, and those are rarely the ones on the front of the deck. The refund rate decided this test because there were 3,066 purchases to count it over; net revenue could not, because there were 32 servers. Know which of your metrics the design can actually see before you launch it, and make sure the pipeline carries them, because a fact table that drops the refund flag is a dosimeter that tops out at 3.6.
Rules this demonstrates: 2 — randomize where interference occurs · 3 — verify assignment · 6 — guardrails · 11 — calculate power before launching · 12 — analyze at the level of randomization