A live-service sale offered high-value specialty packs at a discount. The treatment arm got higher per-event purchase caps on those packs; control kept the standard, tighter caps. The question is not whether sale-weekend revenue rose — it always does — but whether the higher caps created net revenue or merely pulled spend forward in time (front-loading), at worst cannibalizing later spend.

This report ① gets the data → ② assesses its distribution → ③ runs the statistical tests → ④ builds a variant performance table → ⑤ visualizes how the revenue gap rose during the sale then fell back to zero. All data is simulated; revenue is measured over a full window that extends past the promo — the only way front-loading shows.

1. Get the data

Field Value
Test exp_specialty_pack_caps
Randomization unit user
Primary KPI revenue_per_player
Sale weekend 2025-11-28 → 2025-11-30
Full window 2025-11-28 → 2025-12-15
Assigned (control / treatment) 3668 / 3642

Hypothesis. Raising the per-event purchase caps on high-value specialty packs during a sale weekend increases event revenue. The real question is whether higher caps grow net spend or merely pull it forward — front-loading, and at worst cannibalization, rather than incremental revenue.

2. Assess the data & its distribution

Game revenue is zero-inflated and heavy-tailed — most players spend nothing and a few “whales” dominate — so a naive t-test alone is not enough; the tests below add Mann-Whitney and a bootstrap. If the sale only shifted timing, the two arms’ full-window spend distributions should be near-identical.

Full-window spend per assigned player, by arm
arm n mean median sd skew pct_zero p99 max top1pct_rev_share
control 3668 28.007 0 53.399 3.626 0.568 246.561 693.98 0.120
treatment 3642 27.685 0 54.476 3.689 0.573 254.933 706.11 0.127

3. Statistical tests

SRM first (assignment sanity), then the revenue comparison on three horizons — sale weekend, full window, post-sale tail — each with Welch’s t (primary), Mann-Whitney U, and a bootstrap CI on the difference in means.

## SRM chi-square: X2 = 0.09, p = 0.761 -> PASS  (control = 3668, treatment = 3642)
## Conversion (any spend, full window): control 43.2% vs treatment 42.7%  (p = 0.691)
Revenue per player by horizon — Welch, Mann-Whitney, and bootstrap
window control treatment abs_lift rel_lift ci_low ci_high welch_p mannwhitney_p boot_low boot_high significant
Sale weekend (Fri–Sun) 5.4155 6.6671 1.2516 0.2311 0.2398 2.2634 0.0153 0.5332 0.2773 2.2929 TRUE
Full window (18 days) 28.0071 27.6848 -0.3223 -0.0115 -2.7959 2.1513 0.7984 0.5159 -2.7294 2.0578 FALSE
Post-sale tail 22.5916 21.0177 -1.5739 -0.0697 -3.7665 0.6186 0.1594 0.2891 -3.8282 0.5708 FALSE

Read the table top-to-bottom. The sale weekend shows a large, significant revenue lift (Welch and the bootstrap CI agree; Mann-Whitney is null because most players never spend, so the median is $0 in both arms — for a revenue mean the mean-based tests are decision-relevant). Over the full window that lift is gone — a fraction of a percent, indistinguishable from zero. The post-sale tail is negative: control out-spends treatment after the sale, because control’s budget was never pulled forward. That is cannibalization — the sale moved spend in time without creating any.

4. Variant performance table

Per-variant performance
metric control treatment
Assigned players 3,668 3,642
Conversion rate 43.2% 42.7%
Rev / player — full window $28.01 $27.68
Rev / player — sale weekend $5.42 $6.67
Rev / player — post-sale tail $22.59 $21.02
Headline: significant sale-weekend lift vs. null full-window effect
window control treatment abs lift rel lift welch p verdict
Sale weekend (Fri–Sun) $5.42 $6.67 $1.25 23.1% 0.015 SIGNIFICANT
Full window (18 days) $28.01 $27.68 -$0.32 -1.2% 0.798 not significant

5. Performance visual — revenue rose during the sale, then fell back

Conclusion

The higher purchase caps produced a large, significant sale-weekend revenue lift that fully washed out over the full window — the cumulative gap rose then returned to ~$0, the per-player spend distributions are near-identical, and the post-sale tail even favors control. The sale front-loaded revenue rather than creating it; with net incrementality at zero this is textbook cannibalization. A ship decision keyed to the sale-weekend number alone would be measuring timing, not value — the full-window horizon is the one that matters.


Generated by the Savepoint Analytics video-game A/B testing case study. Revenue is read from the dbt marts; the analysis unit is the player; effects are reported with confidence intervals over a horizon that extends past the promo.