An A/B test runs on a shared server: control and
treatment are randomized against each other on
shard_01/02. A third group, holdout, is an
equivalent population on a different server
(shard_08/09) that is not in the
experiment — the out-of-sample guardrail, i.e. the true, uncontaminated
baseline.
The treatment feature has no direct effect on spend, but it creates a negative externality on the shared server: treatment players gain an edge in the shared economy that depresses the control players who share that server (interference / SUTVA violation). A naïve treatment-vs-control read shows treatment “winning” — when in truth treatment is merely at baseline and control was harmed.
This report: ① get the data → ② confirm engagement is equal (in-experiment guardrails pass) → ③ run the naïve read → ④ bring in the out-of-sample holdout → ⑤ visualize → ⑥ make the call. The lesson: compare each arm to an out-of-sample holdout, not just the arms to each other.
| variant | players | servers | in_experiment | true_state | arpu_d1 |
|---|---|---|---|---|---|
| control | 10000 | shard_01, shard_02 | True | harmed_by_externality | 0.779 |
| treatment | 10000 | shard_01, shard_02 | True | baseline | 1.076 |
| holdout | 10000 | shard_08, shard_09 | False | baseline | 1.077 |
Retention and session activity are statistically indistinguishable across all three groups. A team watching the usual guardrails sees nothing wrong.
| kpi | control | treatment | holdout | p_treatment_vs_control |
|---|---|---|---|---|
| retained_d1 | 0.5211 | 0.5165 | 0.5142 | 0.5242 |
| retained_d7 | 0.2192 | 0.2193 | 0.2141 | 1.0000 |
| sessions_d1 | 3.3549 | 3.3824 | 3.3874 | 0.2101 |
Day-1 ARPU is much higher in treatment — large and significant, the textbook “ship the winner.”
## treatment ARPU_d1 = $1.076 control ARPU_d1 = $0.779
## lift = +38.1% Welch p = 1.41e-06 -> SIGNIFICANT: looks like a winner, ship it
Compare each arm to the holdout baseline. Treatment is not actually above baseline — its “lift” is zero. Control is below baseline. The entire treatment-vs-control gap is control being harmed, not treatment doing anything.
| comparison | ref_mean | test_mean | rel_lift | p_value | reads_as |
|---|---|---|---|---|---|
| treatment vs control | 0.7792 | 1.0763 | 0.3814 | 0.0000 | the apparent win |
| treatment vs holdout | 1.0767 | 1.0763 | -0.0004 | 0.9956 | no real lift (treatment == baseline) |
| control vs holdout | 1.0767 | 0.7792 | -0.2764 | 0.0000 | control was harmed |
Left — day-1 ARPU by group with the out-of-sample holdout baseline: treatment sits on the baseline, control sits well below it. Right — the measured “lift” depends entirely on the reference: against control it’s +38%; against the true baseline, treatment is flat and control is −28%.
Engagement is equal, and the naïve treatment-vs-control read shows a large, significant +38% day-1 ARPU “win.” But against the out-of-sample holdout, treatment is exactly at baseline (no real lift) and control is −28% (harmed). The gap is control regression from a negative externality on the shared server, not a treatment effect.
Decision: do not ship / investigate. The experiment is contaminated — control is not a clean counterfactual. Without the different-server holdout you would have shipped a change that does nothing (or, if the externality scales with rollout, harms everyone). Keep an out-of-sample / global holdout as a guardrail and compare each arm to it, not just the arms to each other.
Generated by the Savepoint Analytics video-game A/B testing case
study. Companion to data/mocks/ — a teaching fixture on why
out-of-sample guardrails matter.