Game analytics · Part 08 of 10 · August 23, 2026

Content Demand Substitution

Simulated data with recorded structural parameters. The analysis never reads them; a coda at the end scores the recovery.


The question

In a base-builder, players own a roster of units and pour a limited number of investment occasions into a few of them. The roadmap question is not “how much did players spend” but which unit gets the investment, and what happens to every other unit when we change one of them.

If you discount, buff, or add an equipment slot to one unit, does that grow total investment, or just move it off its neighbours? The question interested me because every unit-balance readout I have read answers the first half and declares victory, and the second half is the one the roadmap actually depends on. A discount that grows the pie and a discount that re-slices it produce the same headline.

Why the effect is harder to see than it looks

Think of an XCOM squad. There are six slots, and no discount on rangers makes it seven. Cheapen one class and you get a different squad, not a bigger one. What you observe, a rise in the discounted unit’s uptake, is real, and it tells you almost nothing about whether the change was good, because the interesting quantity is who got left at the barracks.

A unit roster has the same structure. Investment occasions are scarce, every unit competes for them, and a treatment applied to one unit shows up as a positive on that unit and a negative spread across the others. The naive readout sees the positive. It has no column for the negatives. So the effect that matters, the diversion, is invisible unless you build the estimand to show it, and the estimand most teams build (spend on the featured unit) is constructed to hide it.

There is a second difficulty. Comparing investment across units means attributing currency to units, and revenue-attribution systems are exactly the kind of infrastructure that is never quite finished. I wanted an estimand that did not need one.

My approach

The estimand: a share, not a level

The outcome per player is

share_target = occasions spent on the target ÷ all of that player's occasions

with every other owned unit bucketed into same-role (substitutes), complement (units linked by the synergy graph), or other-role.

That choice buys three things. It removes the need for any revenue-attribution system: no currency has to be assigned to a unit for the comparison to be valid. It is bounded, so one heavy investor cannot drag the mean, which after the whale case study I was not going to leave to chance. And it sits cleanly inside a two-stage budgeting story: stage one decides whether to invest today, stage two allocates across the roster, and the estimand lives entirely in stage two.

The sample is restricted to players who own the target and invested at least once. Ownership is fixed before the window opens, so this is not post-treatment conditioning, and a player who cannot invest in the target cannot possibly respond to its price. Of 9,756 players assigned to the cost ladder, 2,675 qualify. The analysis discards roughly 72% of the assignment by design, before any effect is measured, and says so.

Three tests, three levers

Cost ladder Equipment slot Power buff
Randomization unit user user server
Arms ×0.75 / ×1.00 / ×1.25 control / extra slot control / power ×1.25
Lever budget more ways to invest value

The unit changes for the power buff because a per-player balance change in a PvP game is unfair and detectable, so the shard is the smallest defensible unit. The fairness constraint, not the statistics, chooses the randomization unit there.

SRM on the cost ladder: p = 0.077, pass.

The assumptions doing the heavy lifting

The share is the right summary. It is, for “which unit”. The cost is real: a share is bounded and arithmetically coupled, so a bucket that merely holds its ground still loses share. Levels have to be read alongside it, and I do below.

Ownership is pre-treatment. It is, by construction of the window. In a live game, a discount can induce acquisition of the discounted unit, and then the “owners” sample is selected by the treatment. The window here opens after ownership is fixed. In production that would need checking, not assuming.

Substitution is nested. The bucket structure assumes units doing the same job are closer substitutes than units doing different jobs. That is a modelling assumption about the roster, and it turns out to be the finding: the experiment measures how nested substitution actually is, and the answer is “very”.

What the readout says

The cost ladder

Arm Cost Players Target share Change 95% CI p
cost −25% ×0.75 903 0.2364 +0.0806 [+0.057, +0.104] ~0
control ×1.00 923 0.1559 — — —
cost +25% ×1.25 849 0.1241 −0.0317 [−0.052, −0.011] 0.0027

Arc elasticities: −1.45 on the cut, −1.02 on the rise.

The response is symmetric, in both directions, and that matters more than either number. A one-sided test could not separate a real demand response from a novelty effect; players will try anything new, including a unit that got more expensive. Symmetry is the cheapest validity check available on a price test, and the third arm is what buys it.

Diversion, which is the actual finding

Where did the 8 points come from?

Bucket Share change Share of the target’s gain
Target +8.06pp —
Same-role (substitutes) −6.21pp 77.0%
Complements −1.09pp 13.5%
Other roles −0.76pp 9.4%

Roughly three quarters of the gain came from units doing the same job. Under a tenth came from other roles. The discount reshuffled attention inside one slot of the roster rather than changing how players play. The squad is the same size; a different ranger is in it.

That is a completely different decision from a cut that pulls investment in from elsewhere in the game, and the headline lift cannot distinguish them. On a roster, the question is never “did demand go up” but “whose demand went down”.

Where the data had its own opinion

Levels, alongside shares. This is what keeps the share estimand honest:

Bucket Level, control Level, cut Level change
Target 0.571 0.869 +52.3%
Same-role 0.768 0.638 −17.0%
Complement 0.181 0.123 −32.1%
Other roles 2.041 2.094 +2.6%

The complement bucket falls in both share and level, which looks like evidence against complementarity and is not. The units are genuinely linked; investing in one raises the other’s utility. But every unit competes for the same scarce occasions, and the target just got cheaper. Net complements can still be gross substitutes. The bucket that actually holds its level while losing share is other roles, and that is the case the shares-versus-levels caveat is really about. Read only the share table and you would conclude the whole roster suffered. It did not; one slot did.

Who responds. Strategy is latent, so it is inferred from each player’s pre-window investment mix, a legitimate pre-treatment covariate. The label agrees with the simulator’s true archetype for 78.5% of players with a dominant strategy, which is a rare chance to score a segmentation heuristic against truth.

Pre-period focus Control Cost −25% Lift Treated n
support 0.1235 0.3059 +0.182 33
air (the target’s role) 0.1912 0.2901 +0.099 425
artillery 0.0998 0.1839 +0.084 178
armor 0.1360 0.1541 +0.018 210

The support stratum has the largest apparent lift and 33 players; it stays in the table and out of the chart. Among readable strata, players already leaning toward the target’s own role start from the highest base and gain the most; players furthest from it barely move. Which is what nested substitution predicts, and is reassuring rather than surprising.

The choice model, and what it gets wrong. A conditional logit on the offer sets, fitted within the chosen role so the comparison stays among close substitutes, over 21,652 occasions:

Term Coefficient SE 95% CI
log cost −2.365 0.066 [−2.496, −2.235]
log power 1.623 0.078 [1.469, 1.776]
log synergy 0.397 0.029 [0.340, 0.455]

The point of fitting it is portability: the coefficients predict what a different price would do without running the test again.

But converting a logit coefficient into an aggregate share elasticity with the flat formula coef × (1 − share) assumes substitution is proportional across all units. It is not; it is concentrated inside the role.

  • Implied share elasticity from the choice model: −2.00
  • Observed arc elasticity from the arms: −1.45

The observed number is the one to quote. The gap between the two is a measure of how nested the substitution is, which is itself the finding, and it is only visible because the experiment was run. A model fitted to observational data would have reported −2.00 with a tight interval and nobody would have known.

Equipment: diversion or expansion?

The equipment slot adds more ways to invest rather than a cheaper one. The test of expansion against diversion is occasions per player: if the level does not move, the treatment reshuffled rather than grew.

Arm Target share Same-role Other roles Occasions
Control 0.1592 0.2094 0.5846 3.384
Extra slot 0.3234 0.1401 0.4932 3.442

Target share 15.9% → 32.3%. Occasions 3.38 → 3.44, p = 0.48.

A share move twice as large as the price cut’s (+16.4pp against +8.1pp), with the level flat. Diversion, not expansion. And the diversion pattern is different: here other roles gives up 9.1pp alongside same-role’s 6.9pp, so this lever reaches further across the roster than the price cut did.

Value, randomized by server

The power buff moves demand through value rather than budget. Thirty-two servers, 1,914 players, OLS with server-clustered errors:

  • Effect on target share: +0.0599, cluster-robust 95% CI [+0.038, +0.082], p ≈ 0.
  • Naive interval, ignoring clustering: [+0.036, +0.084].
  • The cluster-robust SE came out 0.92× the naive one. Smaller.

Whether clustering widens an interval is an empirical question, not a given. It matters in proportion to how much of the outcome’s variance sits between clusters, and here that turns out to be very little. A finding, not a licence to skip it; the sink case study has the same 32 servers going the other way.

One thing not tested: occasions per player are lower in the buff arm (3.01 → 2.84) and no test is run on that difference. Do not read this lever as level-neutral the way the equipment slot was.

What the estimates actually say

  • A 25% cost change moves the target’s share +8.1pp down and −3.2pp up. A demand response, not a novelty effect, because it is symmetric.
  • The roadmap-relevant number is the diversion: about three quarters of the gain from units doing the same job, under a tenth from other roles.
  • The equipment slot moved share harder while leaving occasions flat. Diversion again.
  • The power buff moves share through value, and must be randomized by server.

And the honest limit: everything here is composition. Nothing here says anything about revenue. The level question needs the spend layer alongside it, and the next case study is where that starts.

What I would change next time

  1. Keep share as the estimand for “which unit”. It needs no revenue attribution at all, and it is immune to whales. But read levels alongside it, or a bucket that merely held its ground will look like it collapsed.
  2. Ask for the diversion, not the lift. A cost cut that steals from close substitutes is a very different decision from one that pulls in from elsewhere, and only the decomposition can tell you which you ran.
  3. Match the randomization unit to the fairness constraint. Cost and equipment can be randomized per player; power cannot.
  4. Check that ownership stayed pre-treatment. The window did the work here. In production, a discount that induces acquisition would break the sample.

Checking the answers

Parameter Estimate Truth Absolute error
Cost sensitivity −2.3655 −2.3636 0.0018
Power sensitivity 1.6226 1.6364 0.0138

Worth noting what “truth” means here, because a naive scorer would report a disaster. The recovered quantity is the structural coefficient divided by the nest scale, not the coefficient itself. Comparing the fitted −2.37 against the raw structural parameter would look like a catastrophic failure of recovery. The nest is the whole difference, and knowing which quantity your estimator actually identifies is the difference between a passing check and a panic.

The broader lesson

On a roster, a lift is not a finding. It is half of one. The other half is the list of units that paid for it, and the naive readout has no column for them. Build the estimand so the diversion is visible, run the arm in both directions so novelty cannot masquerade as demand, and remember that the squad has six slots however cheap the rangers get.


Rules this demonstrates: 2 — randomize where interference occurs · 6 — guardrails, including cannibalization · 16 — what the experiment reveals beyond the win

Continues in: Price and income elasticity

Worked examples

ReportR Markdown · interactive

Unit demand and substitution — where investment goes when you change one unit

The cost ladder, the diversion decomposition, the conditional logit, and the equipment and power arms.

Open full report ↗