Game analytics · Part 10 of 10 · August 23, 2026

Matchmaking, Rewards, and Dose

Simulated data with recorded structural parameters. Where the notebook’s own prose disagrees with its executed output, and in about a dozen places it does, the numbers below are the executed ones.


The question

Matchmaking fills thin rating tiers with designed opponents: content that is harder than the human target it replaces, and that pays a reward premium to compensate. Anyone who climbed a ranked playlist in Halo 2 or Halo 3 knows what a thin tier feels like from the inside. The ranking system did its job, the pool at the top got very small, and the queue got long and repetitive in proportion. Designed content is one answer to that, and the studio wants to know two things about it, one strategic and one operational.

Do players actually engage with the harder content once you price both what it pays and what it costs to clear? And how large does the reward premium have to be per unit of added difficulty, so the next tier can be priced before it ships rather than tuned after?

What interested me is that the second question has a clean answer, an exchange rate, and the first has three plausible estimators that give three different numbers. Working out why they differ turned out to be worth more than any of the numbers.

Why the effect is harder to see than it looks

Supply of the designed content is endogenous. The queue serves it when it cannot find a good human match, so its share rises as the player’s rating-pool density falls. “Saw a lot of it” is therefore substantially “sits in a thin tier”, and any analysis treating exposure as randomly assigned is comparing tail players to median players. Left 4 Dead’s Director turns up the horde when the survivors are doing well, and a naive analyst would conclude that hordes improve team performance. Same structure: the system allocates the treatment in response to the state you are trying to measure.

What is randomized is the reward premium those matches carry. That is the lever the experiment controls, and everything about designed-content exposure itself has to be handled as observational.

Second structural fact: the offer is declinable and inspectable. The player sees roughly what a target holds before committing, so the refusal is the informative event and the offer log is the estimation sample. Attacks are the offers that cleared a reservation value; conditioning on them is post-treatment selection, the same mistake in a different costume.

Target type Offers Difficulty Posted reward Repair cost Take rate
Designed 24,312 0.5715 1,352.87 851.46 0.3573
Human 70,248 0.4405 920.26 757.18 0.3756

Harder, pays more, costs more. The raw take-rate difference is uninterpretable because it nets a difficulty penalty against a reward premium, which is exactly why the choice model below has to separate them.

My approach

Randomization is per player; the analysis grain is the offer. Offers repeat within a player and assignment is per player, so the naive standard error is anti-conservative and every estimate is clustered on the player.

Four arms: a reward ladder, plus one arm that moves both levers, which is what makes the exchange rate identifiable at all:

Arm Reward premium Difficulty scale Identifies
control 1.35 (shipped) 1.0 the operating point
premium_flat 1.00 1.0 the premium’s own contribution
premium_high 1.70 1.0 the reward slope
harder_richer 1.70 2.0 the exchange rate

SRM p = 0.771. 94,560 offers, 4,656 players.

Then two further reads that are not part of the randomized design, and are labelled as such throughout: a dose-response difference-in-differences on the feature’s introduction, and a separate server-randomized on/off test.

The assumptions doing the heavy lifting

Unaffordable offers are rationing, not preference. They are excluded from the choice model. A player who cannot pay the repair bill has not declined the target; they have been priced out of the decision, and folding them into the “declined” pile would manufacture cost sensitivity.

Posted loot is a ceiling, not a payment. Expected reward is constructed through the player’s own outcome distribution rather than posted loot, because a harder target is worth less than its posted number twice over: through a lower clear rate and a bigger repair bill.

The human surface is a valid within-player placebo. Both levers are scoped to designed content. Human targets sit in the same queue, for the same players, in the same window, and carry neither, so their take rate should not move. If it does, something other than the treatment is moving.

For the dose design: parallel trends. I test this rather than assert it, and two of four outcomes fail. That is reported below, not hidden.

What the readout says

Do players like it, net of what it pays?

A choice model on the offer log in reward-equivalent units:

V = α + β_reward·log1p(E[reward]) − β_cost·log(repair) + β_designed·is_designed

β_designed = +0.128, 95% CI [+0.093, +0.163], worth about 13% of expected value in reward terms. Players prefer the harder content at equal reward and equal difficulty, and that engagement is free before any premium is posted.

Now drop the term:

Parameter With the designed-content term Omitting it
β_reward 1.073 1.072
β_cost 0.603 0.534
β_designed 0.128 —

Designed targets are harder, so they carry a bigger repair bill. With the preference omitted, that preference loads onto the cost coefficient and drags it toward zero. The reward coefficient is untouched; the damage is entirely on the cost side. The indicator is a parameter, not a control, and the mis-specified estimate lands outside the correct model’s own confidence interval. This is the omitted-variable lesson from any econometrics course, and it is easy to miss in practice, because the mis-specified model fits perfectly well.

The indifference curve

Not an elasticity: how much reward buys one unit of difficulty. Two arms post nearly identical loot and differ only in difficulty, so the gap between them is the difficulty slope; the reward rungs give the reward slope; the ratio is the exchange rate.

Arm Difficulty Posted Offers Take rate vs control
control 0.5288 1,272 6,285 0.4700 —
premium_flat 0.5339 949 6,164 0.4219 −4.81pp
premium_high 0.5250 1,623 5,381 0.5097 +3.97pp
harder_richer 0.6872 1,591 6,482 0.4391 −3.09pp
  • Reward slope: +12.5pp of take-rate per 1.0 of premium
  • Difficulty slope: −43.5pp per 1.0 of difficulty
  • Exchange rate: 3.47 of reward premium buys 1.0 of difficulty

The shipped configuration adds +0.16 of difficulty and +0.35 of premium. Holding take-rate flat would need +0.56. It is under-compensated, which is why raising the premium alone still buys +3.97pp over control.

The placebo surface

Across all four arms the human-target take rate spans 42.46% to 43.21%, a 0.75pp spread, inside every arm’s own ±0.8pp interval, while the treated surface spans 42.19% to 50.97%. Clean. The levers moved what they were pointed at and nothing else.

Where the data had its own opinion

The mechanism, and the learning curve under it

Repair scales with difficulty, so harder content costs more of the rationed input by construction. Rising consumption is the feature working, not a cost overrun. A version of this feature that did not raise consumption would not be doing anything, and the resource-consumption guardrail I registered was therefore always going to breach. More on that below.

The corollary: efficiency on a designed target starts below the human baseline and recovers as players learn the layouts. So the effect on yield-per-cost is a time path, not a level, and a pooled average over the window describes no week in particular.

Two things move together and have to be separated. Learning applies only to the designed content; progression (better units and equipment) applies to every target. So the human surface is the within-player control, and the difference identifies learning.

Across a player-week panel, within-player demeaned, clustered on player:

week              (progression, both surfaces) : +0.00146  [+0.00037, +0.00255]
week x designed   (LEARNING, designed only)    : +0.00831  [+0.00490, +0.01172]  p = 0.000

The learning increment is about 5.7× the progression trend and its interval excludes zero.

Better still, index the curve on experience rather than the calendar:

Designed targets already attacked Difficulty Repair cost Take rate Offers
0 (first) 0.635 897.0 0.353 8,600
1–2 0.581 858.0 0.344 13,260
3–5 0.534 824.6 0.346 13,003
6–10 0.507 805.1 0.365 10,768
11–20 0.494 795.5 0.387 5,600
21+ 0.483 787.7 0.453 934

897 units of repair on a player’s first designed target, 788 by their twenty-first, while take-rate climbs from 35.3% to 45.3%. The calendar view is diluted because players accumulate experience at very different rates. Index a learning curve on player experience, and carry a progression control on any efficiency claim.

The dose design

The feature’s introduction can be read as a natural experiment, because treatment intensity varied and nobody chose it.

The dose is pre-period pool density, measured over all of a player’s offers before the introduction and split at the median. Three properties make it work, and they are worth naming separately because they are what any dose variable needs:

  1. Fixed before treatment. Measured entirely in the pre-period, so it cannot be contaminated by the treatment it is meant to dose.
  2. Not chosen by the player. Designed content backfills queue requests that would otherwise stall, so the dose is a mechanical consequence of ladder position. Nobody opted in.
  3. Observable live. It is just where you sit on the ladder, a quantity the game already has, needing no extra instrumentation.

Verify the dose before claiming an effect. Total offers per player-week:

Pool Before After Gain
Thick 4.83 5.25 +0.42
Thin 2.50 4.51 +2.01

A differential of +1.58 offers a week, and a 4.8× larger gain for the thin group. That asymmetry is the dose.

The dose-response estimates, two-way fixed effects on player and week, clustered on player:

Outcome Effect 95% CI p
Rationed input per active day +409.14 [+370.3, +448.0] 0.000
PvP acceptance rate +0.0060 [−0.0087, +0.0207] 0.426
Played PvP at all +0.0206 [+0.0022, +0.0390] 0.028
PvP games per active day +0.1488 [+0.1119, +0.1857] 0.000
Games, among those who played +0.1769 [+0.1359, +0.2178] 0.000

The acceptance rate is flat while both participation margins move. That is what identifies the mechanism: players did not become more willing to attack a given human target; they were offered more matches they could act on. The gain is in supply and matching, not preference.

The caveats, which are substantial

Parallel trends is tested, not asserted. Pre-period interaction of week with the dose:

Outcome Pre-trend p
PvP acceptance −0.00332 0.623
Played PvP +0.00502 0.541
PvP games per active +0.04288 0.007
Rationed input per day +78.19 < 0.0001

Two of four fail, and one of them, games per active day, is a headline positive result. The two groups were already diverging before the feature shipped, so those estimates are contaminated by whatever was driving that. They are reported anyway, as directional, because hiding a failed assumption is worse than reporting one.

The estimand is not what it looks like. Every player in the panel had the feature; only the dose varied. So this estimates the effect of more exposure among the exposed, not the effect of having the feature at all. It is a dose-response contrast wearing an average-treatment-effect’s clothes.

And the obvious successor inherits the same defect. Propensity matching would fix composition: model the probability of high exposure from pre-period observables, then match and rerun. Where a difference-in-differences says “trust the trend”, matching says “fix the composition first”, and they fail in different ways, which is why running both beats either. But with no untreated group, no amount of covariate balancing recovers the missing counterfactual. Matching fixes composition, not the absence of a control.

The randomized answer

A separate test randomizes the feature per server. Matchmaking pool composition is a shared-world property, and a per-player holdout would put held-out players into a queue the treated players had already changed: both unfair and interfered with.

Thirty-two servers, 4,770 players. Manipulation check: designed-content share 0.388 in control, 0.000 with it off.

Effect of switching the feature off:

Outcome Effect 95% CI p
All attacks per active day −0.5732 [−0.605, −0.542] 0.0000
PvP attacks per active day +0.0044 [−0.025, +0.034] 0.770
PvP acceptance rate +0.0214 [+0.011, +0.032] 0.0001
Rating error −0.0006 [−0.003, +0.002] 0.675
Rationed input per active day −429.32 [−455.0, −403.7] 0.0000

Designed content adds engagement without displacing PvP: switching it off costs about 0.57 attacks per active day of total engagement while PvP volume is unchanged. What moves is PvP acceptance, about 2.1pp higher without it, and the mechanism is the budget, because harder targets consume the same rationed input human attacks draw on. It costs acceptance, not games.

A metric-definition trap worth its own paragraph

“Did the player ever attack a human target” comes out +1.9pp with the feature off (p = 0.019), which reads as the feature suppressing PvP participation. Significant, plausible, and the kind of finding that survives peer review.

It is almost entirely definitional. With the feature off, every offer is a human target, so the chance of at least one human attack rises mechanically: the share of players receiving at least one human offer goes 0.981 → 1.000. Condition on having received a human offer and the effect collapses to +0.001, p = 0.897.

The rule: when a treatment changes the composition of what a player is offered, any “did they ever do X” outcome is partly definitional. Report it conditioned on opportunity, or do not report it.

What this design cannot see

The strategic case for designed opponents is the ladder: a reliable rated match calibrates players whose human pool is too thin. This design cannot test it, because every arm of a reward-premium test has the feature. The premium moves take-rate but barely moves match volume, so it barely moves calibration. Across arms, rating error and matches played are flat.

That is not a null about the on-ramp. It is the wrong instrument, and reporting it as a null would be the most damaging thing in the analysis.

There is a related trap in the observational read. Binned by rated matches already played, rating error rises (0.0723 → 0.1029) while the opponent gap falls (0.1844 → 0.1297) and take-rate rises (0.4017 → 0.4502). Two plausible calibration metrics move in opposite directions, and only one of them supports the story. Choose the metric before you see which way it points.

Multiplicity, honestly

Twelve contrasts across four KPIs, Benjamini-Hochberg corrected within the primary and secondary families separately. Three survive, all three on the primary take-rate KPI. Everything on PvP attack rate, rating error, and resource consumption fails correction.

One of nine registered guardrails breached: the reward faucet, in the arm where the premium was removed (−30%), not where it was raised. Retention and refund guardrails are clean in every arm. A note on that guardrail’s design, because I got it wrong: the faucet will breach in every treated arm because the reward is the lever, so it should have been labelled a monitoring metric rather than a stop condition. A stop condition that always fires trains everyone to ignore the panel, and I built one.

The retention read is explicitly directional: the achieved MDE on churn is 0.057 absolute and the observed effects sit well inside it.

What the estimates actually say

Players prefer the harder content at equal reward and equal difficulty, worth about 13% of expected value, and that engagement is free before any premium is posted. The exchange rate is 3.47 units of reward per unit of difficulty, and the shipped configuration is under-compensated. Rising resource consumption is the mechanism, not a guardrail breach. The randomized test adds the ship decision: the feature is close to purely additive, costing PvP acceptance rather than PvP games.

The three estimators disagree because they estimate three things. The choice model gives the preference; the dose design gives the effect of more exposure among the exposed, with two failed pre-trends attached; the server test gives the effect of having the feature at all. Only the last one answers “should we ship it”, and it is the only one that needed 32 servers to run.

Checking the answers

Parameter Estimate Truth 95% CI Covers
β_reward 1.0732 1.10 [1.028, 1.119] yes
β_cost 0.6034 0.66 [0.529, 0.678] yes
β_designed 0.1280 0.15 [0.093, 0.163] yes
Outcome skill coefficient 3.2310 3.20 [3.122, 3.340] yes
Outcome difficulty coefficient −2.6033 −2.60 [−2.702, −2.505] yes

Five of five intervals cover the truth. Note the ordering that makes the omitted-variable story land: the true cost coefficient is 0.66, the correct specification recovers 0.603 and covers it, and the mis-specified one returns 0.534, 19% below truth and outside the correct model’s own interval.

What I would change next time

  1. Treat the exposure indicator as a parameter, not a control. Omitting it manufactures cost sensitivity that is not there.
  2. Index a learning curve on experience, not the calendar, and carry a progression control on any efficiency claim.
  3. Label the faucet a monitoring metric. A guardrail that must breach by construction is not a guardrail.
  4. Do not ask a dose design whether to have the feature. That needs an untreated arm, and in a shared-pool system it has to be randomized by server.
  5. Price the next tier off the exchange rate, then test whether it held. That is the cheapest validation of a demand model there is.

The broader lesson

Three estimators, three estimands, and the temptation throughout was to treat them as three estimates of one thing and average the disagreement away. They are not. The choice model, the dose design and the randomized holdout each answer a different question, each has assumptions the others do not, and the value of running all three is that where they disagree, the disagreement tells you which assumption is doing the work. A studio that ran only the dose design would have shipped a confident number with two failed pre-trends under it. A studio that ran only the server test would have the ship decision and no idea what to charge for the next tier.


Rules this demonstrates: 2 — randomize where interference occurs · 12 — analyze at the level of randomization · 13 — neither peek nor multiply conclusions · 15 — vary the dose when the effect is not a switch