Game analytics · Part 02 of 10 · August 23, 2026

The Experiment Module

What it takes to make the rules in this series enforceable rather than aspirational.


Why a module at all

Every rule in the rules guide is a rule someone has to remember at the right moment, and people reliably do not. Not from carelessness. The moment a rule matters is the moment a result is on the screen and a meeting is in forty minutes, and that is precisely the moment at which “did we log exposure separately from assignment” is the least welcome question in the building.

So the module’s job is to move each rule from remember to check this to the pipeline will not produce a readout unless this was checked. Specifically:

  • Assignment and exposure are first-class events, logged separately, with eligibility fixed from pre-treatment information (rules 3 and 4). A readout computed on the exposed population is possible, but it is labelled as such, and the intention-to-treat estimate is the one the pipeline produces first.
  • Guardrails are defined once, in the modelling layer, rather than re-derived by hand for every test (rule 6). A guardrail is a metric definition with a threshold attached; if the definition lives in a notebook, it drifts.
  • The primary metric is declared before launch and computed by the pipeline, not by whoever writes the deck (rules 1 and 5). The declaration is a config entry with a timestamp; the deck reads from the output, not the other way round.
  • The decision log is written back (rule 10): hypothesis, metric, guardrails, power assumption, result, and what was done about it, in a table the next analyst can query rather than a slide nobody can find.

None of this is clever. It is the boring insight that a rule enforced by a pipeline gets followed on a bad day, and a rule enforced by memory does not.


How it fits together

The experiment module's data pipelineTelemetry contracts, game server metadata, and a synthetic data simulator feed ingestion. Ingestion splits into two coupled transform paths: an ELT modelling layer that also holds the experiment guardrails, and an ETL heavy-compute lane that produces the primary metrics. Both feed a reporting layer and a combined segmentation-and-modelling block used for post-experiment deep dives. Data quality, orchestration, and privacy governance run across the whole thing.Maturity (node fill)DesignPrototypeTestingPilot-readyPathBuiltIn progressValidation harness◐ Data quality & validationfreshness · schema · volume · business rulesdbt tests · Great ExpectationsInbound sourcesTelemetry & contractsclient / server event schemasassignment · exposure events◕ Pilot-readyGame server metadataentity · config · economy state◑ TestingData simulatorknown-answer synthetic data◑ TestingIngestIngestionstream · batch → data lakeassignment logs land raw◐ PrototypeTwo coupled transform pathsELT — data modelingstaging → marts · KPI taxonomyguardrails are defined hereretention · crashes · economySRM & balance checks◑ TestingcoupledETL — heavy computeprimary metric computationattribution · reprocessingvariance reduction · CUPED◑ TestingOutputsReporting layerreadouts · guardrail scoresdecision log◐ PrototypeSegmentation & modellingpost-experiment deep divesheterogeneous effectsuplift & propensity modelselasticity · substitutionsurvival / time-to-event◐ Prototype◑ Orchestration railscheduling · dependencies · run-time reliabilityPrefect◇ Privacy, consent & PII governanceconsent at collection · pseudonymization · retention · DSAR suppressionExcluded from this view: third-party feeds, identity resolution, and Labs — none of themsit on the critical path for an experiment readout.
Read the diagram as a list
  1. Inbound sources. Telemetry design & contracts(pilot-ready) — client and server event schemas, including the experiment assignment and exposure events. Game server metadata(testing) — entity, config, and economy state. Data simulator(testing) — known-answer synthetic data, the reason anything downstream can be validated before a live game is connected.
  2. Ingestion (prototype) — stream and batch into the data lake; assignment logs land raw and immutable.
  3. ELT — data modeling (testing) — staging through marts and the KPI taxonomy. The experiment guardrails are built inside this block: retention, crash and latency, economy, and competitive-balance definitions, plus the SRM and covariate-balance checks, all versioned next to the metric definitions they protect.
  4. ETL — heavy compute (testing) — theprimary metric grabs and processing, alongside revenue attribution, large reprocessing, and variance-reduction work that does not belong in warehouse SQL. Coupled to the ELT path rather than independent of it.
  5. Reporting layer (prototype) — the readout, the guardrail scorecard, and the decision log.
  6. Segmentation & modelling (prototype) — one block rather than two, because post-experiment work is a single activity: heterogeneous treatment effects, uplift and propensity models, elasticity and substitution estimates, and survival analysis, all of which only make sense once a result exists.
  7. Cross-cutting. Data quality & validation(prototype); the orchestration rail (testing); and privacy, consent & PII governance (design).
The experiment module as it currently stands. Node fill deepens as a module hardens; solid paths are validated, dashed ones are still being wired. Everything here is Savepoint's own reference architecture, generalized methods rather than a client's system, data, or schema.

The two paths, and why they are two

The diagram splits the middle of the pipeline into an ELT path and an ETL path, and the split is not architectural fashion.

ELT, where the guardrails live

The ELT path is warehouse SQL: staging models, marts, and a versioned KPI taxonomy on top. Guardrails belong here because a guardrail is a metric definition. Retention, crash rate, refund rate, session days, total wallet: each has one definition in the taxonomy, and the guardrail check references that definition rather than re-implementing it.

The reason to be strict about this is that a guardrail that drifts from the metric it guards is worse than no guardrail. In Chernobyl, the AZ-5 button was the reactor’s emergency shutdown, and pressing it was what finished the job, because the safety system and the reactor had quietly stopped agreeing about what the control rods did on their way in. A refund guardrail computed from a slightly different join than the finance refund metric will pass on the day the finance metric fails, and the fact that it was labelled “guardrail” will make the failure harder to believe, not easier.

ETL, where the primary metric is computed

The ETL path is heavy, non-SQL compute: the primary-metric grab and its processing, revenue attribution, reprocessing, and variance reduction (CUPED against a pre-period covariate). It is separate because the primary metric for an experiment frequently cannot be expressed as a warehouse aggregate without one of two sins. Either the aggregate treats rows as independent when the randomization was at a coarser grain, which is pseudo-replication (rule 12), or it reaches for a mean on a distribution where rule 14 says the mean is a choice rather than a measurement.

Clustered standard errors, bootstraps, winsorization with a pre-registered cap, and rank tests are not things a mart should be doing, and pretending otherwise produces marts that are slow, wrong, or both. So the heavy path reads from the marts, computes the decision statistic properly, and writes the result back.

The two paths are labelled coupled in the diagram because the ETL path depends on the ELT path’s definitions. It does not get its own idea of what a purchase is.


Segmentation and modelling as one block

The platform-level diagram keeps segmentation and machine learning as separate outputs. The experiment module merges them, on purpose.

For an experiment, they are not two parallel outputs. They are one activity that happens strictly after a result exists. Heterogeneous treatment effects, uplift models, elasticity and substitution estimation, survival analysis: every one of them takes the experiment’s assignment as an input, and every one is exploratory by construction (rules 9 and 14). Putting them in one block, downstream of the readout, is a structural statement that they cannot change the headline. The headline was computed before they ran, from a metric declared before the test launched.

This block is also where the reusable asset gets built (rule 16). The elasticity and substitution case studies both end with a parameter the economy team can carry into the next decision; this is where that parameter is estimated, and where it is stored so the next test can check whether it held.


What it does not do

Scope, stated plainly, because the temptation with a diagram is to let the reader assume the boxes that are not drawn.

  • It is not an assignment service. It consumes assignment; it does not decide it. The studio’s feature-flag or experimentation platform does the coin flip; the module logs the outcome and checks it. This is a design decision rather than a gap. Owning assignment would mean owning the client SDK, and that is a different product.
  • No reverse-ETL activation. Results are not pushed back into the game or the CRM. The module produces a decision record; acting on it is somebody else’s pipeline.
  • No identity resolution. It inherits whatever player identity the studio already has, with the consequence that if the studio double-counts players, so does the module. Identity is one of the corners Savepoint treats as must-be-right, but it lives in its own module, not this one.
  • No sequential testing. The current design assumes a fixed horizon, set before launch. Valid always-on monitoring (sequential tests, alpha spending) is a genuine gap rather than a boundary, and the one on this list most likely to change.

Status

The module is in final testing and close to pilot-ready. The maturity labels in the diagram are per-component and deliberately unflattering. Telemetry contracts are the furthest along; ingestion, the reporting layer, and the segmentation-and-modelling block still say prototype, and they say it because they are prototypes. The labels are there so a reader can tell the difference, which, come to think of it, is the whole point of everything in this series.