1. The problem CUPED solves
Most of the variation in an A/B test metric has nothing to do with your change. Some users spend a lot and some spend little, regardless of which variant they saw, and that background spread buries a small real effect. The usual remedy is more traffic or more time. CUPED, short for controlled experiments using pre-experiment data and introduced in a 2013 Microsoft paper, takes a different route: it removes the part of each user's outcome that you could already have predicted from their behavior before the test began. The cleaner signal that is left needs fewer users to detect the same effect, which is exactly the lever the experiment velocity dashboard is trying to pull.
2. The idea in plain terms
Take a pre-test measurement for each user, such as their spend in the four weeks before the experiment, and call it X. Their in-test outcome is Y. The two usually move together. CUPED computes an adjusted outcome, Y minus a coefficient times (X minus the average X), where the coefficient comes from how strongly X predicts Y. The adjusted metric has the same expected effect as the raw one, because X was measured before assignment and cannot depend on the variant, but a smaller variance. The stronger the correlation between X and Y, the bigger the reduction.
3. When it helps and when it cannot
It helps most for metrics that are stable per user and for populations with history: returning customers, subscribers, repeat purchasers. It does little when the pre-test measurement says nothing about the outcome, and it cannot help users who have no pre-period at all, since a brand-new visitor has no X to adjust with. A site dominated by first-time traffic will see a small gain. Consent also bites here: if you cannot reliably link a returning visitor to their earlier activity, the pre-period data is incomplete, a problem touched on in the consent testing playbook.
4. Why this lives in BigQuery, not the GA4 interface
The GA4 interface has no way to run this adjustment. You need user-level data for both the pre-period and the experiment window, which means the BigQuery export and a query along the lines of the event analysis workflow. The practical steps are to define the pre-period before the experiment starts, compute X for every user in either variant, estimate the coefficient on the pooled data, and compare adjusted means. Decide the pre-period in advance and write it down, in the same spirit as pre-registering a hypothesis, so nobody chooses the window that flatters the result.
5. Where I would and would not use it
The first time I applied it to a revenue metric on a returning-customer test, the confidence interval narrowed enough that a result we would have called inconclusive became readable. My take: it is worth the effort for metrics with high per-user variance on audiences with history, and not worth it for a quick test on mostly new traffic. A limitation to be straight about: it makes the estimate more precise, not more correct, so it does nothing for a broken experiment; check that the split is healthy first, as in the sample ratio mismatch check, because a precise answer to a corrupted test is still wrong.
How different teams plug in
CUPED sits between data engineering and experiment analysis, so ownership needs to be explicit:
- Data engineering maintains the user-level tables with a consistent pre-period for every user.
- Analytics defines the covariate and the window before launch and documents the adjustment in the readout.
- Product accepts that the headline number is the adjusted one and understands why it differs from the raw average.
- Leadership sees shorter tests as a result of better method, not a lowered bar for evidence.