1. What a sample ratio mismatch is
You plan a 50/50 test and expect roughly half the users in each variant. A sample ratio mismatch, or SRM, is when the observed split differs from the planned one by more than chance can explain. It matters more than it sounds, because it does not mean one variant is slightly unlucky. It means something in the pipeline treated the two groups differently before any outcome was measured, and whatever did that has probably also biased the metric you care about.
2. How to run the check
Use a chi-square goodness-of-fit test on the user counts per variant against the planned ratio. A worked example: you plan 50,000 users split evenly, so 25,000 each, and observe 25,600 in one variant and 24,400 in the other. Each side is off by 600, so the statistic is 600 squared divided by 25,000, which is 14.4, doubled for two variants to 28.8. With one degree of freedom that corresponds to a p-value on the order of one in ten million. A 52/48 split at that scale is not noise. Many teams alert at a strict threshold such as p below 0.001 rather than 0.05, because you are running the check on every experiment and do not want constant false alarms.
3. The usual suspects
When the check fails, work through the plumbing before touching the analysis. Caching can serve one variant to many users, as covered in the server-side experimentation article. A redirect or a slower variant can lose the exposure event for some users before it fires, so that variant looks smaller. Bot filtering can remove more traffic from one side. An identity change, such as anonymous to logged in, can move a user between buckets. And consent can differ by variant: if one variant delays the banner or the tag, fewer of its users are measured at all, which is the instrumentation trap described in the consent testing playbook.
4. Where it sits in a readout
Run the check first, before you look at the metric, and record the result in the readout. If the split is fine, one line is enough. If it is not, stop: do not report the lift, and do not try to correct for it by reweighting, because you do not know what the missing users looked like. I once spent most of a day debating a surprising positive result before someone ran the split and found it was well off the plan; the debate was about a number that was never trustworthy. This belongs alongside the exposure and readout practices in the GA4 A/B readout article.
5. Limits and a sensible default
My take: make SRM a standing gate on every experiment, automated if you can, rather than something someone remembers to do when a result looks odd. The test only detects a mismatch; it cannot tell you the cause, and a passing check does not prove the experiment is clean, since a bias can exist without moving the ratio. It also needs enough users to have power, so on a tiny test a real problem may not be flagged. Treat a pass as one reassuring signal, not a certificate.
How different teams plug in
A mismatch usually has an owner outside the analytics team, so the check needs a route to them:
- Analytics runs the check, records the result, and blocks a readout when it fails.
- Engineering investigates assignment, caching, redirects, and exposure logging when a mismatch appears.
- Product holds the decision until the experiment is confirmed healthy rather than arguing over a suspect lift.
- Leadership treats a failed check as a reason to rerun, not a reason to trust a convenient result.