1. What peeking does to your error rate
A standard significance test assumes you pick a sample size in advance and look once. If you check repeatedly and stop the first time the p-value falls below 0.05, you are giving randomness many chances to look like a winner. A commonly cited illustration is that five looks at a 0.05 threshold push the real false-positive rate to roughly 14 percent, and it keeps climbing with more looks. The reasoning is laid out in the Johari and colleagues paper on peeking at A/B tests, which is worth reading if you want the primary source. The practical meaning is simple: a result you stopped on is less trustworthy than the p-value printed next to it.
2. Why it keeps happening
Nobody peeks out of bad faith. The dashboard is open, a stakeholder asks how it is going, and a green number is exciting. The tooling often encourages it by showing a live significance badge from day one. The cost is a program that quietly ships changes that did nothing, and then wonders why the combined lift never shows up in the top-line metric, the pattern the velocity dashboard is meant to expose.
3. The simplest fix: decide the end date first
Before launch, compute the sample size you need for the smallest effect worth detecting, convert it into a duration that covers whole weeks, and write both in the hypothesis. Then look at the result once, at the end. This is the discipline in the hypothesis writing field guide, extended to the stopping rule. It feels slow, and it is the most reliable way to keep your false-positive rate where you think it is.
4. If you must look early, use a method built for it
Some teams genuinely need to stop early, for example to cut a harmful variant quickly. Sequential testing methods, including always-valid p-values, are designed so you can check continuously without inflating errors, and several experimentation platforms offer them. They are not a free lunch: they usually need somewhat more data to reach the same power, and you must use the platform's own stopping rule rather than eyeballing a standard p-value. I have watched a team switch to a sequential tool and keep reading the old fixed-horizon number out of habit, which defeated the point entirely.
5. Handling the daily-update request
Give stakeholders something safe to look at: sample size reached so far, the planned end date, and guardrail metrics for harm, but not the significance of the primary metric. The wording in the readout storytelling tips helps here. My take: the cultural fix matters more than the statistical one, because the habit returns the moment someone asks for a daily number. One limitation: a fixed horizon can mean running a clearly losing variant longer than you would like, so decide in advance which guardrail breaches justify stopping early.
How different teams plug in
Peeking is a team habit, so the fix has to be agreed across roles:
- Analytics sets the sample size and end date and publishes the stopping rule with the hypothesis.
- Product resists asking for a verdict before the planned end date and accepts guardrail updates instead.
- Engineering keeps experiment tooling from showing live significance badges that invite early stopping.
- Leadership rewards a correct decision at the planned time over a fast one that cannot be trusted.