Pranesh Negi

Intermediate

Does the lift last? Novelty effects and long-running holdouts

A winning short test may only have measured curiosity; a small holdout and a planned re-check tell you whether the lift lasts.

1. What a novelty effect is

People click a new button partly because it is new. A redesigned banner, a rearranged menu, or a fresh layout can lift a metric for the first days and then fade as visitors get used to it. A standard two-week test can end during the peak and report a win that the business never sees in the long run. The opposite also happens: users resist change at first and a variant that eventually helps looks like a loser early on. Either way, a short test measures a mix of the change and the surprise of the change.

2. How to spot a decaying lift in a readout

Plot the difference between variants by day or by week instead of reporting only the cumulative average. A lift that starts large and shrinks steadily is a warning sign, as is one that is concentrated among returning users who noticed the change. Add a segment cut for new versus returning visitors, since genuinely new visitors have nothing to be surprised by. This is one more reason to keep the segment breakdown that the GA4 readout article recommends, rather than reporting a single topline number.

3. The holdout: keep a few users in the old experience

After you ship a winner, keep a small group in the original experience, commonly 5 to 10 percent of users, and compare them with everyone else weeks or months later. If the gap persists, the effect is real; if it has closed, you shipped a novelty. The cost is that a slice of your users does not get the improvement, so it is worth it for changes with large or expensive consequences and not for every test. Assignment for a long-running holdout should be deterministic, as in the server-side experimentation article, so the same people stay in the holdout.

4. Word the status honestly

I stopped writing "shipped, +4 percent" in summaries after a win quietly decayed over two months and nobody noticed until a quarterly review. Now the status reads "shipped, lift still being verified against a holdout, re-check on this date". That sentence sets the expectation that the number may move, which is easier to live with than a correction later. The framing fits the async summaries in the stakeholder-friendly summaries article.

5. When it is not worth the trouble

My take: for low-stakes copy or color tests, a holdout is overkill, and a re-check of the metric a month later is enough. Reserve the structured holdout for pricing, navigation, onboarding, and anything touching revenue. A limitation to keep in mind: a holdout compares groups that have lived through different histories, so other changes shipped in the meantime can contaminate the comparison unless they are applied to both groups. Record what else changed during the holdout period.

How different teams plug in

Long-running measurement needs someone to own it after the launch excitement fades:

  • Analytics sets the re-check date, tracks the holdout, and reports whether the lift held.
  • Product decides which launches justify a holdout and accepts that some users wait for the change.
  • Engineering keeps holdout assignment stable and excludes it from unrelated rollouts where possible.
  • Leadership treats a first-month lift as provisional until the re-check confirms it.
Shipped a win that faded later? Feel free to drop me a mail!