Qm

Twenty peeks at a checkout test and a suspiciously fast win

To build trust in their A/B tooling, an e-commerce team runs an A/A test: both arms get the exact same checkout flow, so any measured difference is pure noise. They watch the conversion dashboard once a day for 20 days and record whenever the p-value dips below 0.05. To their surprise, it crosses 0.05 on two separate days.

Explain why an A/A test with no real difference still produces "significant" days, and roughly how likely at least one false crossing is over 20 daily looks.

Your answer

Solving needs a free account

Answers, streaks and solutions unlock when you are signed in. Reading the question and the hint stays free.

Discussion

Sign in to join the discussion · reading is open to everyone

💡 Discussion rules

  1. No full solutions here. Hints and approaches only.
  2. Complexity, edge cases and intuition are the point.
  3. Interview experiences are welcome. Respect your NDAs.

Loading discussion…

Learn the concepts

The theory behind this question.

Related questions

Why you can't stop an A/B test when it "hits significance"One guardrail metric out of twenty went redThe test isn't significant, can we just run it longer?Why stopping an A/B test at first significance backfiresAn automated alert that pings the moment p drops below 0.05
All questions →