Twenty peeks at a checkout test and a suspiciously fast win
To build trust in their A/B tooling, an e-commerce team runs an A/A test: both arms get the exact same checkout flow, so any measured difference is pure noise. They watch the conversion dashboard once a day for 20 days and record whenever the p-value dips below 0.05. To their surprise, it crosses 0.05 on two separate days.
Explain why an A/A test with no real difference still produces "significant" days, and roughly how likely at least one false crossing is over 20 daily looks.
Your answer
Solving needs a free account
Answers, streaks and solutions unlock when you are signed in. Reading the question and the hint stays free.
Discussion
Sign in to join the discussion · reading is open to everyone
💡 Discussion rules
- No full solutions here. Hints and approaches only.
- Complexity, edge cases and intuition are the point.
- Interview experiences are welcome. Respect your NDAs.
Loading discussion…
Learn the concepts
The theory behind this question.
Related questions
Why you can't stop an A/B test when it "hits significance"One guardrail metric out of twenty went redThe test isn't significant, can we just run it longer?Why stopping an A/B test at first significance backfiresAn automated alert that pings the moment p drops below 0.05
All questions →