Staggered Adoption and Two-Way Fixed Effects Bias
When different units adopt a treatment at different times, the standard two-way fixed effects regression can produce badly biased estimates because it implicitly uses already-treated units as a comparison group.
Prerequisites: Difference-in-Differences
Two-way fixed effects (TWFE) regression is the workhorse tool for difference-in-differences: it controls for unit fixed effects and time fixed effects, then reads the treatment effect off a single coefficient. That works cleanly when every treated unit switches on at the same date. It breaks down when adoption is staggered — some states pass a law in 2015, others in 2018, others never — because the regression starts using early adopters as an implicit control group for late adopters, and vice versa.
The problem is that "already-treated" units are not a valid comparison group if the treatment effect changes over time (grows, fades, or varies by cohort). The regression's single coefficient turns out to be a weighted average of many two-group, two-period comparisons, and some of those comparisons carry negative weights — meaning a genuinely positive effect in the data can flip the sign of the reported estimate.
Under staggered treatment timing, the standard TWFE estimate is a weighted average of many comparisons, some weighted negatively, so it can be badly biased or even wrong-signed even when every unit's true effect is positive.
Where it shows up
A researcher studying a policy that rolled out state by state over a decade runs one TWFE regression expecting a clean "average treatment effect." The Goodman-Bacon decomposition of that same coefficient reveals it is actually built from comparisons between early- and late-adopting states, using each as the other's control at different points, with weights driven by group size and treatment timing rather than by anything a researcher would choose on purpose.
What to do instead
Modern estimators — Callaway–Sant'Anna, Sun–Abraham, de Chaisemartin–D'Haultfœuille — explicitly build the comparison group only from not-yet-treated or never-treated units, and let the treatment effect vary by adoption cohort and time since treatment. They report a cleaner, interpretable average effect at the cost of needing more data structure (cohort indicators, staggered event-time dummies) than a single TWFE regression. The practical fix is not a regression tweak but a different estimator family — one that never lets an already-treated unit serve as another unit's control.
Related concepts
Practice in interviews
Further reading
- Goodman-Bacon, 'Difference-in-Differences with Variation in Treatment Timing'