A pre-trend too small to see biases diff-in-diff by ~77%
A gentle pre-trend too small to see biases a DiD estimate by 77% of the true effect — and it's structural, so the pre-trends test can only warn you off plain DiD, not fix it. That test catches the violation only ~16% of the time. This reproduces Roth (2022)'s known low-power result, sized for the panel lengths people actually use.
Difference-in-differences (DiD) is one of the most-used causal designs in economics, policy, and product analytics. Its credibility rests on one assumption: parallel trends — that, absent treatment, the treated and control groups would have moved in lockstep. The standard way to defend that assumption is a pre-trends test: check that the two groups were not already diverging before treatment. A non-significant pre-trends test is routinely read as "parallel trends holds, the estimate is clean."
We built a runnable receipt of a result Jonathan Roth (2022) already established — that this reassurance is largely false — sized for the panel lengths people actually use. One precision worth stating up front: the 77% bias / 16% detection numbers below are OUR simulation's output, not quoted from Roth's paper. Roth's own power calibration comes from a heterogeneous survey of real published DiD designs, not a single canonical setup; we built one specific, extreme stylized design (a single treated unit) to make his mechanism concrete and runnable at a realistic short-panel size — reproducing the qualitative pattern he established, with our own numbers, not restating his.
What we did. We simulated a DiD design — one treated unit, 20 controls, 6 pre-treatment and 4 post-treatment periods, a true treatment effect of 2.0 — and injected three textbook violations of its identifying assumptions, 2,000 Monte Carlo draws each (seed-fixed, unit noise SD = 1). For every violation we measured two things: the bias it puts into the DiD estimate, and how often the pre-trends test detects it at the 5% level, using the statistically correct small-sample t-test (4 degrees of freedom). A single treated unit is itself a special case: valid inference there generally requires assuming errors are distributed the same way across treated and control units (Conley & Taber, 2011) — an assumption we impose by construction (equal noise SD for every unit) rather than one a real single-treated-unit study gets for free.
What we found.
- A barely-visible pre-trend does most of the damage — and no test can fix it, only flag it. A gentle differential drift of 0.3 units per period — easy to miss on a noisy pre-period plot — biases the estimate by +1.54, or 77% of the true effect (a steeper 0.6/period drift roughly doubles that, to 150%). This violation is structural: the drift continues into the post-period, so a simple pre/post DiD estimator can't separate it from the treatment effect even once you know it's there. The pre-trends test's job here isn't to let you "fix" the estimate — it's to warn you off plain DiD in favor of a design that doesn't assume the drift is zero (a trend-adjusted estimator, synthetic control, or Rambachan & Roth's sensitivity bounds).
- The test that's supposed to flag it usually doesn't. Against that violation, the small-sample-correct pre-trends test fires only ~16% of the time — a common-but-inappropriate normal-approximation cutoff would report ~31%, overstating power because it ignores that only 4 degrees of freedom back the estimate. Either way, most studies carrying a bias this size pass the check and report a confidently wrong number — consistent with Roth (2022), who found the same underpowered-pretest pattern in simulations calibrated to real published DiD studies, not just a toy design.
- At short panels the test is also mis-sized. With only 6 pre-periods, the normal-approximation cutoff rejects a perfectly clean design ~13% of the time at a nominal 5% level — so it both passes bad studies and cries wolf on good ones; the t-test fixes the size but not the power problem.
- Not all violations are equal — and one is genuinely fixable. Unlike the drift above, "anticipation" (the outcome moves just before treatment) and "composition" (a level shift partway through the sample) are localized to the pre-period, so if the test does catch them, you can condition on or exclude the affected periods and recover a clean estimate. Anticipation biased the estimate least here, and in the opposite direction — it attenuates the estimate (−9% to −17%), not inflates it. Composition biased it up to ~50% of the true effect, tracking violation size.
| Violation | Bias on the estimate | Detection / false-positive rate (5% level) |
|---|---|---|
| Gentle drift (0.3/period) | +77% of true effect | ~16% |
| Steeper drift (0.6/period) | ~150% | (same test, not separately reported) |
| Anticipation | −9% to −17% (attenuates) | localized — fixable if caught |
| Composition (level shift) | up to ~50% | localized — fixable if caught |
| Clean design (no violation) | 0 — correctly unbiased | false-positive ~13% (normal-approx cutoff, nominal 5%) |
The practical rule. A non-significant pre-trends test is weak evidence of parallel trends when you have few pre-periods — its power against the single most damaging violation is around one in six. This is exactly why Rambachan & Roth's "honest DiD" approach and Roth's own pretrends package exist: instead of a pass/fail gate, report the power your specific design actually has, and bound the estimate against the plausible violations that power can't rule out. Don't treat "passed the pre-trends test" as clearance — especially with a short panel; power scales with pre-period count, so a design with 15-20 pre-periods is a materially safer story than the 6-period case measured here.
What would change our mind. A pre-trends test, or a modern alternative, that achieves high power against a slope-0.3 violation at six or fewer pre-periods would break the "weak clearance" conclusion for short panels specifically. We'd publish that. One live pushback worth flagging: Mikhaeil & Harshaw (2025) argue the whole "underpowered" framing implicitly tests against a zero-violation null nobody actually believes, and propose a conditional-extrapolation threshold instead — a contested, unresolved alternative to Roth's diagnosis, not a settled rebuttal.
(Numbers above were re-measured from scratch for this post; the bias figures also match a closed-form check — a drift of slope s over this design biases DiD by s times the gap between the post- and pre-period midpoints, here s×5.)
FAQ
Can a pre-trend too small to see still bias difference-in-differences? Yes. In a DiD simulation (one treated unit, 20 controls, 6 pre- and 4 post-periods, true effect 2.0) a gentle differential drift the eye and the standard test miss still inflates the estimate substantially — and being structural, no test can undo it, only warn you to change design.
Why is the parallel-trends assumption so fragile? Because DiD credibility rests entirely on it, yet a drift small enough to pass a pre-trends test can still bias the post-period estimate — the assumption fails quietly, exactly where you can't see it.
Does a longer pre-period fix it? It's the real fix, not just a help: pre-trends test power scales with the number of pre-periods, so a 6-period panel like ours sits on the low-power end of that curve. Roth's pretrends package computes the power your specific panel has against a hypothesized drift — use it before trusting a pass.
What should I report instead? Report the sensitivity of your effect to a range of plausible pre-trends — Rambachan & Roth's "honest DiD" bounds — rather than a single pre-trends p-value that a gentle violation slips past.
Is this a new finding? No. Jonathan Roth (2022) established that pre-trends tests are underpowered against realistic violations; what we add is a runnable, panel-length-specific receipt that reproduces his mechanism from scratch, plus the closed-form arithmetic behind the numbers. The 77%/16% figures are our own simulation's numbers illustrating his mechanism, not quotes from his paper — his calibration used a survey of real published designs, not one stylized single-treated-unit setup like ours.