CausalSmith · seminar slides

Forbidden Comparisons in Fixed-Effect Poisson Difference-in-Differences

We characterize when pooled fixed-effect Poisson pseudo-maximum likelihood (PPML) in staggered-adoption difference-in-differences (DiD) delivers a negative population treatment coefficient under strictly positive proportional effects.

Overview

  • The object is β(δ)\beta^\star(\delta), the limiting pooled PPML treatment coordinate.
  • Under heterogeneous proportional effects, β(δ)\beta^\star(\delta) is a misspecified projection.
  • Its sign is governed by fitted-mean-weighted residualized treatment comparisons.
  • We give an explicit four-cohort design with β(δ)<0\beta^\star(\delta)<0 while every treated cohort-time proportional effect is positive.
  • The counterfactual-share proportional treatment-on-the-treated target (PTT) remains positive in the same population environment.

Motivation

  • Staggered adoption makes a single treatment indicator tempting.
  • Linear DiD taught us that pooled two-way fixed effects can compare already-treated and newly treated cohorts in hard-to-interpret ways.
  • Goodman-Bacon (2021) gives the linear decomposition.
  • de Chaisemartin and D'Haultfoeuille (2020) show how heterogeneous effects can receive signed weights.
  • Many applied settings use multiplicative means and PPML, especially trade and count outcomes.
  • We ask how the same concern appears on the PPML scale.

Running Example

  • Think of a trade policy adopted by different country pairs at different dates.
  • Outcomes are nonnegative flows, so applied work often uses PPML with high-dimensional fixed effects.
  • The policy effect is naturally proportional: a treated cell has a multiplicative change relative to its untreated mean.
  • A single pooled PPML coefficient is often read as the policy direction.
  • Our results characterize the population object behind that coefficient.
Cohort 2 adopts at date 2 Cohort 3 adopts at date 3 Cohort 4 adopts at date 4 Never-treated no adoption Treated cells cohort 2 periods Treated cells cohort 3 periods Treated cells cohort 4 periods Pooled PPML one coefficient
illustrative Boxes for cohorts 2, 3, 4, and never-treated, arrows from each cohort box to its treated cohort-period cells, and arrows from all cells into one pooled PPML coefficient.

Setup

  • Units belong to adoption cohorts GiG_i, including the never-treated cohort \infty.
  • Treatment is absorbing, so cohort gg is treated in period tt when Dgt=1D_{gt}=1.
  • The untreated mean has a unit baseline and a calendar component.
  • Treated outcomes follow cohort-time proportional log multipliers δgt\delta_{gt}.
  • Cohort shares stay positive in the large-array limit.
  • Within-cohort baseline averages converge, so the panel collapses to cohort-time cells.

Assumptions

The load-bearing structure is multiplicative untreated means, proportional treatment effects, and a full-rank collapsed fixed-effect design.

Assumption A-2 (Untreated exponential mean)

For every unit i=1,,Ni=1,\ldots,N and period t=1,,Tt=1,\ldots,T, EPN ⁣[Yit(0)]=biNexp(γt0), E_{\mathcal P_N}\!\left[Y_{it}(0)\right] = b_{iN}\exp(\gamma_{t0}), with the normalization γ10=0\gamma_{10}=0.

Assumption A-4 (Proportional treatment effects)

For every unit i=1,,Ni=1,\ldots,N and period t=1,,Tt=1,\ldots,T with Dit=1D_{it}=1, EPN ⁣[Yit(1)]=EPN ⁣[Yit(0)]exp(δGit). E_{\mathcal P_N}\!\left[Y_{it}(1)\right] = E_{\mathcal P_N}\!\left[Y_{it}(0)\right]\exp(\delta_{G_i t}).

Assumption A-5 (Collapsed design rank)

The collapsed second-moment matrix gCt=1Tqgtrgtrgt \sum_{g\in\mathcal C}\sum_{t=1}^T q_{gt} r_{gt} r_{gt}' is positive definite.

Positive-effect Scope

  • The sign-reversal question is asked under the strongest sign benchmark.
  • Every treated cohort-time cell has a positive proportional log multiplier.
  • The multicohort scope contains early, middle, late, and never-treated cohorts.
Assumption A-7 (Strict Positive Effects)

For every cohort-period cell (g,t)(g,t) with Dgt=1D_{gt}=1, δgt>0. \delta_{gt}>0.

Assumption A-6 (Multicohort frontier scope)

The cohort set satisfies {2,3,4,}C. \{2,3,4,\infty\}\subseteq \mathcal C.

Projection Target

  • We study the population PPML projection, evaluated at cohort-time means.
  • The fitted mean μgt(δ)\mu^\star_{gt}(\delta) is the PPML fit in the collapsed cohort-time table.
  • The treatment coordinate β(δ)\beta^\star(\delta) is the single coefficient produced by pooling.
  • The key comparison object is W~gt(δ)\widetilde W_{gt}(\delta), the fitted-mean-weighted residual of treatment after partialling out cohort and time fixed effects.
  • A cell with negative W~gt(δ)\widetilde W_{gt}(\delta) acts like an already-treated comparison cell in the PPML projection.

Related Literature

  • Classical DiD builds untreated counterfactual trends from repeated observations: Ashenfelter and Card (1985), Angrist and Pischke (2009), and Imbens and Wooldridge (2009).
  • Modern staggered DiD clarifies heterogeneous-effect aggregation: Goodman-Bacon (2021), Callaway and Sant'Anna (2021), Sun and Abraham (2021), and Borusyak et al. (2024).
  • Nonlinear DiD and functional-form work frame the multiplicative setting: Wooldridge (2023) and Roth and Sant'Anna (2023).
  • PPML practice is central in multiplicative mean models: Santos Silva and Tenreyro (2006, 2011), Correia et al. (2020), and Yotov et al. (2016).
  • Moreau-Kastler (2025) supplies the closest positive proportional PTT benchmark.

Main Result: Local Sign

informal · Theorem T-1 Under collapsed rank, increasing one treated-cell proportional effect moves β(δ)\beta^\star(\delta) in the sign direction of that cell's fitted-mean-weighted residualized treatment.

Theorem T-1 (Sharp PPML Sign)

Let TT be positive and let C\mathcal C be a nonempty finite cohort set. Fix cohort shares πg(0,1)\pi_g\in(0,1), baseline levels bˉg>0\bar b_g>0, time effects γt\gamma_t, and a proportional-effect array δ\delta. Assume:

  • (Rank.) The collapsed design satisfies Assumption A-5.
  • (Cell.) The cohort kk belongs to C\mathcal C, the period s{1,,T}s\in\{1,\ldots,T\}, and the cohort-period cell is treated: Dks=1D_{ks}=1.

Write W~ks(δ) \widetilde W_{ks}(\delta) for the pseudo-true qgtμgt(δ)q_{gt}\mu^\star_{gt}(\delta)-weighted FWL residual of the treatment indicator in cell (k,s)(k,s), and define E(δ)=gCt=1Tqgtμgt(δ)W~gt(δ)2. \mathcal E(\delta) = \sum_{g\in\mathcal C}\sum_{t=1}^T q_{gt}\,\mu^\star_{gt}(\delta)\, \widetilde W_{gt}(\delta)^2 . Then E(δ)>0\mathcal E(\delta)>0, and the one-coordinate path xβ ⁣(δ with δks replaced by x) x\mapsto \beta^\star\!\bigl(\delta \text{ with } \delta_{ks}\text{ replaced by }x\bigr) is differentiable at x=δksx=\delta_{ks} with derivative qksBksexp(δks)W~ks(δ)E(δ). \frac{ q_{ks}\,B_{ks}\exp(\delta_{ks})\,\widetilde W_{ks}(\delta) }{ \mathcal E(\delta) }. Consequently, β(δ)δks<0    W~ks(δ)<0,β(δ)δks=0    W~ks(δ)=0, \frac{\partial \beta^\star(\delta)}{\partial \delta_{ks}}<0 \iff \widetilde W_{ks}(\delta)<0,\qquad \frac{\partial \beta^\star(\delta)}{\partial \delta_{ks}}=0 \iff \widetilde W_{ks}(\delta)=0, and β(δ)δks>0    W~ks(δ)>0. \frac{\partial \beta^\star(\delta)}{\partial \delta_{ks}}>0 \iff \widetilde W_{ks}(\delta)>0.

Intuition

  • PPML fits cohort and time fixed effects first through the multiplicative mean score.
  • The remaining treatment variation is the residual after that weighted fit.
  • The weights are fitted means, so high-mean cells carry more curvature in the score.
  • Increasing a treated-cell effect changes the pooled coefficient through that cell's residualized treatment value.
  • A negative residual means a larger positive effect in that cell pushes the pooled coefficient downward.

Homogeneous Benchmark

informal · Theorem T-3 Under the untreated mean restriction, collapsed rank, and a common treated-cell log multiplier, the pooled PPML coefficient recovers that common multiplier exactly.

Theorem T-3 (Homogeneous Effect Reduction)

Fix TT, a finite cohort set C\mathcal C, sampling laws {PN}N\{\mathcal P_N\}_N, potential outcomes Yit(d)Y_{it}(d), positive unit baselines biNb_{iN}, cohort shares πg(0,1)\pi_g\in(0,1), positive limiting cohort baselines bˉg\bar b_g, untreated time components γt0\gamma_{t0}, cohort-time log effects δgt\delta_{gt}, and a scalar δ0\delta_0. Suppose that

  • (Untreated means.) Assumption A-2 holds.
  • (Collapsed rank.) Assumption A-5 holds for C\mathcal C and {πg}g\{\pi_g\}_g.
  • (Common treated-cell effect.) For every gCg\in\mathcal C and every period tt, if Dgt=1D_{gt}=1, then δgt=δ0\delta_{gt}=\delta_0.

Then β(δ)=δ0, \beta^\star(\delta)=\delta_0, and, for every gCg\in\mathcal C and every period tt, μgt(δ)=Bgtexp(Dgtδ0). \mu^\star_{gt}(\delta) = B_{gt}\exp(D_{gt}\delta_0).

Primitive Frontier

  • The derivative result is local.
  • We also give a global sign diagnostic.
  • Φ\Phi, the primitive sign index, is built from cohort shares, untreated means, treatment timing, and proportional multipliers.
  • Under the stated multicohort positive-effect conditions, Φ\Phi has exactly the same sign as β(δ)\beta^\star(\delta).
  • In the same environment, Φ<0\Phi<0 implies β(δ)<0<PTT\beta^\star(\delta)<0<PTT.
Theorem T-4 (Primitive Sign Frontier)

Fix a panel length TT, a finite cohort support C\mathcal C, triangular-array sampling laws PN\mathcal P_N, potential outcomes Yit(d)Y_{it}(d), deterministic cohort labels GiG_i, positive unit baselines biNb_{iN}, limiting shares πg(0,1)\pi_g\in(0,1), positive limiting cohort baselines bˉg\bar b_g, untreated time effects γt0\gamma_{t0}, and log proportional effects δgt\delta_{gt}. Suppose:

  • (Horizon and support.) T>0T>0, T4T\ge 4, the never-treated cohort belongs to C\mathcal C, and every finite supported cohort has adoption date different from the first period.
  • (Array support.) For every NN and every unit ii, GiCG_i\in\mathcal C.
  • (Share limits.) Assumption A-1 holds for (C,Gi,πg)(\mathcal C,G_i,\pi_g).
  • (Untreated means.) Assumption A-2 holds for (PN,Yit(d),biN,γt0)(\mathcal P_N,Y_{it}(d),b_{iN},\gamma_{t0}).
  • (Baseline limits.) Assumption A-3 holds for (C,Gi,biN,bˉg)(\mathcal C,G_i,b_{iN},\bar b_g).
  • (Proportional effects.) Assumption A-4 holds for (PN,Yit(d),Gi,δgt)(\mathcal P_N,Y_{it}(d),G_i,\delta_{gt}).
  • (Rank.) Assumption A-5 holds for (C,πg)(\mathcal C,\pi_g).
  • (Frontier scope.) Assumption A-6 holds for (T,C)(T,\mathcal C).
  • (Positive effects.) Assumption A-7 holds for (C,δgt)(\mathcal C,\delta_{gt}).

Define hgt=πgbˉgexp(γt0)exp(Dgtδgt),Rg=t=1Thgt,Ct=gChgt, h_{gt} = \pi_g\,\bar b_g\exp(\gamma_{t0})\exp(D_{gt}\delta_{gt}), \qquad R_g=\sum_{t=1}^T h_{gt}, \qquad C_t=\sum_{g\in\mathcal C} h_{gt}, M=gCRg,A=gCt=1TDgthgt,Φ=MAgCt=1TDgtRgCt. M=\sum_{g\in\mathcal C}R_g, \qquad A=\sum_{g\in\mathcal C}\sum_{t=1}^T D_{gt}h_{gt}, \qquad \Phi = MA-\sum_{g\in\mathcal C}\sum_{t=1}^T D_{gt}R_gC_t . Then β(δ)<0    Φ<0,β(δ)=0    Φ=0,0<β(δ)    0<Φ. \beta^\star(\delta)<0\iff \Phi<0,\qquad \beta^\star(\delta)=0\iff \Phi=0,\qquad 0<\beta^\star(\delta)\iff 0<\Phi . Moreover, the primitive tuple induced by (πg,bˉg,γt0,δgt)(\pi_g,\bar b_g,\gamma_{t0},\delta_{gt}) belongs to RT\mathcal R_T if and only if Φ<0\Phi<0. For every x>1x>1 and y>1y>1, in the four-period support {2,3,4,}\{2,3,4,\infty\} with equal shares, unit limiting baselines, zero untreated time effects, multiplier xx on treated cells other than (2,4)(2,4), and multiplier yy at (2,4)(2,4), Φ=5x2+12x2y1516. \Phi=\frac{5x^2+12x-2y-15}{16}. Also, 5(101/100)2+12(101/100)152=44414000, \frac{5(101/100)^2+12(101/100)-15}{2}=\frac{4441}{4000}, and the explicit witness W4W_4 has Φ=1155932000. \Phi=-\frac{11559}{32000}. Let mgt=bˉgexp(γt0)exp(Dgtδgt),H={(g,t):gC, Dgt=1}. m_{gt}=\bar b_g\exp(\gamma_{t0})\exp(D_{gt}\delta_{gt}), \qquad H=\{(g,t):g\in\mathcal C,\ D_{gt}=1\}. For every (g,t)H(g,t)\in H, Bgtobs=Bgt,τgt=exp(δgt)1,ωgt>0. B^{obs}_{gt}=B_{gt},\qquad \tau_{gt}=\exp(\delta_{gt})-1,\qquad \omega_{gt}>0. The counterfactual-share weights sum to one: (g,t)Hωgt=1. \sum_{(g,t)\in H}\omega_{gt}=1. The counterfactual-share PTT satisfies PTT=(g,t)Hqgtmgt(g,t)HqgtBgtobs1. PTT = \frac{\sum_{(g,t)\in H} q_{gt}m_{gt}} {\sum_{(g,t)\in H} q_{gt}B^{obs}_{gt}} -1. Finally, Φ<0β(δ)<0<PTT. \Phi<0 \quad\Longrightarrow\quad \beta^\star(\delta)<0<PTT .

Four-cohort Witness

informal · Theorem T-2 In the equal-share four-cohort witness with flat untreated means, every treated cell has a positive proportional effect, the largest effect is in cell (2,4)(2,4), and the primitive belongs to the sign-reversal region.

Theorem T-2 (Four-Cohort Sign Reversal)

The four-cohort witness W4W_4 is constructed with cohort support C={2,3,4,}\mathcal C=\{2,3,4,\infty\}, equal shares πg=1/4\pi_g=1/4, limiting baselines bˉg=1\bar b_g=1, untreated time component γt0=0\gamma_{t0}=0, and proportional-effect log multipliers δgt={log4,(g,t)=(2,4),log(101/100),Dgt=1 and (g,t)(2,4),0,Dgt=0. \delta_{gt}= \begin{cases} \log 4, & (g,t)=(2,4),\\ \log(101/100), & D_{gt}=1 \text{ and } (g,t)\neq(2,4),\\ 0, & D_{gt}=0 . \end{cases} Then:

  • (Region membership.) The primitive W4W_4 belongs to the sign-reversal region R4\mathcal R_4 of Definition P-5.
  • (Positive effects.) Every treated supported cell has a strictly positive proportional-effect log multiplier: if gCg\in\mathcal C, t{1,2,3,4}t\in\{1,2,3,4\}, and Dgt=1D_{gt}=1, then δgt>0\delta_{gt}>0.
  • (Unique largest treated effect.) The treated cell (2,4)(2,4) has the strictly largest proportional-effect log multiplier: if gCg\in\mathcal C, t{1,2,3,4}t\in\{1,2,3,4\}, Dgt=1D_{gt}=1, and (g,t)(2,4)(g,t)\neq(2,4), then δgt<δ2,4\delta_{gt}<\delta_{2,4}.
  • (Weighted FWL residuals.) At the no-effect vector, the weighted FWL residual W~gt\widetilde W_{gt} of Definition P-4 equals 18(3311133111333113), \frac{1}{8} \begin{pmatrix} -3 & 3 & 1 & -1\\ -1 & -3 & 3 & 1\\ 1 & -1 & -3 & 3\\ 3 & 1 & -1 & -3 \end{pmatrix}, with rows g=2,3,4,g=2,3,4,\infty and columns t=1,2,3,4t=1,2,3,4. In particular, W~2,4=1/8\widetilde W_{2,4}=-1/8.
  • (Negative late-cell derivative at zero.) Holding all other effects at zero, ddxβ ⁣(δ2,4=x, δgt=0 for (g,t)(2,4))x=0=110. \left.\frac{d}{dx}\, \beta^\star\!\left(\delta_{2,4}=x,\ \delta_{gt}=0\text{ for }(g,t)\neq(2,4)\right) \right|_{x=0} =-\frac{1}{10}.
  • (Local negative derivative under positive effects.) There exists ε>0\varepsilon>0 such that, for every effect vector δ\delta' satisfying δgt<ε|\delta'_{gt}|<\varepsilon for all gCg\in\mathcal C and t{1,2,3,4}t\in\{1,2,3,4\}, if δ\delta' has strictly positive treated effects in the sense of Assumption A-7, then ddxβ ⁣(δ with δ2,4 replaced by x)x=δ2,4<0. \left. \frac{d}{dx}\, \beta^\star\!\left(\delta' \text{ with } \delta'_{2,4}\text{ replaced by }x\right) \right|_{x=\delta'_{2,4}} <0 .

Proof Sketch

  • Collapse the unit fixed-effect population criterion to cohort-time cells using cohort shares and within-cohort baseline limits.
  • Use the PPML first-order conditions to express local coefficient changes through a weighted residualized treatment.
  • The full-rank condition keeps residual treatment variation positive.
  • Eliminate fixed effects from the collapsed score to obtain the primitive sign index Φ\Phi.
  • In the four-cohort design, the late cell of the early-treated cohort has a negative residual and the explicit primitive index is negative.

Also in the Paper

informal · Lemma L-1 Under the stated support, share, untreated-mean, baseline-limit, proportional-effect, and rank conditions, the unit fixed-effect population coefficient equals the collapsed finite-array treatment coordinate and converges to β(δ)\beta^\star(\delta).

informal · Lemma L-2 Under collapsed rank, the limiting PPML projection is unique and satisfies the nuisance and treatment score equations.

Interpretation

  • Under common proportional effects, pooled fixed-effect PPML recovers the common log multiplier.
  • Under heterogeneous proportional effects, the pooled coefficient is a projection summary shaped by fixed-effect residual comparisons.
  • The four-cohort witness shows a negative limiting pooled coefficient under strictly positive granular proportional effects.
  • Positive-weight proportional targets aggregate the granular effects directly.
  • The sign of the pooled coefficient and the sign of the proportional PTT can therefore diverge in the same primitive environment.

Takeaways

  • We characterize the population coefficient targeted by pooled fixed-effect PPML in staggered-adoption multiplicative DiD.
  • The sharp sign formula links local movements in β(δ)\beta^\star(\delta) to fitted-mean-weighted residualized treatment.
  • The primitive frontier gives an exact global sign diagnostic through Φ\Phi.
  • The four-cohort witness establishes sign reversal with equal shares, flat untreated means, and strictly positive treated-cell effects.
  • The empirical message is to interpret pooled PPML coefficients as projection summaries and use granular proportional effects or positive-weight proportional aggregates for causal sign statements.