Forbidden Comparisons in Fixed-Effect Poisson Difference-in-Differences
Abstract
This paper studies the population treatment coefficient from pooling a staggered-adoption panel in a unit-and-time fixed-effect Poisson pseudo-likelihood. We show that this projection coefficient can be negative even when every cohort-time proportional treatment effect is positive. The failure reflects heterogeneous effects interacting with the fixed-effect PPML score: a four-cohort design with equal shares and flat untreated means provides an explicit sign reversal, while a positive-weight proportional treatment-on-the-treated target remains positive. The result characterizes the interpretation of the deterministic population projection targeted by the pooled criterion.
Introduction
Poisson pseudo-maximum likelihood with high-dimensional fixed effects is a standard way to estimate multiplicative mean models. It is especially common in gravity and other settings where the conditional mean is naturally proportional, zeros are empirically important, and fixed effects absorb large parts of the untreated outcome structure (Santos Silva et al., 2006; Santos Silva et al., 2011; Correia et al., 2020). At the same time, many policy applications have staggered adoption: some units are treated early, others later, and some never. In linear difference-in-differences, it is now well understood that a pooled two-way fixed-effect coefficient can fail to aggregate heterogeneous effects with interpretable positive weights (Goodman-Bacon, 2021; de Chaisemartin et al., 2020; Callaway et al., 2021; Sun et al., 2021; Borusyak et al., 2024). This paper studies the corresponding issue for a different object: the treatment coordinate of a pooled fixed-effect PPML population projection when treatment effects are proportional rather than additive.
The object of interest is deliberately narrow. Consider a finite cohort-time staggered-adoption design with positive limiting cohort shares, an exponential untreated mean with unit and time components, within-cohort baseline limits, proportional cohort-time treatment effects, and a full-rank collapsed fixed-effect design. The paper defines in Definition 3 as the treatment coordinate of the limiting collapsed Poisson pseudo-likelihood projection. This is the population analogue of the coefficient an applied pooled unit-and-time fixed-effect PPML regression is trying to summarize. Its projection definition permits signed comparison weights and distinguishes it from causal average treatment effects built as convex aggregates of granular proportional effects.
The central question is whether the sign of this pooled PPML coefficient must agree with the signs of the underlying proportional effects. Sign reversal is possible. For any supported design satisfying the horizon, no-first-period-adopter, rank, and maintained population restrictions and whose cohort support contains , Theorem 3 gives a primitive sign index . Its sign is exactly the sign of . The theorem’s closed-form polynomial calculation is a separate specialization to the exact four-cohort support . The general sign equivalence also shows that implies where is the counterfactual-share proportional treatment-on-the-treated target. Thus the pooled fixed-effect PPML coefficient can have the opposite sign from every granular proportional effect while a positive-weight proportional target remains positive.
The mechanism is local as well as global. Theorem 1 shows that, under collapsed rank, the derivative of with respect to a treated cell’s log multiplier has the sign of the treatment indicator residual after projecting treatment on the fixed effects with weights . These weights are generated by the fitted PPML mean through the pseudo-likelihood score and its curvature, producing a comparison geometry distinct from linear least squares. A treated cell with a negative weighted fixed-effect residual pushes the pooled coefficient downward when its proportional effect is increased. This is the multiplicative PPML analogue of a forbidden-comparison concern.
The paper also records the benchmark in which the pooled coefficient is well behaved. Under the same exponential untreated mean and collapsed rank conditions, if every treated cohort-time cell has the same log proportional effect , Theorem 2 shows that and that the fitted collapsed mean equals the true collapsed mean in every cell. This calibration locates the sign problem in the interaction between heterogeneous proportional effects and their compression into one pooled treatment coordinate in a staggered-adoption design.
The explicit witness is intentionally simple. Definition 6 considers four periods and four cohorts, , with equal cohort shares, flat untreated means, and strictly positive treated-cell log effects. All treated cells have multiplier except cell , which has multiplier . Theorem 4 shows that this primitive belongs to the sign-reversal region in Definition 5. The special four-cohort calculation in Theorem 3 gives and the witness has . Hence the negative coefficient occurs with equal shares, strictly positive effects, and flat untreated time components.
The distinction between the pooled coefficient and a positive proportional target is important for interpretation. Counterfactual-share proportional treatment-on-the-treated targets, such as those studied by Moreau-Kastler (2025), aggregate granular proportional effects using positive exposure-based weights. Under strictly positive granular effects, those targets preserve the positive sign. The pooled fixed-effect PPML coefficient studied here is different: it is a pseudo-true projection coefficient in the sense of misspecified likelihood theory (White, 1982; Gouriéroux et al., 1984; Gouriéroux et al., 1984). Its sign is governed by the fixed-effect fit and the induced score equation, not by positive aggregation weights.
The contribution is a sharp population diagnostic. It characterizes the deterministic projection target, derives its local and global sign formulas, and proves that under stated staggered-adoption and multiplicative-mean restrictions its sign can oppose every underlying proportional effect. These results identify the interpretation problem that granular-effect estimation and sampling inference must address in subsequent work.
The rest of the paper proceeds as follows. The next section relates the result to linear and nonlinear DiD, proportional treatment-on-the-treated targets, misspecified likelihood, and PPML applications. The following section defines the finite-array and collapsed population projections and the sign-reversal region. A later section gives the local derivative result, the homogeneous benchmark, the primitive sign frontier, and the four-cohort witness. The final section discusses interpretation and extensions. The appendix contains the collapse lemma, the pseudo-true projection lemma, and proofs of the main results.
Related Literature
This paper contributes to the econometric literature on difference-in-differences by studying a particular population object: the treatment coordinate of a pooled fixed-effect Poisson pseudo-likelihood in a staggered-adoption design. The classical program-evaluation literature established difference-in-differences as a comparison of changes across treated and untreated groups (Ashenfelter et al., 1985; Angrist et al., 2009; Imbens et al., 2009). Semiparametric and nonlinear extensions clarify how identifying restrictions and target parameters change beyond additive linear conditional-mean models (Abadie, 2005; Athey et al., 2006; Sant’Anna et al., 2020). The present paper studies the projection target generated by a deliberately pooled multiplicative regression and asks when that common coefficient can have the wrong sign even though all cohort-time proportional effects are positive.
A large recent literature shows that two-way fixed-effect linear regressions in staggered-adoption designs can aggregate heterogeneous treatment effects through signed or hard-to-interpret weights. Goodman-Bacon (2021) decomposes the linear two-way fixed-effect coefficient into two-by-two comparisons and highlights the role of comparisons between already-treated and newly treated groups. de Chaisemartin et al. (2020), Callaway et al. (2021), Sun et al. (2021), and Borusyak et al. (2024) develop alternative estimands and estimators that use interpretable comparisons under suitable identifying restrictions. Those results are primarily about additive treatment effects and linear least-squares projections. Here the regression is Poisson pseudo-maximum likelihood with unit and time fixed effects, and treatment effects are proportional. Pooling across cohorts and periods therefore operates through a different mechanism: the pseudo-true Poisson coefficient responds to the fixed-effect fit and the induced comparison weights in the multiplicative mean model.
The analysis also relates to nonlinear difference-in-differences. Wooldridge (2023) specifies nonlinear conditional-mean models whose parameters or average partial effects are given a causal interpretation under maintained identifying restrictions. Roth et al. (2020) and Roth et al. (2023) emphasize how identifying assumptions, functional-form restrictions, and target parameters must be separated. The present paper holds the multiplicative untreated-mean restriction fixed and asks what sign is carried by the common treatment coordinate when heterogeneous cohort-time multipliers are pooled in a misspecified fixed-effect PPML criterion. Its object is the pseudo-true projection coefficient , which is distinct from a nonlinear-DiD effect parameter selected for causal interpretability.
The closest proportional-effect comparison is the counterfactual-share treatment-on-the-treated target studied by Moreau-Kastler (2025). The objects differ in four steps. Its primitive effects are granular proportional effects; its target weights those effects by untreated counterfactual exposure; its sign is positive whenever every granular effect is positive; and its analysis is organized around identifying that causal aggregate. Here the same granular effects enter observed cohort-time means, but the target is the treatment coordinate solving the pooled fixed-effect PPML score equations. Those score equations induce fitted-mean curvature weights and residualized comparisons in place of counterfactual shares. Consequently, the positive proportional PTT and the potentially sign-reversing coexist as distinct summaries, and the sign-reversal result requires the pooled-score characterization developed here.
Methodologically, the paper builds on the theory of misspecified maximum likelihood and pseudo-likelihood. White (1982) formalizes pseudo-true parameters under misspecification, and Gouriéroux et al. (1984); Gouriéroux et al. (1984) develop the pseudo-maximum-likelihood properties of Poisson estimators. Andrews (1988) provides general laws of large numbers for dependent and heterogeneous settings that underlie many asymptotic arguments. The formal results here characterize the deterministic population target and its causal interpretation, separating that question from convergence of an empirical estimator to its pseudo-true target.
Fixed-effect Poisson estimation is standard in applied microeconometrics and count-data models (Hausman et al., 1984; Cameron et al., 2013). In international trade, PPML with high-dimensional fixed effects is a central empirical tool because it accommodates zeros and multiplicative gravity equations (Santos Silva et al., 2006; Santos Silva et al., 2011; Correia et al., 2020). Gravity applications often combine multiplicative conditional means, rich fixed effects, and policy variation across country pairs or time (Anderson et al., 2003; Head et al., 2014; Yotov et al., 2016; Nagengast et al., 2025). Those features make PPML attractive, but they also make the pooled treatment coefficient easy to overinterpret when treatment timing is staggered and proportional effects vary across treated cells.
Recent work continues to refine fixed-effect and staggered-adoption estimators, including imputation and two-stage approaches for heterogeneous treatment effects (Gardner, 2022; Baker et al., 2025). This paper complements that literature with a population characterization of when the coefficient produced by the familiar pooled fixed-effect PPML regression is negative despite strictly positive proportional effects. The result isolates the projection itself as the source of an unsafe causal summary and provides a diagnostic foundation for subsequent estimator and inference work.
Setup and Assumptions
Consider a finite- panel with units observed over calendar periods . Each unit belongs to an adoption cohort in the finite support ; the never-treated cohort is denoted by . Treatment is absorbing: indicates whether cohort is treated in period , and . Let and . Expectations under the triangular-array population law are written .
The first restrictions keep the large-array design stable and impose the multiplicative untreated mean structure. Potential outcomes are and . The untreated mean has a positive unit baseline and a calendar component ; is the within-cohort average baseline and its limit.
| Notation | Economic object |
|---|---|
| treatment status of cohort in period | |
| untreated cohort-time mean | |
| treated-cell log proportional effect | |
| observed cohort-time mean | |
| finite-array population PPML treatment coordinate | |
| limiting collapsed population PPML treatment coordinate | |
| fitted-mean-weighted residualized treatment | |
| primitive index with the sign of |
Assumption 1 is the positive limiting cohort shares condition: every cohort remains represented asymptotically and contributes positive mass to the collapsed comparison (de Chaisemartin et al., 2020). The limit share is denoted .
For every unit and period , with the normalization .
⊢ LeanAssumption 2 is a unit-fixed-effect multiplicative parallel-trends restriction in an exponential mean model (Wooldridge, 2023). The normalization fixes the location of the time effects and has no substantive content.
For every , the deterministic triangular-array average satisfies
⊢ LeanAssumption 3 is specific to this analysis. It is the condition that makes the cohort-time collapse deterministic: after averaging within a cohort, the only limiting untreated heterogeneity that remains is the cohort baseline .
For every unit and period with ,
⊢ LeanAssumption 4 is the cohort-time proportional treatment effects condition (Moreau-Kastler, 2025). The log multiplier may vary across treated cohorts and calendar periods, so the pooled coefficient below is not imposed to equal a common structural effect.
These assumptions induce a collapsed cohort-by-time representation. Define the untreated cohort-time mean , the collapsed cell mass , and the observed collapsed mean , where denotes the cohort-time log multiplier on treated cells. Let collect the collapsed fixed-effect regressors and the treatment indicator, and let be the same vector with the treatment coordinate removed.
The collapsed second-moment matrix is positive definite.
⊢ LeanAssumption 5 is the full-rank Poisson fixed-effect design condition (White, 1982). It rules out exact collinearity in the collapsed regressors, so the population projection has a well-defined treatment coordinate rather than an unidentified linear combination of fixed effects and treatment.
The population object is a Poisson pseudo-likelihood projection defined through the conditional-mean criterion. For a finite array, let be the unit-time regressor vector containing unit effects, time effects, and treatment; let be the corresponding coefficient vector; and let be the observed unit-time mean generated by the untreated mean and proportional effects.
Define as the last coordinate of a maximizer over of
⊢ LeanDefinition 1 is the finite-array population analogue of the unit fixed-effect PPML regression. The criterion is evaluated at means rather than realized outcomes, and is the treatment coordinate of that population projection.
Let be the set of admissible extended maximizers of over unit-effect coordinates in , time-effect coordinates in satisfying the period-1 normalization, and , with the convention . Admissibility means that the period-1 time effect is normalized to zero and that a unit-effect coordinate may equal only for a unit whose whole observed outcome path satisfies for every ; for such a unit the extended objective contribution is defined to be zero. If contains an element whose squared Euclidean norm over the finite time-effect and coordinates is no larger than that of every element of , then is the -coordinate of such a selected element. If no such minimum-finite-coordinate-norm admissible extended maximizer exists, set .
⊢ LeanThe limiting projection works directly with the collapsed cohort-time cells. With denoting the collapsed coefficient vector, write the limiting collapsed Poisson criterion as . The finite collapsed analogue, used only to connect the array to this limit, is denoted .
Define where The coefficient is the last coordinate of , and the fitted collapsed mean is
⊢ LeanDefinition 3 gives the main object studied in the paper. The maximizer , its treatment coordinate , and the fitted mean are pseudo-true quantities in the sense of misspecified likelihood theory (White, 1982; Gouriéroux et al., 1984; Gouriéroux et al., 1984). The target is therefore a deterministic population projection defined by the Poisson mean criterion.
The analysis centers on this deterministic population projection. A sample fixed-effect PPML regression replaces the means by observed outcomes and may also confront separation (Santos Silva et al., 2006; Santos Silva et al., 2011; Correia et al., 2020); its sampling theory is a distinct extension. Accordingly, references below to a “pooled coefficient” mean the population coordinate , unless explicitly stated otherwise.
The sign analysis uses the residualized treatment after partialling out the collapsed fixed effects with weights given by the fitted PPML mean. Let be the weighted least-squares coefficient from that auxiliary projection.
Define The weighted Frisch–Waugh–Lovell residualized treatment is
⊢ LeanDefinition 4 gives the comparison weights that enter the local sign formula in the next section. The residual is the fixed-effect projection residual of treatment under the curvature weights of the Poisson criterion and serves as a projection comparison weight.
The sign-reversal result is stated on designs with at least the four adoption statuses needed for the explicit construction. The next two assumptions isolate that scope and the positive-effect case of interest.
The cohort set satisfies
⊢ LeanAssumption 6 is specific to this analysis. It ensures that the design contains early, middle, late, and never-treated cohorts, the minimal structure used by the four-cohort construction.
For every cohort-period cell with ,
⊢ LeanAssumption 7 is also specific to this analysis. It deliberately restricts attention to the strongest sign benchmark: every treated cell has a positive proportional effect, so any negative pooled coefficient is a property of the projection rather than a mixture of positive and negative underlying effects.
The global parameter region collects the maintained restrictions under which the limiting PPML treatment coordinate is negative despite those positive cell-level effects.
Define as the set of all triples with , indexed by a cohort support that contains and contains no first-period finite cohort, such that there exist for with for every supported cell , for every treated supported cell with , the collapsed design has full weighted rank in the sense that and
⊢ LeanDefinition 5 names the population configurations in which the pooled limiting coefficient has the opposite sign from every granular proportional treatment effect. The notation is only a shorthand for this primitive set of shares, untreated means, and treated-cell multipliers.
The paper uses one explicit triangular array to show that this region is nonempty. Its role is evidentiary rather than calibrational: all baselines and untreated time effects are set to one, and the only large treated multiplier is placed in the period-4 cell of the cohort first treated in period 2.
Let be the cofinal triangular subsequence indexed by , , with and For each such , for every , and for every . The treatment effects satisfy for every treated cell , while
⊢ LeanDefinition 6 defines the witness . Equal cohort shares and flat untreated means remove scale and composition as explanations for the sign reversal; the later result depends only on the staggered treatment pattern and the heterogeneous positive proportional effects encoded here.
Finally, the primitive sign characterization uses an algebraic index based on observed mean components. Let denote the set of treated cohort-time cells when that notation is needed later.
Define where and The relation between and is stated separately.
⊢ LeanDefinition 7 introduces the scalar used by the primitive sign result. The component is the cohort-share-weighted observed collapsed mean, equivalently , while , , , and are its cohort sums, period sums, grand sum, and treated-cell sum. The main theorem in the next section supplies the sign relation for this scalar.
Main Results
The setup isolates a population object: the treatment coordinate of the collapsed fixed-effect PPML projection in Definition 3. The first result gives the local comparative static of that coefficient with respect to one treated cohort-time log multiplier. Its content is a nonlinear analogue of a forbidden-comparison diagnostic: after the Poisson fixed effects have been fit, the sign of the derivative is exactly the sign of the residualized treatment in that cell, where residualization uses the fitted PPML mean as the curvature weight.
Let be positive and let be a nonempty finite cohort set. Fix cohort shares , baseline levels , time effects , and a proportional-effect array . Assume:
(Rank.) The collapsed design satisfies Assumption 5.
(Cell.) The cohort belongs to , the period , and the cohort-period cell is treated: .
Write for the pseudo-true -weighted FWL residual of the treatment indicator in cell , and define Then , and the one-coordinate path is differentiable at with derivative Consequently, and
⊢ LeanAssumption 5 makes the denominator strictly positive by preserving residual treatment variation after projecting on the fixed effects. Thus the local sign follows directly from the pseudo-true fit. A treated cell with a negative residual in Definition 4 pushes the pooled coefficient downward when its proportional effect is increased; a positive residual pushes it upward. This is the PPML counterpart to the comparison-weight concern in staggered-adoption linear regressions (Goodman-Bacon, 2021; de Chaisemartin et al., 2020), with weights generated by the multiplicative pseudo-likelihood projection instead of least squares on the outcome.
The next result records the exact-calibration benchmark. When all treated cells have the same proportional log effect, the fixed-effect PPML projection recovers that common multiplier exactly.
Fix , a finite cohort set , sampling laws , potential outcomes , positive unit baselines , cohort shares , positive limiting cohort baselines , untreated time components , cohort-time log effects , and a scalar . Suppose that
(Untreated means.) Assumption 2 holds.
(Collapsed rank.) Assumption 5 holds for and .
(Common treated-cell effect.) For every and every period , if , then .
Then and, for every and every period ,
⊢ LeanThe homogeneous case is therefore an exact calibration point. If the proportional effect is common across treated cells, the treatment regressor completes the multiplicative mean specification after cohort and time effects are included, and the pooled coefficient has the expected causal interpretation. Heterogeneous proportional effects generate the sign-reversal mechanism studied here, the setting in which modern staggered-adoption estimands use cohort-specific effects and interpretable aggregation (Callaway et al., 2021; Sun et al., 2021; Borusyak et al., 2024).
The sharp local derivative identifies the cellwise mechanism. The next theorem complements it with a primitive scalar index whose sign is exactly the sign of the limiting pooled PPML coefficient under the maintained multicohort positive-effect scope. It also compares the pooled coefficient with the counterfactual-share proportional treatment-on-the-treated target, which remains positive when every granular proportional effect is positive (Moreau-Kastler, 2025).
Fix a panel length , a finite cohort support , triangular-array sampling laws , potential outcomes , deterministic cohort labels , positive unit baselines , limiting shares , positive limiting cohort baselines , untreated time effects , and log proportional effects . Suppose:
(Horizon and support.) , , the never-treated cohort belongs to , and every finite supported cohort has adoption date different from the first period.
(Array support.) For every and every unit , .
(Share limits.) Assumption 1 holds for .
(Untreated means.) Assumption 2 holds for .
(Baseline limits.) Assumption 3 holds for .
(Proportional effects.) Assumption 4 holds for .
(Rank.) Assumption 5 holds for .
(Frontier scope.) Assumption 6 holds for .
(Positive effects.) Assumption 7 holds for .
Define Then Moreover, the primitive tuple induced by belongs to if and only if .
For every and , in the four-period support with equal shares, unit limiting baselines, zero untreated time effects, multiplier on treated cells other than , and multiplier at , Also, and the explicit witness has
Let For every , The counterfactual-share weights sum to one: The counterfactual-share PTT satisfies Finally,
⊢ LeanTheorem 3 has two distinct scopes. The sign equivalence between and the pooled coefficient applies to every support satisfying the theorem’s assumptions, including supports that strictly contain . The displayed polynomial in and , and the numerical value for , specialize that general characterization to the exact equal-share four-cohort support. The index serves as a population diagnostic conditional on known or specified primitives, while the counterfactual-share proportional is the causal aggregate. In the same primitive configurations, the counterfactual-share weights are positive and sum to one, so the proportional is a positive weighted average whenever all granular proportional effects are positive.
The explicit four-cohort construction makes the sign reversal concrete. The support contains cohorts first treated in periods 2, 3, and 4, together with a never-treated cohort; untreated means are flat and cohort shares are equal. Hence the example establishes the phenomenon under conventional baseline scaling and constant untreated time components.
The four-cohort witness is constructed with cohort support , equal shares , limiting baselines , untreated time component , and proportional-effect log multipliers Then:
(Region membership.) The primitive belongs to the sign-reversal region of Definition 5.
(Positive effects.) Every treated supported cell has a strictly positive proportional-effect log multiplier: if , , and , then .
(Unique largest treated effect.) The treated cell has the strictly largest proportional-effect log multiplier: if , , , and , then .
(Weighted FWL residuals.) At the no-effect vector, the weighted FWL residual of Definition 4 equals with rows and columns . In particular, .
(Negative late-cell derivative at zero.) Holding all other effects at zero,
(Local negative derivative under positive effects.) There exists such that, for every effect vector satisfying for all and , if has strictly positive treated effects in the sense of Assumption 7, then
The residual table in Theorem 4 identifies the period-4 cell of the early-treated cohort as locally negative at the no-effect vector, and the neighborhood statement extends that derivative sign to small strictly positive effects. The actual witness , however, is a global sign-reversal example because Theorem 3 gives for its primitive tuple and Theorem 4 places that tuple in . In this example every treated cell has a positive proportional effect, and the largest effect is in that late calendar cell, yet the limiting pooled PPML coefficient is negative.
The main implication concerns interpretation under heterogeneous proportional effects: the single pooled coefficient is a pseudo-true projection whose sign is governed by fixed-effect residual comparisons. Interpretable causal summaries should therefore be built from granular proportional effects and positive aggregation weights, as in counterfactual-share targets, in place of inference from the sign of the pooled fixed-effect PPML coefficient alone (Moreau-Kastler, 2025; Wooldridge, 2023; Roth et al., 2020).
Discussion and Extensions
The results are motivated by empirical settings where the multiplicative conditional mean is substantively natural and the treatment timing is staggered. Gravity applications are a leading example: PPML specifications with high-dimensional fixed effects are standard because trade flows are modeled through multiplicative resistance terms, exposure variables, and policy shifters, and because zeros are common in the outcome (Anderson et al., 2003; Head et al., 2014; Yotov et al., 2016; Santos Silva et al., 2006; Santos Silva et al., 2011; Correia et al., 2020). In such designs, a pooled treatment indicator may look like a compact summary of a policy change. The results above show that, under heterogeneous proportional effects, the population coefficient of that regression is instead a fixed-effect Poisson projection whose sign can be governed by residualized comparisons across cohort-time cells.
The conclusion pinpoints how PPML summarizes a multiplicative mean model. The homogeneous benchmark in Theorem 2 shows that a common proportional effect is recovered exactly. With heterogeneous cohort-time proportional effects, Theorem 1 shows that increasing a positive treated-cell effect moves the pooled coefficient in the direction determined by the weighted fixed-effect residual for that cell. The interpretation problem therefore arises from compressing a collection of heterogeneous economic effects into a single pooled projection coordinate.
Limitations and future work
The primitive index in Theorem 3 is a population diagnostic conditional on known or specified primitives. It determines whether the limiting fixed-effect PPML treatment coordinate is negative, zero, or positive. Estimation of those primitives, implementation of as a specification test, and construction of positive-weight estimators from granular proportional effects are the main next steps.
That distinction is important for empirical interpretation. A negative pooled PPML coefficient is sometimes read as evidence that the treatment reduced the conditional mean. Under the conditions of Theorem 3, the same primitive configuration can instead have and a positive counterfactual-share proportional treatment-on-the-treated target. The sign reversal therefore identifies the pooled fixed-effect PPML coefficient as a projection summary distinct from positive-weight proportional targets.
The positive counterfactual-share target remains closer to the causal object under the maintained proportional-effect restrictions because it aggregates treated-cell effects with weights tied to untreated counterfactual exposure (Moreau-Kastler, 2025). When each granular proportional effect is positive, such an aggregation preserves the sign by construction. Theorem 3 formalizes this contrast in the same population environment: the pooled coefficient can be negative while is positive. Empirical conclusions should therefore distinguish the regression projection from the proportional treatment-on-the-treated estimand.
The four-cohort witness in Theorem 4 establishes the sign reversal with equal cohort shares, flat untreated time effects, and positive treated-cell effects. Staggered timing and heterogeneous proportional effects place a large positive effect in a cell with a negative residualized treatment value. This makes the phenomenon relevant for conventional staggered-adoption designs, including policy settings where treatment intensity or exposure plausibly varies across treated cohorts and periods.
The population characterization also clarifies the choice that precedes estimation. A researcher may intend a common log multiplier, a collection of granular proportional effects, or a positive-weight proportional aggregate; the pooled projection equals the first under homogeneity and can differ sharply from the third under heterogeneity. The primitive index establishes this distinction theoretically. Future work can turn that diagnostic into an operational procedure by estimating granular effects, constructing a sampling distribution for , assessing separation, and determining how often the sign-reversal region occurs in empirically calibrated designs (Santos Silva et al., 2006; Santos Silva et al., 2011; Correia et al., 2020; Wooldridge, 2023; Nagengast et al., 2025).
Appendices
Proofs, Auxiliary Lemmas, and Verification Note
This appendix records the two auxiliary population results used to connect the unit-level fixed-effect regression to the collapsed cohort-time analysis. The ordering mirrors the logic of the main text. The collapse result first justifies replacing the finite- unit fixed-effect projection by its cohort-time counterpart. The projection result then gives the first-order conditions for the limiting collapsed pseudo-true parameter. The main theorem proofs use these statements before establishing the primitive sign characterization in Theorem 3; the four-cohort sign-reversal result in Theorem 4 is then read as a consequence of that characterization together with the explicit witness construction.
The first auxiliary result formalizes the reduction from a unit fixed-effect PPML population criterion to a collapsed criterion. Its role is to show that, under the same cohort-share, untreated-mean, baseline-limit, proportional-effect, and rank restrictions used in Assumption 1, the treatment coordinate is invariant to within-cohort rearrangements of unit baselines that leave cohort averages unchanged. Thus the finite-array object in Definition 1 and the limiting object in Definition 3 are linked through the cohort-time means that drive the sign analysis.
Let be a number of periods and let be a finite cohort set. Suppose:
(Horizon.) .
(Cohort support.) The never-treated cohort belongs to , and every finite cohort in is dated after the first period.
(Array support.) For every and every unit , the cohort label belongs to .
(Positive primitives.) The unit baselines and the limiting within-cohort baselines are strictly positive, and each limiting cohort share satisfies .
(Share limits.) Assumption 1 holds for and .
(Untreated means.) Assumption 2 holds for , , , and .
(Baseline limits.) Assumption 3 holds for , , and .
(Proportional effects.) Assumption 4 holds for , , , and .
(Collapsed rank.) Assumption 5 holds for and .
Then, for every with , the finite-array unit-FE population criterion of Definition 1 has a unique global maximizer, the collapsed finite-array criterion has a unique global maximizer, and Moreover, Finally, for every with and every alternative positive baseline array , if then replacing by leaves the finite-array unit-FE coefficient unchanged:
⊢ LeanThe uniqueness clauses are population pseudo-likelihood statements in the usual misspecified-likelihood sense (White, 1982; Gouriéroux et al., 1984; Gouriéroux et al., 1984). The final invariance clause is specific to the cohort-time collapse used here: once unit effects absorb individual baseline levels, only the within-cohort baseline average enters the collapsed mean relevant for the treatment coordinate. This is why the later sign results can be stated in terms of cohort shares, untreated cohort-time means, and proportional-effect multipliers rather than the full list of unit baselines.
The second auxiliary result records the first-order conditions for the collapsed pseudo-true PPML projection. These score equations are the fixed-effect and treatment equations behind the derivative formula in Theorem 1 and the algebraic sign calculation in Theorem 3.
Fix a number of periods , a finite cohort set , cohort shares , positive limiting baseline means , time components , and proportional-effect log multipliers . Suppose the collapsed design rank condition in Assumption 5 holds for . Let , , , and be as in Definition 3. Then is the unique global maximizer of . Moreover, for every collapsed nuisance regressor index , and the treatment score also vanishes:
⊢ LeanThroughout, a supported cell is a pair with and , and a collapsed parameter is written with nuisance block and treatment coordinate , so that the collapsed linear index is The collapsed cell mass and observed cohort-time mean are Definition 3 defines the limiting criterion , the collapsed projection , and the fitted collapsed mean
Step 1: the supported-cell table is nonempty. Take the collapsed parameter , whose nuisance block is zero and whose treatment coordinate is one, so . If or is empty, then because the sum is empty; this contradicts Assumption 5, which requires the displayed quantity to be strictly positive at every nonzero collapsed parameter. Hence and is nonempty, and at least one supported cell exists.
Step 2: masses and observed means are strictly positive. For every supported cell, the first because and by Step 1, the second because and the exponential is positive.
Step 3: the collapsed design is injective. If two collapsed parameters satisfy at every supported cell, then satisfies at every supported cell by linearity of , whence By Assumption 5 this forces , i.e. . Thus the linear map from into the space of supported-cell tables is injective.
Step 4: existence and uniqueness of the maximizer. Written cellwise, so is a finite positive-mass, positive-mean Poisson criterion composed with the injective linear design of Step 3.
Two elementary cell facts drive the argument. First, for every and every real , which follows from for and from for .
Second, for every , every , and every pair , by strict convexity of the exponential.
The first display supplies the tail bound used for existence. Applying it cell by cell bounds above by a constant minus the norm of the weighted index table The map sending to this table is injective by Step 3 and the strict positivity from Step 2, so finite-dimensional norm equivalence bounds this table norm below by a positive multiple of . Hence as ; since is continuous, a global maximizer exists on a large enough closed ball. The second display gives the uniqueness step: if were two global maximizers, injectivity would give a supported cell with , and the strict inequality in that cell together with the weak inequality in the others would make the midpoint strictly better, a contradiction. Hence
Step 5: identification of the projection with the unique maximizer. By Definition 3, is a maximizer of . Step 4 shows that this maximizer is unique. Therefore for all , and implies . This is the asserted unique global maximality.
Step 6: every directional score vanishes. Fix a direction and consider . In each cell the index is affine in , and is differentiable with derivative . Differentiating the finite sum term by term at and using , By Step 5 the scalar function attains its maximum at , so its derivative there is zero:
Step 7: the two stated score equations. Take to be the unit vector in nuisance coordinate , so that ; Step 6 gives Take instead to be the pure treatment direction, with zero nuisance block and treatment coordinate one, so that ; Step 6 gives
∎Lemma 1 is the collapsed analogue of the pseudo-true score condition for a Poisson quasi-likelihood projection. The rank condition rules out flat directions in the fixed-effect and treatment regressors, while positivity of the cell means keeps the exponential criterion strictly concave on the relevant design span. The argument uses the mean criterion characteristic of standard pseudo-maximum-likelihood analysis (White, 1982; Gouriéroux et al., 1984).
Combining Proposition 1 and Lemma 1 fixes the population target before the sign results are applied. The derivative result in Theorem 1 uses the score equations to express the response of the treatment coordinate through the weighted residualized treatment. The homogeneous benchmark in Theorem 2 then identifies the correctly specified common-effect case. The primitive sign theorem, Theorem 3, comes next because it supplies the global sign criterion used to place the explicit four-cohort construction in . With that ordering, Theorem 4 becomes the concrete specialization of the sign frontier to the witness in Definition 6.
Reproducibility note.
The formal supplement verifies the displayed population claims in Lean 4 and records the correspondence between mathematical statements and declarations. Its deterministic scope covers the finite-dimensional projection, collapse, derivative-sign, primitive-sign, and four-cohort calculations. Sampling consistency, inference, and empirical identification of the primitive means and proportional effects form the natural empirical extension.
Proofs of the main results
Throughout, a collapsed parameter is written through the levels it induces, so that for every supported cell , , A unit parameter is written the same way, so that Both identities are the statement that the intercept, dummy, and treatment coordinates of the regressor pick out exactly these terms.
Write the finite-array collapsed primitives as and write the unit-time observed mean entering the criterion of Definition 1 as which is the form generated by the untreated mean of Assumption 2 together with the proportional effects of Assumption 4.
Step 0: standing positivity. The horizon assumption gives , and makes nonempty, so at least one supported cell exists. Fix with . Assumption 1 then supplies and for every ; every unit label lies in by the array-support hypothesis. Consequently, for every and every , the last two because they are products of positive baselines and exponentials.
Step 1: the collapsed designs are injective, and has a unique maximizer. If at every supported cell, then has at every supported cell, so , and Assumption 5 forces . Thus is injective.
With positive masses , positive means , and this injective design, the finite collapsed criterion is a finite positive-mass, positive-mean Poisson criterion, so it has exactly one global maximizer : each cell contribution is strictly concave and satisfies , which by injectivity makes continuous, coercive, and strictly concave.
Because maximizes and the criterion is differentiable, every collapsed directional score vanishes: for every ,
Step 2: lifting the collapsed maximizer to unit fixed effects. Define the unit parameter by (Any prescribed levels are realized by a unique choice of intercept, dummy, and treatment coordinates.) By the two level identities,
Exponentiating and using , which holds because and , each unit residual is the corresponding cell residual scaled by the baseline ratio:
Step 3: unit scores at the lift are collapsed scores. For every , the definition of gives
Given a unit direction , define the collapsed direction by its levels
Then Indeed, substituting the residual identity of Step 2 and grouping units by cohort (legitimate because every lies in ), the inner sum over at fixed is and the displayed baseline-ratio identity turns that bracketed sum into , while .
Applying the vanishing collapsed score of Step 1 to , we conclude that every unit directional score at vanishes:
Step 4: the lift maximizes the unit criterion. For a Poisson criterion with nonnegative weights, vanishing of all directional scores is sufficient for global maximality: because the tangent-line bound makes each cell increment at most its linearization, and the sum of the linearizations is the score in the direction , which is zero by Step 3.
Step 5: the unit design is injective. Suppose for every unit and every period . For each choose a representative unit ; this is possible because . Define a collapsed parameter by , , . Then for every supported cell, using . Hence , so Assumption 5 gives ; in particular and for every . Evaluating in the first period and substituting and gives for every unit . All levels of vanish, so , and the unit design is injective.
Step 6: unique unit maximizer and equality of treatment coordinates. The unit criterion is a finite Poisson criterion with the constant positive weights , positive means , and, by Step 5, an injective design; exactly as in Step 1 it therefore has a unique global maximizer. By Step 4 that maximizer is . Hence and each have a unique global maximizer, and i.e. is the treatment coordinate of the unique maximizer of .
Step 7: convergence to the limiting collapsed projection. Fix a supported cell . By Assumption 1, and by Assumption 3, since does not depend on , Write the finite set of supported cells as , put , , and write for the supported-cell linear predictor. The preceding convergence is cellwise convergence of to the strictly positive array , and the map is injective by Step 1. Define Since every is positive, is injective. Finite-dimensional norm equivalence therefore gives a constant such that For all sufficiently large , uniformly over supported cells, Set The one-cell bound , combined with the displayed eventual coefficient bounds, gives for every and all sufficiently large , The same coefficient bounds give . Since maximizes , and therefore eventually. Thus the finite maximizers are eventually contained in the compact ball .
Because there are finitely many supported cells, the cellwise convergence of also gives uniform convergence on this compact set, and is continuous there. By Lemma 1, the limiting criterion has the unique global maximizer , hence the unique maximizer on . The compact-containment bound, compact-uniform convergence, eventual global maximality of , and this uniqueness yield
Taking the treatment coordinate, which is a continuous function of the parameter, gives ; since for every by Step 6,
Step 8: invariance to within-cohort baseline rearrangement. Fix and a positive baseline array with for every . The finite collapsed observed mean depends on the baselines only through the within-cohort average, and the masses do not involve the baselines at all. Hence as functions of , so their unique maximizers coincide. Applying Step 6 once with and once with ,
∎Write , , and A collapsed parameter is written with , and is the weighted nuisance coefficient of Definition 4, so that
Step 1: . Each summand is nonnegative, since , , and . If , every summand vanishes, and positivity of forces at every supported cell. Consider the collapsed parameter whose treatment coordinate is one, so . Its collapsed index is exactly the residual, whence , contradicting Assumption 5. Thus .
Step 2: the perturbed criterion. For let be with the coordinate replaced by , and write . Because and , the perturbation moves exactly one collapsed mean: Setting
Step 3: the score system and its solution branch. For a direction put By Lemma 1 applied at the effect array , the criterion has a unique global maximizer, and the nuisance and treatment score equations imply by linearity that for every direction at that maximizer. Conversely, vanishing of every directional score is sufficient for global maximality of a Poisson criterion with nonnegative weights, by the tangent-line bound .
The system is smooth in , it is solved at , and its derivative in at that point is the bilinear form which is negative definite: if , the injectivity of , obtained from Assumption 5 as in Lemma 1, gives a supported cell with . Hence the derivative is invertible. The implicit function theorem therefore supplies a continuously differentiable defined near , with and for every ; by the characterization above and uniqueness of the maximizer, for near . In particular is differentiable at , with derivative the treatment coordinate of .
Differentiating at and using the one-cell derivative display above gives the linearized score: for every direction ,
Step 4: solving the linearized score by weighted FWL. Put and define the mean-normalized one-cell source and the nuisance part of by so that and lies in the span of the nuisance columns . Because and , each summand of the linearized score factors as so dividing the whole linearized score by rewrites it as Taking to be the pure treatment direction gives ; taking for arbitrary gives for every in the nuisance span.
The coefficient in Definition 4 minimizes a finite weighted least-squares projection of on the nuisance span. Its normal equations give Consequently , and, since also lies in the nuisance span, . Combining the score normal equations with in the nuisance span gives Expanding this last display and using the projection orthogonality identities yields
Now by the positivity of shown above, and, since is supported on the single cell , Dividing,
Step 5: the derivative formula. Combining the differentiability branch with the FWL calculation, which is the asserted differentiability at together with the displayed derivative.
Step 6: the sign equivalences. The derivative is with because , , , and as shown above. Multiplication by a positive scalar preserves and reflects each of the three order relations, so which, with , are exactly the three asserted equivalences.
∎Throughout, , , , , and , so that Treatment is absorbing from the adoption date on, so exactly when , and ; in table form, with rows and columns , A collapsed parameter is written through its levels: the intercept , the cohort dummies for , the period dummies for , and the treatment coordinate , so that with the convention , the period-1 dummy being absent.
Step 1: the collapsed design has full rank. Let satisfy . Every summand is nonnegative and every is nonzero, so Evaluate this at four families of cells, using the displayed level formula and the treatment table. At : all dummies vanish and , so . At for : the period-1 dummy vanishes and , so . At for : the cohort dummy vanishes and , so . Finally at , where : all dummies and the intercept vanish, so . Hence , which is Assumption 5 for this design.
Step 2: region membership. We verify each clause of Definition 5. The horizon clause is . The support clause holds because and the finite cohorts all adopt after the first period; those same three cohorts together with give the frontier scope in Assumption 6. The shares satisfy The untreated-mean clause is witnessed by , since in every supported cell. The positive-effect clause is Step 3 below, and the rank clause is Step 1. It remains to check .
Evaluate the polynomial of Definition 7 at the present primitives. Write for the common treated multiplier off cell and for the multiplier at . Since and , the components are again with rows and columns , so that The treated cells are , so which expands to , while . Subtracting,
The collapsed sign comparison used here is the beta-zero frontier-elimination comparison underlying Theorem 3. In this specialization, let be the fixed-effect parameter whose fitted cell mean is The nuisance scores at vanish by the row and column identities defining , , and . The treatment score at in the limiting collapsed criterion of Definition 3 is Here , since and every is positive. The weights and means are positive, and the rank calculation above makes the collapsed design map injective. If is the maximizer in Definition 3 and , the score at is zero. Whenever , strict monotonicity of the exponential along the nonzero fitted-index direction gives If , then all score coordinates vanish at , so is the unique collapsed maximizer and . Therefore For the present primitives , hence .
Since the primitive tuple entering Definition 5 is exactly the restriction of to the supported and treated cells, its treatment coordinate is this same . All clauses hold, so .
Step 3: strict positivity of the treated effects. If is supported with , then , and both are strictly positive because and .
Step 4: cell has the strictly largest treated effect. The cell is treated, since , and . Every other treated supported cell has , and with the logarithm strictly increasing on the positive reals gives
Step 5: the weighted FWL residual table at the no-effect vector. Take the effect vector identically zero. Theorem 2 applies to it with common treated multiplier : its common-treated-effect hypothesis holds because every effect is zero, its rank hypothesis is Step 1, and its untreated-mean hypothesis is discharged by the auxiliary array in which every unit’s untreated and treated outcomes are deterministically , with unit baselines and the time components already fixed here, so that with the normalization . Its conclusion gives, for every supported cell,
Hence the projection weights of Definition 4 are
Let denote the entry of the displayed table, with rows and columns . Two facts about are checked by direct summation of these sixteen numbers:
First, lies in the fixed-effect nuisance span. Indeed, with one checks cell by cell that and the right-hand side is the value at of the nuisance combination with intercept coefficient , cohort-dummy coefficients (with absorbed by the intercept), and period-dummy coefficients (with ); every such additive intercept-plus-cohort-plus-period array is a linear combination of the columns .
Second, is orthogonal to that span under the weights . Each nuisance regressor is either the constant , the indicator of one cohort, or the indicator of one period; because the weights are the same constant in every cell, the corresponding weighted inner products are , , and , all zero by the row and column sums above; orthogonality extends from these spanning columns to their whole span by linearity:
The two facts together say that is the -weighted orthogonal projection of onto the nuisance span, which by Definition 4 means that the weighted FWL residual is the remainder:
Step 6: the late-cell residual. Reading off the entry in row , column of ,
Step 7: the exact derivative at the no-effect vector. Since every weight is and , the residual energy of Theorem 1 is because each row of consists of two entries of absolute value and two of absolute value , contributing per row and in total.
The cell is treated, and Assumption 5 holds by Step 1, so Theorem 1 applies at and gives, with , , ,
Step 8: a neighborhood of the no-effect vector. The map is continuous at the zero array: the collapsed projection , hence each fitted mean , depends continuously on because it is the unique maximizer of a criterion whose finitely many positive means move continuously with ; under Assumption 5 the weighted nuisance Gram matrix built from those weights is invertible, and the residual is a continuous rational function of its entries.
Since and the cells are finite in number, continuity supplies such that Now take any effect array with for every and and with strictly positive treated effects in the sense of Assumption 7, and define its supported truncation Then at every cell, so . Moreover and agree on all supported cells, and the criterion – hence , the fitted means, the weights, and the residual – involves only supported cells, so
Applying Theorem 1 once more at the treated cell , now at the effect array , the derivative equals with a strictly positive prefactor and denominator, so it has the sign of :
∎For a collapsed parameter and a supported cell , write the collapsed linear index as so that, by Definition 3, The map is linear, so for all .
Step 1: the candidate parameter. Let be the never-treated reference cohort and define coordinatewise by with treatment coordinate .
Assumption 2 carries the normalization . Hence, for every period , since for all period dummies are switched off and , while for exactly the dummy of period contributes. Likewise, for every , because exactly the dummy of cohort contributes when , and none when .
Step 2: the index at . Adding the intercept, cohort, period, and treatment contributions in the two cases and gives, for every and every period ,
Step 3: the homogeneous-effect hypothesis in product form. For every and every , if the common treated-cell hypothesis gives , and if both sides vanish.
Step 4: the criterion is exactly specified at . Using together with Steps 2 and 3, the observed collapsed mean satisfies In particular and
Step 5: the scalar Poisson cell inequality. For every and every real , with strict inequality when . Apply , strict for , at , multiply by , and rearrange using .
Step 6: the cell masses are strictly positive. First : if , then for the nonzero collapsed parameter with zero nuisance block and treatment coordinate one the weighted sum is an empty sum, hence zero, contradicting Assumption 5. Since , it follows that
Step 7: is a global maximizer. Fix any . Applying Step 5 cell by cell with and , and using from Step 4, Multiplying by and summing over supported cells,
Step 8: the maximizer is unique. Suppose and . Then , so Assumption 5 gives and by linearity , so some supported cell has At that cell the strict half of Step 5 applies, and all other cells obey the weak inequality of Step 7. Since , summing gives contradicting the assumed equality. Hence .
Step 9: identification of the projection. By Steps 7 and 8 the criterion has as its unique global maximizer, and Definition 3 selects among the global maximizers, so
Step 10: the two conclusions. The treatment coordinate of is , so and for every and every period , by Step 2,
∎Write and , with . Then the quantities of Definition 7 are A collapsed parameter is written , with nuisance block and treatment coordinate , and . Finally set
Step 1: positivity of the primitive margins. Since , we have , and since , the support is nonempty. Every is a product of strictly positive factors. Hence
Also, summing in the two possible orders gives
Step 2: the zero-treatment row-column fit. Define the collapsed parameter , whose treatment coordinate is zero and whose nuisance levels are Adding the intercept, cohort, and period contributions, with the cohort dummy absent for and the period dummy absent for , gives, for every supported cell, where the logarithms and division use the positivity in Step 1.
Step 3: the conditional residuals and their margins. Define the mass-weighted residual of in cell by
Using and from Step 1, its row and column sums vanish:
Step 4: the scores at . Every collapsed nuisance regressor is either the constant , the indicator of one cohort , or the indicator of one period . The corresponding nuisance score at is therefore , , or , and these vanish by Step 3: Because the nuisance block is finite-dimensional and is the corresponding linear combination of the coordinate regressors, the same conclusion holds for every nuisance direction :
The treatment score at is and by Step 1.
Step 5: the pseudo-true coefficient has the sign of . The finite-dimensional Poisson sign comparison used here is the following. For a finite index set with strictly positive weights , strictly positive means , and an injective linear design map , suppose that a zero-treatment point clears every nuisance-direction score, If and is the treatment coordinate of the unique maximizer of , then Indeed, the score at a maximizer vanishes in every direction. Writing for the maximizer, , and , the score identities and the nuisance-direction condition give The summands on the right are nonnegative, and at least one is strictly positive when , by injectivity of and strict monotonicity of the exponential. If , all scores at vanish, and the tangent-line inequality for the exponential makes globally maximizing; uniqueness gives the zero treatment coordinate.
Apply this comparison with index set , design , and . The weights and means are positive by Step 1, Step 4 supplies the nuisance-direction score condition, and Assumption 5 gives injectivity: if two collapsed parameters have the same fitted-index array, their difference has weighted sum of squared indices equal to zero, so the rank condition forces the difference to be zero. Therefore the treatment coordinate of has the same three-way sign as . Since with ,
Step 6: the observed baseline proxy reproduces the untreated mean. For every supported cohort , the first-period treatment indicator is zero: if this is the never-treated path, and if is finite then would force the adoption date to be the first period, which the support excludes. Also for every , and Assumption 2 gives . Hence, for every and every , where justifies the cancellation.
Step 7: is nonempty and . Assumption 6 supplies a supported finite cohort adopting in the second period; in that period its treatment indicator equals one, so is nonempty. For every , , and Step 6 gives . Hence every summand of is nonnegative and at least one summand is strictly positive:
Step 8: membership in is equivalent to . If the induced primitive tuple lies in , then the last clause of Definition 5 gives negativity of its treatment coordinate. In the restricted primitive coordinates the limiting criterion is the same collapsed criterion , so this coordinate equals , and Step 5 gives .
Conversely suppose . The horizon, support, and frontier-scope clauses of Definition 5 are hypotheses of the present theorem. The simplex clause holds because the finite-array shares sum to one for every , every unit label lies in , and Assumption 1 gives after passage to the limit.
The untreated-mean clause is witnessed by , since the primitive untreated coordinate is . The positive-effect clause is Assumption 7. The rank clause is Assumption 5, after rewriting the sum over supported cells as the double sum over and . Finally Step 5 gives , which is the tuple’s treatment coordinate. Hence the tuple lies in .
Step 9: the four-cohort polynomial. Let , , and consider the four-period support with , equal shares , unit limiting baselines, zero untreated time effects, multiplier on every treated cell except , and multiplier at . With treatment absorbing from the adoption date on, exactly when and , so with rows and columns . Summing rows and columns, The treated cells of row are , those of row are , and the treated cell of row is , so which expands to , while Subtracting gives
Step 10: the numerical identity. With and ,
Step 11: the value at the witness. For of Definition 6, the treated multipliers are off cell and at , so Step 9 applies with and :
Step 12: the counterfactual-share objects on treated cells. Let , so . Step 6 gives , and , so Moreover , , and by Step 7, hence
Step 13: the weights sum to one. By the definition of and the positivity ,
Step 14: the ratio form of . Expanding the definition and using on every treated cell, The second fraction equals , so
Step 15: the sign contrast. Suppose . Step 5 gives . For , each treated cell has by Step 12, and Assumption 7 gives , hence and . The sum defining has nonnegative terms and, by Step 7, at least one strictly positive term. Therefore , and
∎References
- Moreau-Kastler, Ninon (2025). Proportional Treatment Effects in Staggered Settings: An Approach for Poisson Pseudo-Maximum Likelihood. . doi
- Wooldridge, Jeffrey M (2023). Simple approaches to nonlinear difference-in-differences with panel data. The Econometrics Journal. doi
- Goodman-Bacon, Andrew (2021). Difference-in-Differences with Variation in Treatment Timing. Journal of Econometrics. doi
- de Chaisemartin, Cl{\'e}ment and D'Haultfoeuille, Xavier (2020). Two-Way Fixed Effects Estimators with Heterogeneous Treatment Effects. American Economic Review. doi
- White, Halbert (1982). Maximum Likelihood Estimation of Misspecified Models. Econometrica. doi
- Andrews, Donald W. K. (1988). Laws of Large Numbers for Dependent Non-Identically Distributed Random Variables. Econometric Theory. doi
- Nagengast, Arne J. and Yotov, Yoto V. (2025). Staggered Difference-in-Differences in Gravity Settings: Revisiting the Effects of Trade Agreements. American Economic Journal: Applied Economics. doi
- Ashenfelter, Orley and Card, David (1985). Using the Longitudinal Structure of Earnings to Estimate the Effect of Training Programs. The Review of Economics and Statistics. doi
- Abadie, Alberto (2005). Semiparametric Difference-in-Differences Estimators. The Review of Economic Studies. doi
- Athey, Susan and Imbens, Guido W. (2006). Identification and Inference in Nonlinear Difference-in-Differences Models. Econometrica. doi
- Imbens, Guido W. and Wooldridge, Jeffrey M. (2009). Recent Developments in the Econometrics of Program Evaluation. Journal of Economic Literature. doi
- Angrist, Joshua D. and Pischke, J{\"o}rn-Steffen (2009). Mostly Harmless Econometrics: An Empiricist's Companion. Princeton University Press.
- Callaway, Brantly and Sant'Anna, Pedro H. C. (2021). Difference-in-Differences with Multiple Time Periods. Journal of Econometrics. doi
- Sun, Liyang and Abraham, Sarah (2021). Estimating Dynamic Treatment Effects in Event Studies with Heterogeneous Treatment Effects. Journal of Econometrics. doi
- Borusyak, Kirill and Jaravel, Xavier and Spiess, Jann (2024). Revisiting Event-Study Designs: Robust and Efficient Estimation. The Review of Economic Studies. doi
- Sant'Anna, Pedro H. C. and Zhao, Jun (2020). Doubly Robust Difference-in-Differences Estimators. Journal of Econometrics. doi
- Jonathan Roth and Pedro H. C. Sant'Anna (2020). When Is Parallel Trends Sensitive to Functional Form?. Econometrica. doi
- Roth, Jonathan and Sant'Anna, Pedro H. C. and Bilinski, Alyssa and Poe, John (2023). What's Trending in Difference-in-Differences? A Synthesis of the Recent Econometrics Literature. Journal of Econometrics. doi
- Gardner, John (2022). Two-Stage Differences in Differences. . arXiv
- Baker, Andrew and Callaway, Brantly and Cunningham, Scott and Goodman-Bacon, Andrew and Sant'Anna, Pedro H. C. (2025). Difference-in-Differences Designs: A Practitioner's Guide. . arXiv
- Gouri{\'e}roux, Christian and Monfort, Alain and Trognon, Alain (1984). Pseudo Maximum Likelihood Methods: Theory. Econometrica. doi
- Gouri{\'e}roux, Christian and Monfort, Alain and Trognon, Alain (1984). Pseudo Maximum Likelihood Methods: Applications to Poisson Models. Econometrica. doi
- Hausman, Jerry and Hall, Bronwyn H. and Griliches, Zvi (1984). Econometric Models for Count Data with an Application to the Patents-R \& D Relationship. Econometrica. doi
- Cameron, A. Colin and Trivedi, Pravin K. (2013). Regression Analysis of Count Data. Cambridge University Press. doi
- Santos Silva, J. M. C. and Tenreyro, Silvana (2006). The Log of Gravity. The Review of Economics and Statistics. doi
- Santos Silva, J. M. C. and Tenreyro, Silvana (2011). Further Simulation Evidence on the Performance of the Poisson Pseudo-Maximum Likelihood Estimator. Economics Letters. doi
- Correia, Sergio and Guimarães, Paulo and Zylkin, Tom (2020). Fast Poisson estimation with high-dimensional fixed effects. The Stata Journal. doi
- Anderson, James E. and van Wincoop, Eric (2003). Gravity with Gravitas: A Solution to the Border Puzzle. American Economic Review. doi
- Head, Keith and Mayer, Thierry (2014). Gravity Equations: Workhorse,Toolkit, and Cookbook. Handbook of International Economics. doi
- Yotov, Yoto V. and Piermartini, Roberta and Monteiro, Jos{\'e}-Antonio and Larch, Mario (2016). An Advanced Guide to Trade Policy Analysis: The Structural Gravity Model. World Trade Organization and United Nations Conference on Trade and Development.
Comments on earlier versions
Anchored to: