CausalSmith · seminar slides

Minimax ATE Estimation with Many Discrete Cells

With fixed overlap and a known heterogeneity radius, we give finite-sample minimax benchmarks for average treatment effect estimation and construct clipped estimators that match the lower benchmark in the main regimes.

Overview

  • Target: scalar average treatment effect with a finite discrete adjustment variable.
  • Challenge: many cells may be sparsely populated and unevenly weighted.
  • Organizing parameter: σ\sigma, the heterogeneity radius for cell treatment effects.
  • Main output: a finite-sample minimax bracket indexed by n,d,M,σn,d,M,\sigma.
  • Constructive side: a known-radius selector over polynomial, collision, and zero branches.
  • Converse side: binary hard experiments transported into the same real-outcome class.

informal · Theorem T-14 Under fixed overlap, bounded outcome scale, and radius 0σ20\le\sigma\le2, the minimax mean-squared risk lies between the stated lower benchmark and at most the stated selector benchmark.

Motivation

  • Think of adjustment cells as age-by-income-by-region-by-baseline-history groups.
  • Identification says we should compare treated and control outcomes within each cell.
  • Finite samples create a second problem: some cells contain treated observations, controls, or neither.
  • When dd is large, the obstacle is combinatorial as much as statistical.
  • The estimand is still one scalar ATE, but the data arrive through many sparse cells.

Research Question

  • Suppose treatment is ignorable within a finite cell X{1,,d}X\in\{1,\ldots,d\}.
  • Suppose every supported cell has fixed overlap ϵπk1ϵ\epsilon\le\pi_k\le1-\epsilon.
  • Suppose outcomes are real-valued, centered on scale MM, with conditional second moments at most M2M^2.
  • Suppose cell effects differ from the ATE by at most σM\sigma M.
  • What is the best finite-sample mean-squared error as n,d,σn,d,\sigma vary?

Setup

  • pkp_k is the cell mass P(X=k)P(X=k).
  • πk\pi_k is the cell propensity P(A=1X=k)P(A=1\mid X=k).
  • μak\mu_{ak} is the conditional mean outcome in arm aa and cell kk.
  • τk=μ1kμ0k\tau_k=\mu_{1k}-\mu_{0k} is the cell treatment effect.
  • τ(P)=kpkτk\tau(P)=\sum_k p_k\tau_k is the average treatment effect.
  • δk=τkτ(P)\delta_k=\tau_k-\tau(P) is the cell-effect deviation.
Definition P-4 (Average treatment effect \(\tau(P)\))

For a real-outcome finite-cell law PP indexed by k=1,,dk=1,\ldots,d and an overlap parameter ϵ\epsilon with 0<ϵ<1/20<\epsilon<1/2, assume consistency, conditional exchangeability of (Y(0),Y(1))(Y(0),Y(1)) and AA given XX for measurable potential-outcome events, overlap ϵπk1ϵ\epsilon\le \pi_k\le 1-\epsilon whenever pk>0p_k>0, and first-moment integrability of each arm-cell outcome law for every a{0,1}a\in\{0,1\} and positive-mass cell. With pkp_k the cell mass and μak\mu_{ak} the arm-cell conditional outcome mean, define τ(P):=k=1dpk(μ1kμ0k). \tau(P) := \sum_{k=1}^{d}p_k(\mu_{1k}-\mu_{0k}).

Assumptions

  • Consistency and conditional exchangeability give the observed-data causal interpretation.
  • Fixed overlap keeps both treatment arms available in every supported cell.
  • Mean normalization fixes the outcome scale MM.
  • The second-central-moment bound allows real outcomes with variance control.
  • Approximate homogeneity says the largest supported-cell deviation is at most σM\sigma M.
Assumption A-3 (Fixed overlap)

For each cell kk, let pk:=P(X=k), p_k:=P(X=k), and, when pk>0p_k>0, let πk:=P(A=1X=k). \pi_k:=P(A=1\mid X=k). For every kk with pk>0p_k>0, ϵπk1ϵ. \epsilon\le \pi_k\le 1-\epsilon .

Assumption A-5 (Approximate homogeneity)

For every cell kk, let τk:=μ1kμ0k \tau_k:=\mu_{1k}-\mu_{0k} and δk:=τkjpjτj, \delta_k:=\tau_k-\sum_j p_j\tau_j , both well defined on zero-mass cells as well. The cell-effect deviations satisfy maxk:pk>0δkσM. \max_{k:p_k>0}|\delta_k|\le \sigma M .

Model Class

  • We work on one radius-indexed real-outcome model class.
  • The class allows arbitrary cell masses.
  • The radius σ[0,2]\sigma\in[0,2] moves from exact homogeneity to unrestricted heterogeneity on the normalized scale.
  • The minimax criterion asks for worst-case squared error over this class.
Definition P-1 (Real-outcome model class \(\mathcal P_{d,\epsilon,M,\sigma}\))

Pd,ϵ,M,σ:={P:X{1,,d}, A{0,1}, Y,Y(0),Y(1)R,0<ϵ<1/2,M1,0σ2,P satisfies ass:consistency, ass:conditional-exchangeability, ass:overlap, ass:mean-normalization, ass:second-central-moment, ass:approximate-homogeneity}. \mathcal P_{d,\epsilon,M,\sigma} := \left\{ P: \begin{array}{l} X\in\{1,\ldots,d\},\ A\in\{0,1\},\ Y,Y(0),Y(1)\in\mathbb R,\\ 0<\epsilon<1/2,\quad M\ge 1,\quad 0\le \sigma\le 2,\\ P\text{ satisfies }\text{ass:consistency, ass:conditional-exchangeability, ass:overlap, ass:mean-normalization, ass:second-central-moment, ass:approximate-homogeneity} \end{array} \right\}.

Definition P-9 (Minimax risk \(\mathsf R_{n,d,\epsilon,M,\sigma}\) and estimator class \(\mathcal T_{n,M}\))

Let Tn,M:={T:On[M,M]  :  T is measurable}. \mathcal T_{n,M} := \bigl\{ T:\mathcal O^n\to[-M,M]\;:\;T\ \text{is measurable} \bigr\}. The minimax mean-squared risk over Pd,ϵ,M,σ\mathcal P_{d,\epsilon,M,\sigma} is Rn,d,ϵ,M,σ:=infTTn,MsupPPd,ϵ,M,σEQP(n) ⁣[(Tτ(P))2]. \mathsf R_{n,d,\epsilon,M,\sigma} := \inf_{T\in\mathcal T_{n,M}} \sup_{P\in\mathcal P_{d,\epsilon,M,\sigma}} E_{Q_P^{(n)}}\!\left[(T-\tau(P))^2\right].

Related Literature

  • Rubin (1974), Rubin (1979), and Rosenbaum and Rubin (1983) supply the potential-outcomes adjustment logic.
  • Hahn (1998) and Hirano et al. (2003) give classical efficiency benchmarks for ATE estimation.
  • Belloni et al. (2017), Chernozhukov et al. (2018), and Kennedy (2022) study high-dimensional nuisance estimation.
  • Wager and Athey (2018), Nie and Wager (2021), and Semenova and Chernozhukov (2021) focus on heterogeneous-effect learning.
  • Zeng et al. (2024) identify the closest sparse discrete-adjustment collision phenomena for binary outcomes.
  • Jiao et al. (2015) and Wu and Yang (2016) supply the large-alphabet polynomial-estimation toolkit.

Key Idea

  • Sparse cells create two useful regimes.
  • Rare-cell regime: estimate a large-alphabet functional with a polynomial approximation.
  • Crossed-cell regime: use cells where both treatment arms appear and borrow through the homogeneity radius.
  • The known radius σ\sigma tells us which branch has the better benchmark.
  • The zero branch handles fully saturated scales through clipping.
Observed data Pilot split classify cells: heavy vs light Heavy cells plug-in arm contrasts Light cells Chebyshev factorial device Polynomial branch heavy + light parts summed only the light part is polynomial u ≍ d²/(n²log²(en)) Collision branch crossed cells, occupancy weights h ≍ σ²+d/n² Zero branch constant 0, uses no data Known-radius selector compares u, h, 1 via σ each branch already in [-M, M]
illustrative Box-and-arrow schematic showing observed data flowing into heavy-light classification, polynomial estimation, collision estimation, zero estimation, and a known-radius selector.

Estimators

  • Heavy cells: use plug-in treated-control contrasts on the estimation split.
  • Light cells: use a signed Chebyshev factorial device at degree KnK_n.
  • Crossed cells: average treated-control contrasts over cells where both arms are observed.
  • Selector: compare un,d=d2/[n2log2(en)]u_{n,d}=d^2/[n^2\log^2(en)], hn,d,σ=σ2+d/n2h_{n,d,\sigma}=\sigma^2+d/n^2, and 11.
  • All branches are clipped to [M,M][-M,M].

informal · Theorem T-10 The polynomial branch has risk at most the parametric term plus the rare-cell polynomial remainder, and the collision branch has risk at most the parametric term plus σ2+d/n2\sigma^2+d/n^2.

Definition P-7 (Selector \(\widehat\tau_n^{\star}\))

Given candidate estimators TnpolyT_n^{\mathrm{poly}} and TncolT_n^{\mathrm{col}}, each a total measurable map On[M,M]\mathcal O^n\to[-M,M], define τ^n\widehat\tau_n^{\star} by the following steps.

  1. Define the rare-cell polynomial remainder un,d:=d2n2log2(en). u_{n,d}:=\frac{d^2}{n^2\log^2(en)}.
  2. Define the collision remainder hn,d,σ:=σ2+dn2. h_{n,d,\sigma}:=\sigma^2+\frac{d}{n^2}.
  3. Output the measurable selector τ^n:={Tncol,hn,d,σmin{1,un,d},Tnpoly,un,d<min{1,hn,d,σ},0,otherwise. \widehat\tau_n^{\star}:= \begin{cases} T_n^{\mathrm{col}},&h_{n,d,\sigma}\le\min\{1,u_{n,d}\},\\ T_n^{\mathrm{poly}},&u_{n,d}<\min\{1,h_{n,d,\sigma}\},\\ 0,&\text{otherwise}. \end{cases}

Upper Bound

  • The selector inherits the better available branch.
  • Polynomial estimation pays the rare-cell scale.
  • Collision estimation pays the heterogeneity radius and the crossed-cell occupancy scale.
  • The constant-zero branch protects the all-alphabet statement — validity at every alphabet size dd, not just a restricted dimension range.

informal · Theorem T-11 The known-radius selector has worst-case mean-squared error at most CϵM2rn,d,σC_\epsilon M^2 r_{n,d,\sigma} for all n,d1n,d\ge1, M1M\ge1, and 0σ20\le\sigma\le2.

Theorem T-11 (All-alphabet frontier upper bound)

For every overlap parameter ϵ\epsilon with 0<ϵ<1/20<\epsilon<1/2, there exist a constant Cϵ>0C_{\epsilon}>0 and a polynomial calibration handle (Nϵ,ρϵ)(N_{\epsilon},\rho_{\epsilon}) such that:

  • (Handle calibration.) ρϵ>0\rho_{\epsilon}>0, and for every pair of integers n,d1n,d\ge1, Nϵn,dρϵnlog(en)2Kn, N_{\epsilon}\le n,\qquad d\le \rho_{\epsilon} n\log(en) \quad\Longrightarrow\quad 2\le K_n , where KnK_n is the polynomial degree used by the polynomial estimator.
  • (Index range.) The guarantee applies to every pair of integers n,d1n,d\ge1, every M1M\ge1, and every 0σ20\le\sigma\le2.
  • (Model class.) The law PP ranges over Pd,ϵ,M,σ\mathcal P_{d,\epsilon,M,\sigma} from Definition P-1.

For these indices, let TnpolyT_n^{\mathrm{poly}} be the admissible polynomial estimator calibrated by (Nϵ,ρϵ)(N_{\epsilon},\rho_{\epsilon}), let TncolT_n^{\mathrm{col}} be the collision estimator, and let τ^n\widehat\tau_n^{\star} be the known-radius selector in Definition P-7 applied with radius σ\sigma to these two estimators, and write rn,d,σ=1n+min{1,min[d2n2log2(en),σ2+dn2]} r_{n,d,\sigma} = \frac{1}{n} + \min\left\{ 1,\, \min\left[ \frac{d^2}{n^2\log^2(en)},\, \sigma^2+\frac{d}{n^2} \right] \right\} for the selector benchmark. Then τ^n\widehat\tau_n^{\star} is measurable, satisfies τ^n(s)[M,M]\widehat\tau_n^{\star}(s)\in[-M,M] for every sample ss, and supPPd,ϵ,M,σEQP(n) ⁣[(τ^nτ(P))2]CϵM2rn,d,σ. \sup_{P\in\mathcal P_{d,\epsilon,M,\sigma}} E_{Q_P^{(n)}}\!\left[\left(\widehat\tau_n^{\star}-\tau(P)\right)^2\right] \le C_{\epsilon}M^2 r_{n,d,\sigma}.

Lower Bound

  • The converse uses binary sparse-cell experiments as source problems.
  • An affine map preserves cell masses and propensities while putting outcomes on scale MM.
  • An exact-homogeneity source gives the collision baseline.
  • A radius channel attenuates binary contrasts by σ/2\sigma/2 and fits the radius constraint.
  • Data processing transfers the binary testing difficulty through the channel.
  • The two components combine inside the same real-outcome model class.
Exact source binary homogeneity collision baseline Affine embedding real outcomes scale M Exact lower baseline lower bound data processing Radius source binary alternatives Attenuation σ/2 channel data processing Radius family real outcomes inside σM radius Radius lower channel lower bound
illustrative Box-and-arrow schematic showing exact binary sources and radius-channel binary sources transported into real-outcome lower bounds.

informal · Theorem T-5 Binary source classes embed into the real-outcome classes, transferring exact-homogeneity and radius-channel lower bounds on the MM scale.

informal · Theorem T-9 For every allowed n,d,M,σn,d,M,\sigma, the minimax risk is at least the all-alphabet lower benchmark with exact-homogeneity and radius-channel components.

Theorem T-9 (All-d radius converse)

For every 0<ϵ<1/20<\epsilon<1/2, there exist constants cϵc_{\epsilon} and bϵradb_{\epsilon}^{\mathrm{rad}} such that 0<cϵ1,0<bϵrad. 0<c_{\epsilon}\le 1, \qquad 0<b_{\epsilon}^{\mathrm{rad}}. For all integers n,d1n,d\ge 1, all M1M\ge 1, and all 0σ20\le \sigma\le 2, the minimax mean-squared risk Rn,d,ϵ,M,σ\mathsf R_{n,d,\epsilon,M,\sigma} of Definition P-9 satisfies cϵM2{1n+min(1,dn2)+σ2min(1,d2n2log2(en))}Rn,d,ϵ,M,σ. c_{\epsilon}M^2 \left\{ \frac1n+\min\left(1,\frac d{n^2}\right) +\sigma^2\min\left(1,\frac{d^2}{n^2\log^2(en)}\right) \right\} \le \mathsf R_{n,d,\epsilon,M,\sigma}. Moreover, for the same constants and every such n,d,M,σn,d,M,\sigma, there are a capped alphabet size dradd_{\mathrm{rad}}, a source family, a padding map into {1,,d}\{1,\ldots,d\}, an embedding into Pd,ϵ,M,σ\mathcal P_{d,\epsilon,M,\sigma}, and a coupling realizing the radius-channel construction with source constant bϵradb_{\epsilon}^{\mathrm{rad}}. Along this realization, for all source laws P0,P1P_0,P_1, τ(P1)τ(P0)=Mσ2{ψ(P1)ψ(P0)}, \tau(P_1)-\tau(P_0) = M\frac{\sigma}{2} \{\psi(P_1)-\psi(P_0)\}, and for every estimator τ^\widehat\tau taking values in [M,M][-M,M], some source law PP satisfies cϵM2σ2min(1,d2n2log2(en))EP{(τ^τ(P))2}. c_{\epsilon}M^2\sigma^2 \min\left(1,\frac{d^2}{n^2\log^2(en)}\right) \le \mathbb E_P\{(\widehat\tau-\tau(P))^2\}.

Main Result

  • The lower side adds two irreducible costs: crossed-cell scarcity under homogeneity and radius-sensitive rare-cell difficulty.
  • The upper side takes the best of polynomial estimation, collision estimation, and the zero branch.
  • The bracket is stated over one same-index model class.
Theorem T-14 (All-alphabet minimax bracket)

For every overlap parameter ϵ\epsilon with 0<ϵ<1/20<\epsilon<1/2, there are real constants cϵc_{\epsilon} and CϵC_{\epsilon} such that 0<cϵ1,1Cϵ,cϵCϵ. 0<c_{\epsilon}\le 1,\qquad 1\le C_{\epsilon},\qquad c_{\epsilon}\le C_{\epsilon}. For every n,dNn,d\in\mathbb N, MRM\in\mathbb R, and σR\sigma\in\mathbb R, suppose that

  • (Sample and alphabet sizes.) n1n\ge 1 and d1d\ge 1.
  • (Outcome envelope.) M1M\ge 1.
  • (Radius range.) 0σ20\le \sigma\le 2.

Let Rn,d,ϵ,M,σ\mathsf R_{n,d,\epsilon,M,\sigma} be the minimax mean-squared risk in Definition P-9 over the real-outcome class Pd,ϵ,M,σ\mathcal P_{d,\epsilon,M,\sigma} with fixed overlap ϵ\epsilon, conditional exchangeability, consistency, conditional mean normalization, conditional second central moment envelope MM, and approximate-homogeneity radius σM\sigma M. Define un,d=d2n2log2(en),hn,d,σ=σ2+dn2, u_{n,d}=\frac{d^2}{n^2\log^2(en)}, \qquad h_{n,d,\sigma}=\sigma^2+\frac{d}{n^2}, and write n,d,σ=1n+min{1,dn2}+σ2min{1,un,d},rn,d,σ=1n+min{1,un,d,hn,d,σ}. \ell_{n,d,\sigma} = \frac{1}{n} + \min\left\{1,\frac{d}{n^2}\right\} + \sigma^2\min\left\{1,u_{n,d}\right\}, \qquad r_{n,d,\sigma} = \frac{1}{n} + \min\left\{1,u_{n,d},h_{n,d,\sigma}\right\}. Then cϵM2n,d,σRn,d,ϵ,M,σCϵM2rn,d,σ. c_{\epsilon}M^2\ell_{n,d,\sigma} \le \mathsf R_{n,d,\epsilon,M,\sigma} \le C_{\epsilon}M^2r_{n,d,\sigma}.

Endpoints

  • At σ=0\sigma=0, exact homogeneity gives the collision scale n1+min(1,d/n2)n^{-1}+\min(1,d/n^2).
  • At σ=2\sigma=2, the radius-indexed class equals the unrestricted-radius class.
  • The unrestricted endpoint gives the large-alphabet polynomial scale n1+min{1,d2/[n2log2(en)]}n^{-1}+\min\{1,d^2/[n^2\log^2(en)]\}.

informal · Theorem T-12 The bracket reduces to the exact-homogeneity rate at radius zero and the unrestricted large-alphabet rate at radius two.

Theorem T-12 (Endpoint rate reductions)

There exist constants c,C(0,)c,C\in(0,\infty), with cCc\le C, such that the following hold.

  • (Exact homogeneity.) For every pair of integers n,d1n,d\ge 1, c{n1+min(1,dn2)}n1+min{1,d2n2log2(en),dn2}C{n1+min(1,dn2)}, c\left\{n^{-1}+\min\left(1,\frac d{n^2}\right)\right\} \le n^{-1}+\min\left\{1,\frac{d^2}{n^2\log^2(en)},\frac d{n^2}\right\} \le C\left\{n^{-1}+\min\left(1,\frac d{n^2}\right)\right\}, and c{n1+min(1,dn2)}n1+min(1,dn2)C{n1+min(1,dn2)}. c\left\{n^{-1}+\min\left(1,\frac d{n^2}\right)\right\} \le n^{-1}+\min\left(1,\frac d{n^2}\right) \le C\left\{n^{-1}+\min\left(1,\frac d{n^2}\right)\right\}.
  • (Unrestricted radius.) For every pair of integers n,d1n,d\ge 1, c{n1+min(1,d2n2log2(en))}n1+min(1,d2n2log2(en))C{n1+min(1,d2n2log2(en))}, c\left\{n^{-1}+\min\left(1,\frac{d^2}{n^2\log^2(en)}\right)\right\} \le n^{-1}+\min\left(1,\frac{d^2}{n^2\log^2(en)}\right) \le C\left\{n^{-1}+\min\left(1,\frac{d^2}{n^2\log^2(en)}\right)\right\}, and c{n1+min(1,d2n2log2(en))}n1+min(1,dn2)+4min(1,d2n2log2(en))C{n1+min(1,d2n2log2(en))}. c\left\{n^{-1}+\min\left(1,\frac{d^2}{n^2\log^2(en)}\right)\right\} \le n^{-1}+\min\left(1,\frac d{n^2}\right) +4\min\left(1,\frac{d^2}{n^2\log^2(en)}\right) \le C\left\{n^{-1}+\min\left(1,\frac{d^2}{n^2\log^2(en)}\right)\right\}.
  • (Class equality at radius two.) For every alphabet size dd, every ϵR\epsilon\in\mathbb R, and every MRM\in\mathbb R, the laws represented by Pd,ϵ,Munr\mathcal P_{d,\epsilon,M}^{\mathrm{unr}} in Definition P-2 are exactly the laws represented by Pd,ϵ,M,2\mathcal P_{d,\epsilon,M,2} in Definition P-1.
  • (Endpoint identity.) For every pair of integers n,d1n,d\ge 1, the radius-two upper frontier expression is exactly n1+min(1,d2n2log2(en)). n^{-1}+\min\left(1,\frac{d^2}{n^2\log^2(en)}\right).

Regimes

  • For every fixed positive radius, the upper and lower benchmarks agree in order.
  • They also agree in the saturated region.
  • They agree when the polynomial term is dominated by the exact-homogeneity baseline.
  • They agree when the squared radius is dominated by the exact-homogeneity baseline.
  • The remaining region is a shrinking-radius wedge with intermediate alphabet size.
Phase plane uₙ,d × σ² Polynomial rare-cell bias pays uₙ,d Collision uncrossed-cell bias pays σ²+d/n² Zero zero branch Homogeneity σ²=0 d/n² scale Unrestricted large alphabet d²/{n² log²(en)} Selector chooses branch no global commit Benchmark selector order
illustrative Box-and-arrow schematic showing the phase plane, polynomial branch, collision branch, zero branch, endpoint regimes, selector, and benchmark.

informal · Theorem T-13 Fixed positive radii, saturation, and parametric-dominance elbows have minimax risk within constants of the selector benchmark.

Proof Sketch

  • Upper bound: split the sample, classify cells, and analyze heavy and light contributions conditionally on the pilot split.
  • Light-cell control: Chebyshev coefficients approximate the rare-cell functional, and factorial statistics estimate the required moments.
  • Collision control: crossed cells supply direct contrasts, while σ\sigma bounds the effect of cells that fail to cross.
  • Lower bound: binary hard instances are embedded into real outcomes without changing the adjustment structure.
  • Radius channel: attenuation by σ/2\sigma/2 keeps alternatives inside the radius class and scales target separation.

Also in the Paper

informal · Theorem T-6 On the restricted dimension range, the polynomial and collision branches satisfy the same two constructive upper bounds.

informal · Theorem T-2 On the restricted dimension range, the known-radius selector has risk at most the stated frontier benchmark.

informal · Theorem T-4 On the restricted dimension range, the minimax risk is at least the parametric, exact-collision, and radius-channel lower benchmark.

informal · Theorem T-3 On the restricted dimension range, the endpoint expressions reduce to the exact-homogeneity and unrestricted-radius rates.

informal · Theorem T-7 On the restricted dimension range, fixed interior radii and the parametric elbows match the selector benchmark.

informal · Theorem T-8 On the restricted dimension range, the same two-sided minimax bracket holds with the range condition.

informal · Theorem Under the published binary collision guarantee, the selector remainder is at most the collision remainder, and it is asymptotically smaller when the polynomial remainder is negligible relative to the collision remainder.

Conclusion

  • We give a finite-sample minimax bracket for scalar ATE estimation with many discrete adjustment cells.
  • The model allows arbitrary cell masses and real outcomes under fixed overlap, bounded conditional means, and bounded conditional second moments.
  • The heterogeneity radius σ\sigma organizes the attainable risk.
  • The known-radius selector is constructive and all-alphabet.
  • The bracket is order matched at exact homogeneity, unrestricted radius, every fixed positive radius, saturation, and the parametric-dominance elbows.