CausalSmith · seminar slides
Sharp Rates for Discrete Confounding
For fixed interior overlap, we characterize the minimax ATE risk as 1/n+d2/(n2(logn)2) and construct a count-based hybrid estimator that attains it.
slides for Sharp Minimax Rates for Average Treatment Effects with Discrete Confounding under Fixed Overlap
Overview
- The setting is observational ATE estimation with a finite covariate alphabet of size d.
- Each category has its own propensity score and outcome means.
- Fixed overlap keeps every positive-mass category inside ϵ≤πk≤1−ϵ.
- The sharp mean-squared-error rate is 1/n+d2/(n2(logn)2).
- The estimator reaches the lower-bound scale by combining ratios on abundant categories with polynomial estimation on sparse categories.
Motivation
- Discrete covariates are a clean benchmark for high-dimensional adjustment.
- A category can be a covariate profile: age bin, location, diagnosis group, baseline score bin, or their full interaction.
- With many profiles, many categories have few observations even when n is large.
- Standard cellwise ATE estimators then pay for unstable treatment-control ratios.
- The question is: what is the best possible error when the categorywise nuisance functions are unrestricted?
Related Literature
- Rubin (1974) and Rosenbaum and Rubin (1983) organize ATE identification around potential outcomes, exchangeability, and overlap.
- Hahn (1998), Hirano, Imbens, and Ridder (2003), and Robins, Rotnitzky, and Zhao (1994) give the semiparametric ATE benchmark.
- Robins and Ritov (1997) highlight the inferential difficulty created by unrestricted high-dimensional adjustment.
- Jiao et al. (2015), Wu and Yang (2016), and Han et al. (2020) show how polynomial approximation sharpens large-alphabet functional estimation.
- Zeng et al. (2024) give the direct discrete-covariate causal predecessor and the lower-bound scale we attain.
Setup
- We observe Oi=(Xi,Ai,Yi) for i=1,…,n.
- X takes one of d categories, while A and Y are binary.
- pk is the category mass, πk is the treatment probability, and μak is the outcome mean in arm a.
- The overlap parameter ϵ fixes how close category-level assignment can be to deterministic.
- The target is the category-adjusted ATE, written as τ(P).
The observed data are Oi=(Xi,Ai,Yi), i=1,…,n, with O1,…,On∼iidP, where the covariate X takes values in {1,…,d}, the treatment A and outcome Y are binary. For each category k∈{1,…,d} write pk=P(X=k) for the category mass, πk=P(A=1∣X=k) for the treatment probability, and, for each treatment arm a∈{0,1}, μak=EP[Y∣A=a,X=k] for the conditional mean outcome.
Causal Interpretation
- Consistency links the observed outcome to the potential outcome under the realized treatment.
- Conditional exchangeability says treatment is as good as random after conditioning on the finite category X.
- Fixed overlap gives both treatment arms population support inside every positive-mass category.
- Together, these assumptions make τ(P) the finite-alphabet adjusted ATE.
For the overlap level ϵ, every category k with category mass pk>0 satisfies ϵ≤πk≤1−ϵ.
The observed outcome satisfies Y=Y(A) almost surely.
The potential outcomes satisfy (Y(0),Y(1))⊥A∣X.
Risk Target
- We evaluate worst-case mean-squared error over the overlap-restricted iid experiment class.
- The decision problem lets the distribution vary freely across all category-treatment-outcome cells subject to fixed overlap.
- This isolates the statistical price of unrestricted discrete confounding.
Rn,d,ϵ:=τinfP⊗n∈En,d,ϵsupEP[(τ−τ(P))2]. The infimum ranges over all measurable estimators τ based on O1,…,On.
Central Insight
- Abundant categories and sparse categories need different estimators.
- Heavy categories have enough observations for empirical treatment-control ratios.
- Light categories have unstable denominators, so direct ratios create the large-alphabet cost seen in plug-in methods.
- We approximate the reciprocal structure in the light-cell contribution by a Chebyshev polynomial of degree about logn.
- Falling-factorial moments turn that polynomial into an estimable count statistic.
- The logarithmic degree creates the d2/(n2(logn)2) large-alphabet term.
Estimator Pipeline
- Split the sample in half.
- Use the pilot split to classify categories as heavy or light at the L=log(en) scale.
- Estimate heavy categories by second-split empirical ratios.
- Estimate light categories by a Chebyshev reciprocal polynomial lifted through factorial moments.
- Add the two branches and clip the result to [−1,1].
Light Cells
informal · Lemma L-1 Under fixed overlap and d≤cϵnlogn, the light-cell polynomial branch has mean-squared error at most a constant times 1/n+d2/(n2(logn)2).
- The polynomial branch targets the sparse categories where empirical ratios are most fragile.
- Chebyshev approximation controls the bias of the reciprocal terms.
- Factorial moments estimate each polynomial monomial directly from counts.
- Computation is linear in d and polynomial in the logarithmic degree M(n).
Heavy Cells
informal · Lemma L-3 Under fixed interior overlap and d≤ρϵnlogn, the heavy-cell ratio branch has mean-squared error at most a constant times 1/n+d2/(n2(logn)2).
- The pilot split certifies that heavy categories have population mass at the logarithmic threshold.
- The estimation split then handles treatment-control ratios with controlled denominator events.
- The heavy branch contributes at the same scale as the light branch.
Lower Bound
informal · Lemma L-4 For every fixed 0<ϵ<1/2, the minimax risk is at least a constant times 1/n+d2/(n2(logn)2) when n≥Nϵ and d≤bϵnlogn.
- Zeng et al. (2024) provide the large-alphabet lower-bound technology for this causal experiment.
- The transfer step places that lower bound inside the fixed-overlap ATE risk.
- The result says the polynomial estimator reaches the correct benchmark.
Main Result
informal · Theorem T-2 For fixed 0<ϵ<1/2, n≥Nϵ, and d≤ρϵnlogn, the minimax MSE is bracketed above and below by constants times 1/n+d2/(n2(logn)2).
Fix an overlap level ϵ with 0<ϵ<1/2. Then there are constants aϵ,ρϵ,Cϵ>0 and an integer Nϵ<∞ such that, writing rn,d:=n1+n2(logn)2d2, the following statements hold.
- (Matched minimax envelope.) For every n,d, every single-observation law P, and every sample law μn in the experiment class En,d,ϵ of Definition P-1, if d>0,n≥Nϵ,d≤ρϵnlogn, then the minimax risk Rn,d,ϵ of Definition P-6 and the truncated balanced ratio-polynomial hybrid estimator τnhyb of Definition P-5 satisfy aϵrn,d≤Rn,d,ϵ≤P⊗n∈En,d,ϵsupEP[(τnhyb−τ(P))2]≤Cϵrn,d.
- (Parametric interior.) For every K>0 there is a constant CK>0 such that, for every n,d, every single-observation law P, and every sample law μn in En,d,ϵ, if d>0,n≥Nϵ,d≤Knlogn,d≤ρϵnlogn, then naϵ≤Rn,d,ϵ≤nCK.
- (Consistency threshold.) For every pair of integer sequences nj,dj with nj→∞, dj>0 eventually, and dj≤ρϵnjlognj eventually, the minimax risk converges to zero exactly along the sequences satisfying the vanishing normalized dimension condition: Rnj,dj,ϵ→0⟺njlognjdj→0.
Consequences
- The ordinary sampling term is 1/n.
- The unrestricted discrete-confounding term is d2/(n2(logn)2).
- In the parametric interior, d=O(nlogn) gives risk of order 1/n.
- Along sequences in the theorem’s range, consistency is equivalent to d/(nlogn)→0.
- In the running profile example, adjustment remains consistent across nearly nlogn unrestricted profiles up to the theorem’s vanishing normalized dimension condition.
Endpoint Comparison
informal · Theorem T-3 The same hybrid estimator is computable and attains the fixed-overlap rate, while a centered estimator gives the near-randomization envelope and the exact randomization endpoint has minimax risk between 1/(100n) and 1/n.
With τnhyb as in Definition P-5 and Rn,d,ϵ as in Definition P-6, the following assertions hold.
- Computability. There is a universal integer K>0 such that, for every n,d∈N, a real-arithmetic program based on the hybrid count input has operation count at most KdM(n)4 and, for every sample O1,…,On, evaluates exactly to τnhyb.
- Fixed-overlap hybrid calibration. For every overlap level ϵ with 0<ϵ<1/2, there are constants Cϵ,ρϵ>0 and an integer Nϵ such that, for all n,d∈N with n>0, d>0, n≥Nϵ, and d≤ρϵnlogn, the hybrid estimator satisfies P⊗n∈En,d,ϵsupEP[(τnhyb−τ(P))2]≤Cϵ{n1+n2(logn)2d2}. For the deterministic estimator τCϵ,ϵ={τnhyb,τctr,if Cϵ{n1+n2(logn)2d2}≤n1+4(1/2−ϵ)2,otherwise, one also has Rn,d,ϵ≤P⊗n∈En,d,ϵsupEP[(τCϵ,ϵ−τ(P))2]≤max{Cϵ,4}[n1+min{n2(logn)2d2,(1/2−ϵ)2}].
- Centered randomization bound. For every overlap level ϵ with 0<ϵ≤1/2 and every n,d∈N with n>0 and d>0, the centered estimator τctr=n1i=1∑n2(2Ai−1)(Yi−21) satisfies P⊗n∈En,d,ϵsupEP[(τctr−τ(P))2]≤n1+4(1/2−ϵ)2.
- Endpoint randomization bracket. For every n,d∈N with n>0 and d>0, 100n1≤Rn,d,1/2≤n1.
Proof Sketch
- First, the pilot split sorts categories with polynomially small classification error.
- Second, light cells use approximation: the Chebyshev reciprocal polynomial keeps sparse-cell bias at the target scale.
- Third, factorial moments estimate the polynomial terms while controlling variance and cross-category covariance.
- Fourth, heavy cells use ratio residual bounds and missing-arm controls.
- Finally, the lower bound and the two upper branches meet at the same fixed-interior rate.
Takeaways
- We characterize the fixed-interior minimax MSE rate for unrestricted finite-alphabet ATE estimation.
- We construct a computable hybrid estimator that attains the rate with universal numerical tuning.
- The rate gives a parametric window d=O(nlogn) and a consistency frontier d=o(nlogn) within the calibrated theorem range.
- The endpoint results connect the fixed-overlap analysis to exact randomization, where the risk is parametric.