CausalSmith · seminar slides
Interior Dose-Response Lower Bounds
Under Hölder smoothness, local positivity, bounded outcomes, and strict baseline slack, the minimax mean-squared error for an interior continuous-treatment dose-response partial mean is at least the one-dimensional treatment-regression rate.
slides for A Minimax Lower Bound for Interior Dose-Response Estimation
Overview
- We study the target-dose partial mean for a continuous treatment.
- The target asks: what is the average outcome if the dose is fixed at an interior value t0?
- The lower bound is over the same Hölder dose class used in the risk statement.
- The obstruction comes from learning the outcome regression locally in the treatment coordinate.
- The lower-bound exponent is independent of the treatment-density smoothness β, for each fixed β>0.
informal · Theorem T-1 Under the stated smoothness, positivity, boundedness, and strict-slack conditions, the minimax MSE is at least a constant times n−2α/(2α+1).
Motivation
- Continuous-treatment dose-response curves are common in policy, medicine, and economics.
- A policymaker may ask for the mean outcome at a particular interior dose, rather than for a binary treatment contrast.
- Regression adjustment estimates the conditional mean at that dose and averages over the covariate distribution.
- The hard part is local: observations near t0 carry the information about the regression value at t0.
- The minimax question asks how small the worst-case squared error can be.
Running Example
- Think of A as dose intensity, Y as an outcome, and X as pre-treatment covariates.
- We evaluate an interior dose t0, such as a moderate policy intensity or medication level.
- Local positivity says each covariate group has enough probability of receiving doses near t0.
- Hölder smoothness says the regression and density vary regularly near that dose.
- The target averages the dose-t0 regression over the population distribution of X.
Target
- One observation is O=(Y,A,X). - A is a continuous treatment in [0,1]. - μP(a,x) is the conditional mean of Y at dose a and covariates x. - pX,P is the covariate density. - The estimand is θP(t0)=∫[0,1]dμP(t0,x)pX,P(x)dx.
Identification
- The statistical target is an observed-data partial mean.
- With consistency, the observed outcome equals the potential outcome at the realized dose.
- With no unmeasured confounding, treatment assignment is conditionally independent of potential outcomes given X.
- With local positivity, the data contain information near the target dose for every covariate value.
- These conditions give the causal reading of θP(t0) as the dose-response mean at t0.
Model
- Outcomes are uniformly bounded by M.
- The dose t0 has a full local window inside (0,1).
- The treatment density is bounded below by c0 on that window.
- The treatment regression has Hölder smoothness α in the dose coordinate.
- The treatment density has Hölder smoothness β in the dose coordinate.
- The regression, treatment density, and covariate density have Hölder smoothness s in the covariate coordinate.
Baseline Slack
- The lower-bound construction starts from baseline densities p0 and q0.
- Strict slack means these baselines sit inside the smoothness, boundedness, and positivity restrictions with positive margin.
- That margin lets us add a local perturbation while staying in the same model class.
- In the running example, the baseline population and dose assignment mechanism are regular enough that a small local regression change remains admissible.
There exist a covariate density p0 on [0,1]d, a conditional treatment density q0 on [0,1], a constant η0>0, and a constant B0 with 0<B0<M, such that p0(x)≥0 for all x∈[0,1]d, q0(a)≥0 for all a∈R, ∫[0,1]dp0(x)dx=1and∫01q0(a)da=1, ∥p0∥Hs([0,1]d)≤M−η0, q0 has Hölder norm at most M−η0 of order β on [t0−ε0,t0+ε0], q0(a)≥c0+η0for all a∈[t0−ε0,t0+ε0], and p0(x)≤M−η0for all x∈[0,1]d.
Related Literature
- Rubin (1974), Rosenbaum and Rubin (1983), and Robins (1986) provide the causal potential-outcome and adjustment foundations.
- Imbens (2000), Hirano and Imbens (2004), and Imai and van Dyk (2004) develop continuous-treatment propensity and dose-response ideas.
- Kennedy et al. (2017), Lee (2018), and Colangelo and Lee (2020) study nonparametric and debiased continuous-treatment estimation.
- Bonvini and Kennedy (2022) give the higher-order influence-function benchmark used for comparison.
- Stone (1982) and Tsybakov (2009) provide the nonparametric minimax lower-bound toolkit.
Main Result I
informal · Theorem T-1 For every fixed positive β, the same-class minimax MSE is at least a constant times the one-dimensional treatment-regression rate under the stated slack-baseline conditions.
For every covariate dimension d and every real constants α,β,s,M,c0,ε0,t0, suppose that
- (Positive smoothness.) α>0, β>0, and s>0.
- (Regime constants.) M>0, c0>0, t0∈(0,1), ε0∈(0,1/2), and the local window determined by (t0,ε0) is contained in (0,1).
- (Strict-slack baseline.) Assumption A-12 holds for (d,β,s,M,c0,ε0,t0).
Then there exists a constant c>0 such that, for all sufficiently large sample sizes n, cn−2α/(2α+1)≤Rn(Pα,β,s(M,c0,ε0,t0),t0), where Pα,β,s(M,c0,ε0,t0) is the model class of Definition P-1 and Rn is the minimax risk of Definition P-3.
Key Idea
- Hold the covariate density and treatment density fixed.
- Perturb only the outcome regression in a narrow neighborhood of t0.
- The perturbation changes θP(t0) because the target evaluates the regression exactly at t0.
- The induced data laws remain statistically close because the perturbed region is narrow.
- The best test between the two laws cannot reliably detect the target shift.
Intuition
- A direct plug-in view says the problem is to learn μP(t0,x) from observations with doses near t0.
- A very narrow neighborhood gives low bias but few effective observations.
- A wider neighborhood gives more observations but cannot resolve α-Hölder local variation at t0.
- The lower-bound construction chooses the local bump width that balances these two forces.
- Because the treatment density is fixed inside the construction, β-smoothness remains part of the class while the hard pair is governed by α.
Smooth Covariates
- The published benchmark rate is ρn=n−2α/(2α+1)∨n−2/(1+d/(4s)+1/α). - When d≤4s, the benchmark collapses to the treatment-regression term. - In that regime, the same-class lower floor has the same exponent as ρn.
informal · Theorem T-2 In the smooth-covariate regime d≤4s, the published benchmark equals the treatment-regression rate, and the minimax risk is at least a constant times that benchmark.
Main Result II
informal · Theorem T-4 When d≤4s, the same-class lower bound is at least a constant times n−2α/(2α+1), and this rate equals ρn.
Let d∈N and let α,β,s,M,c0,ε0,t0∈R. Suppose that:
- (Smoothness.) α>0, β>0, and s>0.
- (Regime constants.) M>0, c0>0, t0∈(0,1), ε0∈(0,1/2), and the window [t0−ε0,t0+ε0] is contained in (0,1).
- (Smooth-covariate regime.) d≤4s.
- (Strict-slack baseline.) Assumption A-12 holds for (d,β,s,M,c0,ε0,t0).
Then there exists a constant c>0 such that, for all sufficiently large n, cn−2α/(2α+1)≤Rn(Pα,β,s(M,c0,ε0,t0),t0), where Rn is the minimax risk of Definition P-3 over the same-class model of Definition P-1. Moreover, n−2α/(2α+1)=ρn, with ρn as in Definition P-4.
Low Covariate Smoothness
- When 4s<d, the published benchmark is governed by the covariate-smoothness term.
- The lower-bound construction still gives the treatment-regression floor.
- The two exponents are strictly ordered in this regime.
- The result identifies the algebraic gap between the same-class lower floor and the published benchmark sequence.
informal · Theorem T-5 When 4s<d, the minimax risk is still at least the treatment-regression lower floor, while the published benchmark has a strictly smaller exponent.
Fix a covariate dimension d∈N and constants α,β,s,M,c0,ε0,t0∈R. Assume:
- (Regime constants.) α>0, β>0, s>0, M>0, c0>0, t0∈(0,1), ε0∈(0,1/2), and the window [t0−ε0,t0+ε0] lies in (0,1).
- (Deficient covariate smoothness.) 4s<d.
- (Strict-slack baseline.) Assumption A-12 holds for (d,β,s,M,c0,ε0,t0).
Then there exists a constant c>0 such that, for all sufficiently large sample sizes n, cn−2α/(2α+1)≤Rn(Pα,β,s(M,c0,ε0,t0),t0), where the minimax risk is as in Definition P-3 and the class is Definition P-1. Moreover, for the published HOIF benchmark of Definition P-4, ρn=n−2/(1+d/(4s)+1/α) and its exponent is strictly smaller than the certified lower-bound exponent: 1+d/(4s)+1/α2<2α+12α.
- This separates the certified same-class lower floor from the published benchmark exponent in the low-covariate-smoothness regime.
- The comparison is exact at the level of the displayed benchmark algebra.
Also in the Paper
informal · Gate L-17 Under the stated Bonvini-Kennedy localized-regularity conditions, the cited higher-order influence-function estimator attains conditional mean-squared error at most a constant times ρn.
informal · Definition P-5 The beta comparison handle records the all-β lower floor, the two benchmark regimes, and the possible same-class upper-frontier alternatives.
Proof Sketch
- Build two laws with the same covariate density and the same treatment density.
- Add a local Hölder-compatible regression bump near t0 under one law.
- Use bounded two-point outcome channels so the target shift is carried by the conditional mean.
- Calibrate the bump so the n-sample laws have bounded Kullback-Leibler divergence.
- Apply Le Cam’s two-point argument to convert indistinguishability into a squared-error lower bound.
Takeaways
- The interior dose-response partial mean has a same-class minimax lower floor at n−2α/(2α+1).
- The floor holds for every fixed β>0 under the corresponding strict-slack baseline condition.
- When d≤4s, this exponent matches the published higher-order influence-function benchmark exponent.
- When 4s<d, the benchmark exponent and the certified same-class lower-bound exponent are strictly ordered.
- The comparison clarifies exactly which rate facts are established for the Hölder dose class and which upper-side questions remain for the low-s regime.