CausalSmith · seminar slides

Interior Dose-Response Lower Bounds

Under Hölder smoothness, local positivity, bounded outcomes, and strict baseline slack, the minimax mean-squared error for an interior continuous-treatment dose-response partial mean is at least the one-dimensional treatment-regression rate.

Overview

  • We study the target-dose partial mean for a continuous treatment.
  • The target asks: what is the average outcome if the dose is fixed at an interior value t0t_0?
  • The lower bound is over the same Hölder dose class used in the risk statement.
  • The obstruction comes from learning the outcome regression locally in the treatment coordinate.
  • The lower-bound exponent is independent of the treatment-density smoothness β\beta, for each fixed β>0\beta>0.

informal · Theorem T-1 Under the stated smoothness, positivity, boundedness, and strict-slack conditions, the minimax MSE is at least a constant times n2α/(2α+1)n^{-2\alpha/(2\alpha+1)}.

Motivation

  • Continuous-treatment dose-response curves are common in policy, medicine, and economics.
  • A policymaker may ask for the mean outcome at a particular interior dose, rather than for a binary treatment contrast.
  • Regression adjustment estimates the conditional mean at that dose and averages over the covariate distribution.
  • The hard part is local: observations near t0t_0 carry the information about the regression value at t0t_0.
  • The minimax question asks how small the worst-case squared error can be.

Running Example

  • Think of AA as dose intensity, YY as an outcome, and XX as pre-treatment covariates.
  • We evaluate an interior dose t0t_0, such as a moderate policy intensity or medication level.
  • Local positivity says each covariate group has enough probability of receiving doses near t0t_0.
  • Hölder smoothness says the regression and density vary regularly near that dose.
  • The target averages the dose-t0t_0 regression over the population distribution of XX.

Target

- One observation is O=(Y,A,X)O=(Y,A,X). - AA is a continuous treatment in [0,1][0,1]. - μP(a,x)\mu_P(a,x) is the conditional mean of YY at dose aa and covariates xx. - pX,Pp_{X,P} is the covariate density. - The estimand is θP(t0)=[0,1]dμP(t0,x)pX,P(x)dx. \theta_P(t_0) = \int_{[0,1]^d} \mu_P(t_0,x)\,p_{X,P}(x)\,dx .

Observed data O=(Y,A,X) Regression μ_P(a,x) conditional mean Target dose interior t₀ Evaluation μ_P(t₀,X) Covariates population distribution Target mean θ_P(t₀) averages over X Assumptions consistency, no confounding local positivity Interpretation causal dose-response
illustrative Box-and-arrow schematic from observed data to the regression, target dose, covariate averaging, target mean, and causal dose-response interpretation under consistency, no confounding, and local positivity.

Identification

  • The statistical target is an observed-data partial mean.
  • With consistency, the observed outcome equals the potential outcome at the realized dose.
  • With no unmeasured confounding, treatment assignment is conditionally independent of potential outcomes given XX.
  • With local positivity, the data contain information near the target dose for every covariate value.
  • These conditions give the causal reading of θP(t0)\theta_P(t_0) as the dose-response mean at t0t_0.

Model

  • Outcomes are uniformly bounded by MM.
  • The dose t0t_0 has a full local window inside (0,1)(0,1).
  • The treatment density is bounded below by c0c_0 on that window.
  • The treatment regression has Hölder smoothness α\alpha in the dose coordinate.
  • The treatment density has Hölder smoothness β\beta in the dose coordinate.
  • The regression, treatment density, and covariate density have Hölder smoothness ss in the covariate coordinate.

Baseline Slack

  • The lower-bound construction starts from baseline densities p0p_0 and q0q_0.
  • Strict slack means these baselines sit inside the smoothness, boundedness, and positivity restrictions with positive margin.
  • That margin lets us add a local perturbation while staying in the same model class.
  • In the running example, the baseline population and dose assignment mechanism are regular enough that a small local regression change remains admissible.
Assumption A-12 (Baseline submodel slack)

There exist a covariate density p0p_0 on [0,1]d[0,1]^d, a conditional treatment density q0q_0 on [0,1][0,1], a constant η0>0\eta_0>0, and a constant B0B_0 with 0<B0<M0<B_0<M, such that p0(x)0p_0(x)\ge 0 for all x[0,1]dx\in[0,1]^d, q0(a)0q_0(a)\ge 0 for all aRa\in\mathbb R, [0,1]dp0(x)dx=1and01q0(a)da=1, \int_{[0,1]^d} p_0(x)\,dx=1 \qquad\text{and}\qquad \int_0^1 q_0(a)\,da=1, p0Hs([0,1]d)Mη0, \|p_0\|_{\mathcal H^s([0,1]^d)}\le M-\eta_0, q0q_0 has Hölder norm at most Mη0M-\eta_0 of order β\beta on [t0ε0,t0+ε0][t_0-\varepsilon_0,t_0+\varepsilon_0], q0(a)c0+η0for all a[t0ε0,t0+ε0], q_0(a)\ge c_0+\eta_0 \qquad\text{for all }a\in[t_0-\varepsilon_0,t_0+\varepsilon_0], and p0(x)Mη0for all x[0,1]d. p_0(x)\le M-\eta_0 \qquad\text{for all }x\in[0,1]^d.

Related Literature

  • Rubin (1974), Rosenbaum and Rubin (1983), and Robins (1986) provide the causal potential-outcome and adjustment foundations.
  • Imbens (2000), Hirano and Imbens (2004), and Imai and van Dyk (2004) develop continuous-treatment propensity and dose-response ideas.
  • Kennedy et al. (2017), Lee (2018), and Colangelo and Lee (2020) study nonparametric and debiased continuous-treatment estimation.
  • Bonvini and Kennedy (2022) give the higher-order influence-function benchmark used for comparison.
  • Stone (1982) and Tsybakov (2009) provide the nonparametric minimax lower-bound toolkit.

Main Result I

informal · Theorem T-1 For every fixed positive β\beta, the same-class minimax MSE is at least a constant times the one-dimensional treatment-regression rate under the stated slack-baseline conditions.

Theorem T-1 (Sharp Pointwise Lower Bound)

For every covariate dimension dd and every real constants α,β,s,M,c0,ε0,t0\alpha,\beta,s,M,c_0,\varepsilon_0,t_0, suppose that

  • (Positive smoothness.) α>0\alpha>0, β>0\beta>0, and s>0s>0.
  • (Regime constants.) M>0M>0, c0>0c_0>0, t0(0,1)t_0\in(0,1), ε0(0,1/2)\varepsilon_0\in(0,1/2), and the local window determined by (t0,ε0)(t_0,\varepsilon_0) is contained in (0,1)(0,1).
  • (Strict-slack baseline.) Assumption A-12 holds for (d,β,s,M,c0,ε0,t0)(d,\beta,s,M,c_0,\varepsilon_0,t_0).

Then there exists a constant c>0c>0 such that, for all sufficiently large sample sizes nn, cn2α/(2α+1)Rn ⁣(Pα,β,s(M,c0,ε0,t0),t0), c\, n^{-2\alpha/(2\alpha+1)} \le R_n\!\left( \mathcal P_{\alpha,\beta,s}(M,c_0,\varepsilon_0,t_0), t_0 \right), where Pα,β,s(M,c0,ε0,t0)\mathcal P_{\alpha,\beta,s}(M,c_0,\varepsilon_0,t_0) is the model class of Definition P-1 and RnR_n is the minimax risk of Definition P-3.

Key Idea

  • Hold the covariate density and treatment density fixed.
  • Perturb only the outcome regression in a narrow neighborhood of t0t_0.
  • The perturbation changes θP(t0)\theta_P(t_0) because the target evaluates the regression exactly at t0t_0.
  • The induced data laws remain statistically close because the perturbed region is narrow.
  • The best test between the two laws cannot reliably detect the target shift.
Fixed density covariates frozen treatment frozen Local bump width h near t₀ Hölder height Law minus observed data shared densities Law plus observed data bumped regression Close laws statistically close Target shift θ_P(t₀) separated β obstruction every fixed β
illustrative Box-and-arrow schematic showing fixed covariate and treatment densities, a local bump near t0t_0, two observed-data laws, statistical closeness, target separation, and the resulting all-β\beta lower-bound obstruction.

Intuition

  • A direct plug-in view says the problem is to learn μP(t0,x)\mu_P(t_0,x) from observations with doses near t0t_0.
  • A very narrow neighborhood gives low bias but few effective observations.
  • A wider neighborhood gives more observations but cannot resolve α\alpha-Hölder local variation at t0t_0.
  • The lower-bound construction chooses the local bump width that balances these two forces.
  • Because the treatment density is fixed inside the construction, β\beta-smoothness remains part of the class while the hard pair is governed by α\alpha.

Smooth Covariates

- The published benchmark rate is ρn=n2α/(2α+1)n2/(1+d/(4s)+1/α). \rho_n = n^{-2\alpha/(2\alpha+1)} \vee n^{-2/(1+d/(4s)+1/\alpha)} . - When d4sd\le 4s, the benchmark collapses to the treatment-regression term. - In that regime, the same-class lower floor has the same exponent as ρn\rho_n.

informal · Theorem T-2 In the smooth-covariate regime d4sd\le 4s, the published benchmark equals the treatment-regression rate, and the minimax risk is at least a constant times that benchmark.

Main Result II

informal · Theorem T-4 When d4sd\le 4s, the same-class lower bound is at least a constant times n2α/(2α+1)n^{-2\alpha/(2\alpha+1)}, and this rate equals ρn\rho_n.

Theorem T-4 (Smooth Covariate Minimax Floor)

Let dNd\in\mathbb N and let α,β,s,M,c0,ε0,t0R\alpha,\beta,s,M,c_0,\varepsilon_0,t_0\in\mathbb R. Suppose that:

  • (Smoothness.) α>0\alpha>0, β>0\beta>0, and s>0s>0.
  • (Regime constants.) M>0M>0, c0>0c_0>0, t0(0,1)t_0\in(0,1), ε0(0,1/2)\varepsilon_0\in(0,1/2), and the window [t0ε0,t0+ε0][t_0-\varepsilon_0,t_0+\varepsilon_0] is contained in (0,1)(0,1).
  • (Smooth-covariate regime.) d4sd\le 4s.
  • (Strict-slack baseline.) Assumption A-12 holds for (d,β,s,M,c0,ε0,t0)(d,\beta,s,M,c_0,\varepsilon_0,t_0).

Then there exists a constant c>0c>0 such that, for all sufficiently large nn, cn2α/(2α+1)Rn ⁣(Pα,β,s(M,c0,ε0,t0),t0), c\,n^{-2\alpha/(2\alpha+1)} \le R_n\!\left( \mathcal P_{\alpha,\beta,s}(M,c_0,\varepsilon_0,t_0),t_0 \right), where RnR_n is the minimax risk of Definition P-3 over the same-class model of Definition P-1. Moreover, n2α/(2α+1)=ρn, n^{-2\alpha/(2\alpha+1)} = \rho_n, with ρn\rho_n as in Definition P-4.

Low Covariate Smoothness

  • When 4s<d4s<d, the published benchmark is governed by the covariate-smoothness term.
  • The lower-bound construction still gives the treatment-regression floor.
  • The two exponents are strictly ordered in this regime.
  • The result identifies the algebraic gap between the same-class lower floor and the published benchmark sequence.

informal · Theorem T-5 When 4s<d4s<d, the minimax risk is still at least the treatment-regression lower floor, while the published benchmark has a strictly smaller exponent.

Theorem T-5 (Deficient Frontier Bracket)

Fix a covariate dimension dNd\in\mathbb N and constants α,β,s,M,c0,ε0,t0R\alpha,\beta,s,M,c_0,\varepsilon_0,t_0\in\mathbb R. Assume:

  • (Regime constants.) α>0\alpha>0, β>0\beta>0, s>0s>0, M>0M>0, c0>0c_0>0, t0(0,1)t_0\in(0,1), ε0(0,1/2)\varepsilon_0\in(0,1/2), and the window [t0ε0,t0+ε0][t_0-\varepsilon_0,t_0+\varepsilon_0] lies in (0,1)(0,1).
  • (Deficient covariate smoothness.) 4s<d4s<d.
  • (Strict-slack baseline.) Assumption A-12 holds for (d,β,s,M,c0,ε0,t0)(d,\beta,s,M,c_0,\varepsilon_0,t_0).

Then there exists a constant c>0c>0 such that, for all sufficiently large sample sizes nn, cn2α/(2α+1)Rn ⁣(Pα,β,s(M,c0,ε0,t0),t0), c\,n^{-2\alpha/(2\alpha+1)} \le R_n\!\left( \mathcal P_{\alpha,\beta,s}(M,c_0,\varepsilon_0,t_0),t_0 \right), where the minimax risk is as in Definition P-3 and the class is Definition P-1. Moreover, for the published HOIF benchmark of Definition P-4, ρn=n2/(1+d/(4s)+1/α) \rho_n = n^{-\,2/(1+d/(4s)+1/\alpha)} and its exponent is strictly smaller than the certified lower-bound exponent: 21+d/(4s)+1/α<2α2α+1. \frac{2}{1+d/(4s)+1/\alpha} < \frac{2\alpha}{2\alpha+1}.

  • This separates the certified same-class lower floor from the published benchmark exponent in the low-covariate-smoothness regime.
  • The comparison is exact at the level of the displayed benchmark algebra.

Also in the Paper

informal · Gate L-17 Under the stated Bonvini-Kennedy localized-regularity conditions, the cited higher-order influence-function estimator attains conditional mean-squared error at most a constant times ρn\rho_n.

informal · Definition P-5 The beta comparison handle records the all-β\beta lower floor, the two benchmark regimes, and the possible same-class upper-frontier alternatives.

Proof Sketch

  • Build two laws with the same covariate density and the same treatment density.
  • Add a local Hölder-compatible regression bump near t0t_0 under one law.
  • Use bounded two-point outcome channels so the target shift is carried by the conditional mean.
  • Calibrate the bump so the nn-sample laws have bounded Kullback-Leibler divergence.
  • Apply Le Cam’s two-point argument to convert indistinguishability into a squared-error lower bound.

Takeaways

  • The interior dose-response partial mean has a same-class minimax lower floor at n2α/(2α+1)n^{-2\alpha/(2\alpha+1)}.
  • The floor holds for every fixed β>0\beta>0 under the corresponding strict-slack baseline condition.
  • When d4sd\le 4s, this exponent matches the published higher-order influence-function benchmark exponent.
  • When 4s<d4s<d, the benchmark exponent and the certified same-class lower-bound exponent are strictly ordered.
  • The comparison clarifies exactly which rate facts are established for the Hölder dose class and which upper-side questions remain for the low-ss regime.