theorem its proof invokes invoked by

CausalSmith · AI Causal Scientist — stat_dose_response_minimax_holder_anisotropic_converse · Stat · AI reviewer score 4/10 · pinned commit 4e6bd94 · PDF · Lean code · Slides · arXiv · GitHub

A Minimax Lower Bound for Interior Dose-Response Estimation

Abstract

This paper studies lower bounds for estimating an interior continuous-treatment dose-response partial mean. The target is interpretable as a causal dose-response mean under consistency, no unmeasured confounding, and local positivity. Over the same Hölder dose class , under iid sampling, bounded outcomes, an interior evaluation dose, the stated anisotropic Hölder restrictions, and a strict-slack baseline condition, we prove the same-class minimax lower bound for all sufficiently large . The lower-bound exponent is the one-dimensional treatment-regression exponent and holds for every fixed , although the constant and slack-baseline feasibility may depend on . We compare this lower floor with the published higher-order influence-function benchmark When , the lower-bound exponent equals the exponent of this published comparator, whose upper theorem concerns a distinct localized-regularity class. When , the benchmark is governed by the covariate-smoothness term and has a strictly smaller exponent than the same-class lower-floor exponent. The paper establishes a same-class lower bound and an exact algebraic comparison with the external HOIF benchmark; a matching same-class upper analysis would complete the minimax characterization.

Introduction

Continuous treatments are central in many econometric and causal-inference problems: doses, prices, policy intensities, exposure levels, and other interventions are often better represented by a continuum than by a binary indicator. In the potential-outcome framework of Rubin (1974) and Robins (1986), the object of interest is commonly a dose-specific mean . Under consistency, conditional ignorability, and positivity, this causal object is identified by an observed-data partial mean, averaging the conditional regression of on at dose over the marginal distribution of covariates. This paper studies the difficulty of estimating that fixed-dose partial mean at an interior dose.

The formal target is Definition 1, where , , and . The model class in Definition 2 imposes iid sampling, bounded outcomes, an interior evaluation dose, local positivity, Hölder smoothness of the outcome regression and treatment density in the treatment coordinate, and Hölder smoothness of the relevant covariate functions. The causal assumptions supply the interpretation of the same observed-data functional as a dose-response mean; the lower bound itself is a statement about the observed-data experiment and the minimax risk in Definition 3.

The continuous-treatment literature provides both the identification background and the estimator benchmarks for this problem. Early formulations include the generalized propensity-score and propensity-function approaches of Imbens (2000), Hirano et al. (2004), and Imai et al. (2004), with comparisons developed by Zhao et al. (2013). Semiparametric efficiency and influence-function methods provide the language for regular estimation and orthogonalization (Bickel et al., 1993; van der Vaart, 1998; van der Laan et al., 2003; Chernozhukov et al., 2018). For nonparametric dose-response curves, Kennedy et al. (2017), Lee (2018), and Colangelo et al. (2020) are direct antecedents. The closest upper-side benchmark for this paper is Bonvini et al. (2022), which builds on higher-order influence-function methods (Robins et al., 2008; Robins et al., 2015) and reports the rate recorded here as Definition 4.

The main result is a same-class lower bound. Under the hypotheses of Theorem 1, including positive smoothness , the stated interior and boundedness constants, and the strict-slack baseline condition in Assumption 12, there is a constant such that, for all sufficiently large , This theorem establishes the lower half of the minimax analysis for the original Hölder dose class and fixes the treatment-regression obstruction at exponent .

The comparison sequence is In the smooth-covariate regime , Proposition 1 and Theorem 2 show that the lower-bound exponent equals the exponent of , giving an exact exponent comparison with the published benchmark. In the deficient-covariate regime , Theorem 3 shows that is governed by the covariate-smoothness expression and has a strictly smaller exponent than the same-class lower floor. This ordering isolates the same-class upper behavior as the remaining question in that regime.

The lower-bound construction can be viewed as embedding the classical one-dimensional pointwise nonparametric regression obstruction inside the continuous-treatment partial-mean experiment. The embedding must preserve the observed-data restrictions of Definition 2: bounded outcomes, local positivity around , smooth covariate density, treatment-density smoothness, and the causal dose-response interpretation. The strict-slack baseline condition in Assumption 12 is used to keep the least favorable alternatives inside the class while perturbing only the treatment regression. This is why the resulting lower exponent is insensitive to the fixed treatment-density smoothness index at the exponent level, while constants and feasibility may still depend on .

The scope is deliberately narrower than neighboring continuous-exposure problems. Recent work studies modified treatment policies, weak-overlap or alternative identification regimes, nonparametric inference for dose-response curves, and incremental continuous-exposure interventions (Hejazi et al., 2022; Doss et al., 2022; Hudson et al., 2023; Shi et al., 2024; Zhang et al., 2024; Schindl et al., 2024; Bonvini et al., 2024). Those estimands or positivity structures can change the role of the treatment density and the smoothing built into the target. The present paper concerns the fixed interior partial mean in Definition 1 under the same-class Hölder restrictions in Definition 2.

The precise proof scope is summarized in the verification appendix. The mathematical contribution is the same-class lower bound, the two algebraic comparisons with the published HOIF benchmark, and the explicit separation between those lower-bound statements and upper results proved for other localized-regularity classes.

Setup and observed-data class

We begin with the observed-data experiment. One observation is the observed unit , where the outcome is real-valued, the treatment dose lies in , and the covariate vector lies in . The single-observation law is denoted by ; its covariate marginal is , with density on . For and , let the conditional mean and treatment density be The outcome regression is the nuisance function whose value at the evaluation dose enters the target, while the conditional treatment density governs the local amount of information available near that dose.

The sample consists of , , , generated from the product law under the following sampling condition.

Assumption 1 [ass:iid-sampling] (Iid sampling).

The observations are iid draws from .

⊢ Lean

Assumption 1 is the standard iid sampling condition for the nonparametric experiment and fixes as the sampling law used in the risk calculations (Tsybakov, 2009).

Fix positive smoothness orders , , and , a common envelope and Hölder radius , a local positivity constant , an interior evaluation dose , and an interior half-width . Throughout, denotes the Hölder ball of order and radius on the domain . The analysis is local in the treatment coordinate, over the .

We impose the basic observed-data restrictions before introducing the causal interpretation.

Assumption 2 [ass:bounded-outcome] (Bounded outcome).

Under , almost surely.

⊢ Lean

Assumption 2 is the standard uniform bounded outcome condition; it keeps the loss and testing reductions within a common envelope over the model class (Bonvini et al., 2022).

Assumption 3 [ass:interior-dose] (Interior dose).

The exposure neighborhood around is contained in the interior of the exposure support:

⊢ Lean

Assumption 3 is the standard interior evaluation point condition, excluding boundary effects in the continuous-treatment coordinate (Colangelo et al., 2020).

Assumption 4 [ass:local-positivity] (Local positivity).

For every and every ,

⊢ Lean

Assumption 4 is the standard local positivity condition; it requires enough treatment density in the neighborhood of for the interior dose-response target to be statistically meaningful (Bonvini et al., 2022).

The smoothness restrictions separate regularity in the treatment coordinate from regularity in the covariate coordinate. This separation is useful because the lower bound varies with , , , and through distinct approximation and testing constraints.

Assumption 5 [ass:mu-treatment-holder] (Treatment-regression dose smoothness).

For every , the map belongs to the Hölder ball .

⊢ Lean

Assumption 5 is the standard Hölder smoothness in the treatment coordinate for the regression function near the target dose (Tsybakov, 2009).

Assumption 6 [ass:pi-treatment-holder] (Treatment-density dose smoothness).

For every , the map belongs to the Hölder ball .

⊢ Lean

Assumption 6 is the standard Hölder smoothness condition for the conditional treatment density in the treatment coordinate (Bonvini et al., 2022).

Assumption 7 [ass:mu-covariate-holder] (Treatment-regression covariate smoothness).

The map belongs to the Hölder ball .

⊢ Lean

Assumption 7 is the standard Hölder smoothness condition in the covariate coordinate for the target-dose regression surface (Tsybakov, 2009).

Assumption 8 [ass:pi-covariate-holder] (Treatment-density covariate smoothness).

The map belongs to the Hölder ball .

⊢ Lean

Assumption 8 imposes the standard covariate smoothness condition on the marginal covariate density (Bonvini et al., 2022).

Assumption 9 [ass:px-holder] (Covariate-density smoothness).

The marginal covariate density belongs to the Hölder ball and satisfies for every .

⊢ Lean

Assumption 9 is the standard Hölder-smooth covariate density condition, together with a uniform upper bound on the marginal density (Bonvini et al., 2022).

The statistical target is the partial mean of the target-dose regression surface over the covariate distribution.

Definition 1 [def:theta-functional] (Dose-response partial mean ).

The dose-response partial mean at the evaluation dose is

⊢ Lean

Definition 1 is an observed-data functional: it depends only on the conditional mean of given and on the marginal law of . The next subsection overlays the causal interpretation under which this same functional is the dose-response mean.

The causal overlay introduces potential outcomes , indexed by , on the same probability space as the observed data. The following two assumptions are the identification conditions used to interpret Definition 1 causally.

Assumption 10 [ass:consistency] (Consistency).

Under , almost surely.

⊢ Lean

Assumption 10 is the standard consistency condition, linking the observed outcome to the potential outcome at the realized dose (Kennedy et al., 2017).

Assumption 11 [ass:no-unmeasured-confounding] (No unmeasured confounding).

For every , the potential outcome is independent of conditional on under .

⊢ Lean

Assumption 11 is the standard conditional ignorability condition for continuous treatments (Kennedy et al., 2017). Together with local positivity in Assumption 4, it identifies the causal dose-response mean at the interior dose with the observed-data partial mean in Definition 1.

Combining these restrictions gives the observed-data class used for the lower-bound analysis. The minimax risk and the published benchmark rate referenced in the definition are introduced formally in Definition 3 and Definition 4.

Definition 2 [def:holder-dose-class] (Hölder Dose Class ).

With the minimax risk and the published HOIF benchmark as in Definitions 3 and 4, define to be the collection of observed-data laws satisfying all of the following conditions:

  • (Sampling.) The observations satisfy Assumption 1.

  • (Causal identification.) Consistency and no unmeasured confounding hold as in Assumptions 10 and 11.

  • (Outcome boundedness.) The bounded-outcome condition holds as in Assumption 2.

  • (Interior dose.) The evaluation dose and local radius satisfy Assumption 3.

  • (Local positivity.) The treatment density obeys the local positivity condition with constant as in Assumption 4.

  • (Treatment-direction smoothness.) The maps and satisfy the treatment-direction Hölder conditions in Assumptions 5 and 6.

  • (Covariate-direction smoothness and density range.) The maps and satisfy the covariate-direction Hölder conditions in Assumptions 7 and 8, and satisfies the Hölder, nonnegativity, and upper-envelope condition in Assumption 9.

  • (Semantic ties to the observed-data law.) The regression, covariate-density, and treatment-density symbols are tied to : equals the conditional mean of given -almost surely; is the density of the -marginal on ; and is nonnegative on with the -law density relative to Lebesgue measure on .

⊢ Lean

Definition 2 is a statistical model for observed laws equipped with the stated causal interpretation. The lower-bound alternatives are constructed as observed-data laws in this class; the potential-outcome overlay supplies the causal reading of the same target while preserving the statistical experiment.

For the least favorable submodels used later, we also require a baseline law at a positive distance from the model boundary.

Assumption 12 [ass:baseline-submodel-slack] (Baseline submodel slack).

There exist a covariate density on , a conditional treatment density on , a constant , and a constant with , such that for all , for all , has Hölder norm at most of order on , and

⊢ Lean

Assumption 12 is specific to this analysis. It provides an interior baseline density , an interior baseline treatment density , and a positive margin , so that the perturbations used in the lower-bound construction can remain inside Definition 2.

Main results

The target in Definition 1 is evaluated under squared loss over the observed-data class in Definition 2. We first fix the risk criterion. The estimator is allowed to be any measurable function of the iid sample, with range restricted to the outcome envelope; this normalization is immaterial for the rate comparisons below because the target itself lies in the same bounded range.

Definition 3 [def:minimax-risk] (Minimax risk ).

The minimax risk over the model class at the evaluation dose is where is the product law of the iid sample .

⊢ Lean

The benchmark used for comparison is the rate reported for higher-order influence-function methods in the continuous-exposure literature. We use to denote .

Definition 4 [def:published-hoif-rate] (Published HOIF benchmark ).

The published higher-order influence-function benchmark rate is

⊢ Lean

Definition 4 records the comparison sequence from the published localized-regularity upper result. The first term is the one-dimensional nonparametric regression scale in the treatment coordinate, while the second is the higher-order influence-function benchmark involving the covariate dimension and covariate smoothness (Robins et al., 2008; Robins et al., 2015; Kennedy et al., 2017; Bonvini et al., 2022). The lower bound below is the paper’s same-class result.

Theorem 1 [thm:sharp-pointwise-lower-bound] (Sharp Pointwise Lower Bound).

For every covariate dimension and every real constants , suppose that

  • (Positive smoothness.) , , and .

  • (Regime constants.) , , , , and the local window determined by is contained in .

  • (Strict-slack baseline.) Assumption 12 holds for .

Then there exists a constant such that, for all sufficiently large sample sizes , where is the model class of Definition 2 and is the minimax risk of Definition 3.

⊢ Lean

Theorem 1 gives a lower floor for the original Hölder dose class, uniformly over every positive value of the treatment-density smoothness parameter . Its construction varies the treatment regression while keeping the density components within the slack allowed by Assumption 12. The result fixes the same-class treatment-regression obstruction and supplies the lower half of the minimax analysis.

The next statements reduce the comparison with to algebraic regimes. When the covariate smoothness is high enough relative to dimension, the maximum in Definition 4 is attained by the treatment-regression term.

Proposition 1 [prop:oracle-regime-reduction] (Oracle Regime Reduction).

Let be the covariate dimension, and let be constants. Suppose that

  • (Well-formed constants.) , , , , , , , and the window is contained in .

  • (Oracle regime.) .

  • (Baseline slack.) Assumption 12 holds for .

Then there exists a constant such that, for all sufficiently large sample sizes , for the published HOIF benchmark rate of Definition 4, and where the model class is Definition 2 and the minimax risk is Definition 3.

⊢ Lean

Proposition 1 gives the exact comparison statement for : the certified lower-bound exponent agrees with the exponent of the published benchmark sequence. Construction of an estimator over the same Hölder class is the corresponding upper-analysis question.

Theorem 2 [thm:sharp-minimax-smooth-covariate] (Smooth Covariate Minimax Floor).

Let and let . Suppose that:

  • (Smoothness.) , , and .

  • (Regime constants.) , , , , and the window is contained in .

  • (Smooth-covariate regime.) .

  • (Strict-slack baseline.) Assumption 12 holds for .

Then there exists a constant such that, for all sufficiently large , where is the minimax risk of Definition 3 over the same-class model of Definition 2. Moreover, with as in Definition 4.

⊢ Lean

Theorem 2 packages the same conclusion in the smooth-covariate regime: it establishes a lower bound for Definition 2 and an exact exponent comparison with .

When , the published benchmark is governed by the covariate-smoothness term. The lower bound remains on the treatment-regression scale, and the benchmark exponent is strictly smaller.

Theorem 3 [thm:frontier-bracket-deficient] (Deficient Frontier Bracket).

Fix a covariate dimension and constants . Assume:

  • (Regime constants.) , , , , , , , and the window lies in .

  • (Deficient covariate smoothness.) .

  • (Strict-slack baseline.) Assumption 12 holds for .

Then there exists a constant such that, for all sufficiently large sample sizes , where the minimax risk is as in Definition 3 and the class is Definition 2. Moreover, for the published HOIF benchmark of Definition 4, and its exponent is strictly smaller than the certified lower-bound exponent:

⊢ Lean

Thus, in the low-covariate-smoothness regime, Theorem 3 determines the strict ordering between the same-class lower floor and the exponent in the published benchmark sequence. The same-class upper frontier remains open among the benchmark rate, an intermediate rate, and a rate depending more directly on .

For orientation, we record the external higher-order influence-function result used as a literature comparator. The cited bound from Bonvini et al. (2022) is stated over an explicitly defined, more regular subclass of the Hölder dose class in Definition 2 and supplies the corresponding upper-side benchmark.

Cited result 1 [lem:published-upper-bound-cited] (Published HOIF Comparator).

For covariate dimension and parameters , consider any Bonvini–Kennedy HOIF estimator specification consisting of an observed-data law , evaluation dose , nuisance estimators , kernel , bandwidths , projection dimensions , and true and estimated projection kernels and . Suppose the following conditions of Bonvini et al. (2022) hold for these same inputs: is a probability law; , , and are respectively the outcome regression, covariate density, and conditional treatment density of ; there exist such that, for every , and ; there exists such that and -almost surely and for every ; and, for every and , the maps and lie in the one-dimensional Hölder ball of order and radius on , while and lie in the corresponding ball of order . Writing , assume also that there exist and such that , , , for , , , and, for every , ; that there exists with and for every ; and that there exist such that, for every , Assume further that , , , , , , and that there exists such that, for every and , the maps , , , and lie in the -Hölder ball of radius on . Let and . The rate specialization requires eventually, eventually, and eventually. The tuning satisfies, when , and eventually, and, when , and eventually. Then the specified Bonvini–Kennedy HOIF estimator , with its nuisance-training outputs held fixed, attains the published benchmark conditional mean-squared-error rate of Definition 4: there exists such that, for all sufficiently large ,

⊢ Lean

The notation inside the cited display belongs to the imported comparator. In the present paper, it motivates the sequence and provides an upper-side benchmark from a strictly smaller, more regular class; Definition 3 reserves for the same-class risk.

The final two remarks summarize the established rate comparison and identify the remaining same-class upper frontier.

Remark 1 [def:beta-frontier-handle] (Beta frontier handle ).

The beta comparison handle denotes the following frontier statement. For every , there is a constant such that, for all sufficiently large , For every , when , the published benchmark satisfies For every , when and , the published benchmark satisfies and its exponent obeys The upper-frontier alternatives are recorded as the following disjunction: either there is a constant such that, for all sufficiently large , or there are and with such that, for all sufficiently large , or there are , , exponents with , and positive constants such that, for all sufficiently large , and

⊢ Lean

Remark 1 is an interpretive shorthand for the preceding lower-bound and comparison statements. The word “frontier” denotes this exponent-level comparison and highlights the benchmark-aligned smooth-covariate regime.

Discussion, limitations, and extensions

The lower-bound analysis isolates one obstruction that is already present in the standard interior partial mean problem. The target in Definition 1 is an observed-data partial mean at a fixed interior dose, and the risk in Definition 3 is taken over the same Hölder dose class of Definition 2. Within that class, Theorem 1 gives a treatment-regression lower floor. The comparison statements in Theorem 2 and Theorem 3 relate this floor to the published benchmark sequence from Definition 4, while Cited result 1 supplies the external Bonvini–Kennedy comparator on its localized-regularity class.

For each fixed , under the corresponding slack-baseline condition in Assumption 12, the exponent of the same-class lower bound is invariant to the treatment-density smoothness index. The constant in the lower bound, and the feasibility of the slack baseline itself, may depend on and on the other fixed model constants. Thus the conclusion is exponent-level insensitivity to , with constants calibrated separately for each fixed smoothness value.

This distinction matters most when . In that regime, Theorem 3 shows that the published benchmark sequence has a smaller exponent than the same-class lower-bound exponent. The unresolved same-class question is whether the upper behavior in low covariate smoothness follows the published benchmark, an intermediate exponent, or a rate that depends more directly on the smoothness of the treatment density.

The scope is tied to the interior positivity regime. Assumption 4 imposes a fixed lower bound on the treatment density in a neighborhood of , and Assumption 3 selects an interior evaluation window. These restrictions match the usual interior continuous-exposure formulation. Weak-overlap, design-adaptive, and trimmed targets lead to a different statistical experiment in which the amount of information near enters the rate calculation directly. Recent work on continuous exposures and overlap-sensitive causal estimands emphasizes these complementary regimes (Hejazi et al., 2022; Doss et al., 2022; Hudson et al., 2023; Shi et al., 2024; Zhang et al., 2024; Schindl et al., 2024; Bonvini et al., 2024).

Neighboring estimands raise a similar qualification. The partial mean studied here evaluates the regression surface at a single interior dose and averages over the marginal law of . Other continuous-exposure targets average over dose neighborhoods, consider stochastic interventions, or target modified exposure policies. Their definitions can smooth the treatment coordinate or alter the role of the conditional treatment density. The present lower bound is consequently a benchmark for the fixed-dose interior partial mean, while those neighboring functionals invite estimand-specific minimax analyses.

Finally, the comparison with higher-order influence-function methods is a comparison across stated model classes. The imported result in Cited result 1 supplies an upper-side benchmark under additional localized regularity, while the lower bound here is proved over the Hölder class in Definition 2. Two routes can bridge the statements: a same-class upper bound under Definition 2, or an inclusion argument covering the present Hölder class by the localized-regularity class with the needed constants.

Appendices

auxiliary lemmas and proofs

This appendix collects the auxiliary mathematical statements used to support the lower-bound and rate-comparison results in Theorem 1–Theorem 3. The statements are organized according to their role: first the two-point testing ingredients, then the algebra for the benchmark sequence , and finally the combined lower-bound comparison used to package the main conclusions. Throughout this appendix, denotes the Kullback–Leibler divergence from to .

The testing reduction begins with a bounded binary outcome channel. It is convenient because the least favorable alternatives only need to separate conditional means while keeping outcomes inside the common envelope. The following elementary inequality controls the information distance between two such channels by the squared separation of their means.

Lemma 1 [lem:bernoulli-mean-channel-kl] (Bernoulli mean-channel KL bound).

Fix . Let and be distributions on with means and , respectively, so that If and , then

⊢ Lean
Proof of Lemma 1.

Throughout, denotes the two-point law on with Set so that, by the definition of the channels in the statement,

Step 1 (the two success probabilities lie in the central band). Since , and the hypotheses and give and . Hence

Step 2 (transport to the parametrization). Let Because , the map is an affine bijection of the real line whose inverse is again affine, hence measurable, and , . Consequently the image of under puts mass at and mass at , which is exactly ; the same computation with in place of gives as the image of :

Applying one and the same measurable bijection to both arguments leaves the Kullback–Leibler divergence unchanged, since the Radon–Nikodym derivative of the two image laws is the transport of the original one along . Therefore

Step 3 (quadratic band bound for the Bernoulli divergence). For the divergence of two-point laws is the explicit sum and on the central band this is dominated by a quadratic: Indeed, writing for the negative Bernoulli entropy, the left-hand side is the Bregman remainder , so the assertion is that the function is nonnegative at ; and it is, because , , and is convex on , its second derivative there being so that minimizes over that interval.

Step 4 (back to the mean parametrization). Step 1 supplies the band hypothesis needed in Step 3, and so the last inequality because . Chaining Steps 2–4, which is the claim.

Lemma 1 supplies the local information calculation for the outcome part of the two-point experiment. The boundedness condition keeps both Bernoulli probabilities uniformly away from zero and one, so the quadratic control of the divergence is uniform over the perturbations used in the construction.

The next ingredient is the standard two-point minimax testing bound. It translates a separation in the parameter into a lower bound on mean-squared error whenever the two induced sample laws remain close in information distance; this is the usual Le Cam device for nonparametric lower bounds (Tsybakov, 2009).

Lemma 2 [lem:le-cam-two-point-mse-source] (Le Cam Two-Point Bound).

Let and be two laws for the -sample experiment, and let and be the corresponding values of a real parameter. Set If then there exists a constant , depending only on , such that every estimator satisfies

⊢ Lean
Proof of Lemma 2.

Throughout the proof, denotes the extended-real Kullback–Leibler divergence from to ; when this divergence is finite, its real value is Write for total variation. Let and define, before any laws or estimator are chosen, Fix probability laws on a common measurable sample space satisfying the finite-budget hypothesis, and fix a measurable estimator such that the two squared losses under and under are integrable. Put

Step 1 (estimation to testing). Since , the two measurable error sets cover the sample space in the following sense: if fails, then , and the triangle inequality gives Thus . The testing characterization of total variation therefore gives Since a sum of two real numbers is at most twice their maximum,

Step 2 (a testing floor for a finite budget). The budget hypothesis is an inequality in the extended nonnegative reals against the finite value associated with the real number . Hence the divergence is finite, so , and its real value satisfies The Bretagnolle–Huber affinity bound applies under these two side conditions and yields Because , monotonicity of the exponential gives . Combining this inequality with the preceding display and Step 1 gives

Step 3 (testing back to squared error). For , the event is the event . The squared loss is nonnegative and integrable under the corresponding law, so the integral form of Markov’s inequality gives Writing and distinguishing whether or , these two inequalities imply

Step 4 (conclusion). Multiplying the lower bound from Step 2 by and using Step 3, Since , Therefore The constant was chosen from alone, before the laws and the estimator, and the argument applies to every measurable estimator with the two stated integrability properties. This is the claimed two-point bound.

Lemma 2 is used only at the level stated here. In the lower-bound construction, the alternatives are chosen so that their target values differ by of the desired order, while the product-law divergence is bounded by a fixed constant . The resulting mean-squared error lower bound is therefore proportional to .

We next record the algebraic reductions for the comparison sequence from Definition 4. These statements identify exactly which exponent appears in the published benchmark under the two covariate-smoothness regimes.

Lemma 3 [lem:rho-oracle-regime-algebra] (Smooth-Covariate Rate Collapse).

Let be the published HOIF benchmark of Definition 4, evaluated at . Suppose that

  • (Positive treatment smoothness.) .

  • (Positive covariate smoothness.) .

  • (Smooth-covariate regime.) .

  • (Sample size.) .

Then

⊢ Lean
Proof of Lemma 3.

Throughout, is the maximum of the two power terms in Definition 4, so the claim is that the first term dominates.

Case . Every real power of equals , so both terms coincide, and hence , which is the asserted identity.

Case . Since we have , so the smooth-covariate hypothesis can be divided through: Adding to both sides, Both sides are strictly positive, since and by . Dividing the constant by the two sides therefore reverses the inequality:

Multiplying numerator and denominator of the left-hand side by gives the elementary identity so that Since , the map is strictly increasing; negating the two exponents reverses the last inequality and then applying preserves it, so Thus the maximum in Definition 4 is attained by the treatment-regression term, that is,

Lemma 3 states that, when , the maximum defining is attained by the treatment-regression term. This is the algebra behind the smooth-covariate comparison in Proposition 1 and Theorem 2.

Lemma 4 [lem:rho-deficient-regime-algebra] (Deficient Regime Algebra).

For any sample size , treatment smoothness , covariate smoothness , and covariate dimension , suppose that

  • (Positive treatment smoothness.) ;

  • (Positive covariate smoothness.) ;

  • (Deficient-covariate regime.) ;

  • (Nonzero sample size.) .

Then the published HOIF benchmark from Definition 4 satisfies and the deficient-regime exponent is strictly smaller than the oracle exponent:

⊢ Lean
Proof of Lemma 4.

Step 1 (strict comparison of the two exponents). Since we have , so the deficient-regime hypothesis can be divided through: Adding to both sides, Both sides are strictly positive, because by . Dividing the constant by the two sides therefore reverses the strict inequality: Multiplying numerator and denominator by gives the elementary identity and hence the strict exponent comparison asserted in the statement,

Step 2 (identification of the maximum in ). If , every real power of equals , so both terms of Definition 4 coincide, and , as claimed. If instead , the map is strictly increasing; negating the two exponents reverses the strict inequality of Step 1, and applying then preserves the reversed inequality, so In either case the maximum in Definition 4 is attained by the covariate-smoothness term, that is, Together with the strict comparison of Step 1, this is the desired conclusion.

Lemma 4 gives the complementary algebra for . In that regime the published benchmark is governed by the covariate-smoothness expression, and its exponent is smaller than the exponent appearing in the same-class lower floor. The lemma thereby supplies the deficient-regime exponent comparison.

The lower-bound input for the main results is the oracle floor obtained by perturbing the treatment regression while keeping the density components within the slack baseline introduced in Assumption 12. The construction is the usual local nonparametric testing argument in the treatment coordinate (Stone, 1982; Tsybakov, 2009; Goldenshluger et al., 2020), specialized to the interior partial-mean functional.

Lemma 5 [lem:oracle-dose-regression-lower-all-beta] (All-Beta Oracle Floor).

Fix a covariate dimension and constants . Suppose that:

  • (Positive smoothness.) , , and .

  • (Regime constants.) , , , , and the window of radius around lies in the interior dose region.

  • (Slack baseline.) The strict-slack baseline condition of Assumption 12 holds for .

Then there exists a constant such that, for all sufficiently large , Here is the model class of Definition 2, and is the minimax risk of Definition 3.

⊢ Lean
Proof of Lemma 5.

Step 1 (baseline, outcome scale, and the reference bump). By the strict-slack baseline condition of Assumption 12, fix a covariate density on , a conditional treatment density on , and a slack with the properties listed there. Taking the order- bound in the Hölder condition on and combining it with gives, in particular, so . Put Let be a fixed smooth reference bump with in particular , , and has compact support, so all of its derivatives are bounded.

Step 2 (the amplitude gate). There is an amplitude with such that, for every sign and every bandwidth , the scaled signed bump lies in the treatment-direction Hölder ball required by Definition 2: Indeed, write for the largest integer strictly below , let for , let be an -Hölder constant of on , and set For the -th derivative of the scaled bump at is , and because and ; hence its modulus is at most . For the top order, writing and using , which is exactly the top-order Hölder requirement of radius .

Step 3 (the divergence budget and the constant). Set let be the constant supplied by Lemma 2 for this budget, and define

Step 4 (bandwidth). For put Since we have , so ; hence for all sufficiently large both and hold. Moreover, for every , Fix such an from here on.

Step 5 (the two least favourable laws). For define the perturbed treatment regression and let be the observed-data law generated as follows: draw on with density ; independently draw on with density ; and, given , draw from the two-point channel on of Lemma 1. The potential-outcome overlay is By construction the law-side nuisances of are Write and . Because , , and (so ), so in particular and each channel is a genuine probability measure on .

Step 6 (both laws lie in the model class). We check the conditions listed in Definition 2 for , .

  • Sampling. with , with , and each outcome channel is a probability measure by the last display of Step 5; so is a probability law under which and almost surely, and the sample is independent draws from it.

  • Bounded outcome. , so almost surely.

  • Interior dose. is a hypothesis of the present lemma.

  • Local positivity. for every in the window and every , by Assumption 12 and .

  • Treatment-direction smoothness of the regression. This is exactly the conclusion of Step 2, since does not depend on .

  • Treatment-direction smoothness of the density. has Hölder norm at most of order on the window, and , so .

  • Covariate-direction smoothness of the regression. Since , a constant function of of modulus at most by Step 5; a constant of modulus at most lies in , all its derivatives vanishing.

  • Covariate-direction smoothness of the density. Likewise is constant in , of modulus at most , hence in .

  • Covariate density. , and on the cube.

  • Semantic ties. The channel has mean , so ; the -marginal of is times Lebesgue measure on ; and the -law has Lebesgue density on , which is exactly the required factorization.

  • Causal overlay. Consistency holds by the definition of in Step 5. For ignorability, fix ; because has a Lebesgue density the event is null, so almost surely, a deterministic function of alone; hence is conditionally independent of given .

Therefore .

Step 7 (one-observation divergence). The two laws share the same -law, with density , and differ only in the conditional outcome channel; since the observed pair is a function of the observation and each fibre channel is dominated by its counterpart, the chain rule for divergence gives Fibrewise, and by Step 5, so Lemma 1 applies with and ; as it yields The bump concentrates the remaining integral: whenever , so the integrand vanishes outside ; on that interval, which is contained in the window because , we have and by Step 1; and the interval has length . Hence

Using , the two previous displays combine to the one-observation estimate

Step 8 (tensorization and the budget). Each fibre channel puts mass at least on each of the two atoms , because ; consequently and are mutually absolutely continuous with integrable log-likelihood ratio, and the divergence between the -fold product laws tensorizes: the last equality by the calibration of Step 4.

Step 9 (functional separation). Since and , so the two target values are separated by

Step 10 (two-point reduction inside the minimax risk). Let be any estimator admissible in Definition 3, that is, measurable with values in . For every in the model class, and , and has unit Lebesgue measure, so

Thus under every -fold product law the squared loss is bounded by , so it is integrable and the worst-case risk over the class is a finite supremum. Apply Lemma 2 with the -sample laws , , parameter values , budget (available by Step 8) and estimator : Both and belong to the class by Step 6, so the right-hand side is at most the supremum over of the risk of . As was arbitrary, taking the infimum over admissible estimators preserves the bound, and with the separation of Step 9,

Step 11 (conclusion). Finally, by the choice of in Step 4, Combining the last two displays, for all sufficiently large , with depending only on the fixed model constants and the slack baseline, not on .

Lemma 5 is the same-class lower-bound engine. Its conclusion is uniform in the sense relevant for the paper’s exponent comparison: for each fixed admissible , the lower exponent is . The constant may depend on the fixed model parameters and on the slack available in Assumption 12; the stated uniformity concerns each fixed .

verification note

This appendix records the verification scope for the formal statements used in the paper. The observed-data experiment, causal overlay, Hölder dose class, target functional, minimax risk, published benchmark sequence, auxiliary testing statements, algebraic rate comparisons, and lower-bound consequences are represented by the numbered assumptions, definitions, lemmas, propositions, theorems, and remarks displayed in the paper. The corresponding theorem statements and their derivations are machine-checked in Lean 4.

The machine-checked part covers the consequences of the stated assumptions and auxiliary lemmas as they are formulated here. In particular, Assumption 1–Assumption 12 define the statistical and causal restrictions used by the displayed results; Definition 1, Definition 2, Definition 3, and Definition 4 fix the target, model class, risk, and comparison sequence; and the auxiliary appendix lemmas supply the testing and algebraic inputs used to derive Theorem 1, Theorem 2, and Theorem 3 together with Proposition 1. The verification establishes the same-class lower-bound and rate-comparison statements, with the upper frontier identified as a separate research question.

Several ingredients enter as mathematical assumptions or cited comparators. The causal identification primitives in Assumption 10 and Assumption 11 are imposed conditions on the observed-data law with its potential-outcome interpretation. The formal argument uses the classical Le Cam minimax reduction at the level of the auxiliary testing lemma stated in the appendix and verifies its consequences for the present model and rates. The Bonvini–Kennedy result in Cited result 1, imported from Bonvini et al. (2022), motivates the benchmark sequence on its localized-regularity class; Definition 2 specifies the distinct same-class lower-bound model studied here.

The Lean-oriented labels are therefore confined to cross-references and this reproducibility note. In the main text they serve only to identify the precise assumptions and statements being invoked. The econometric content of the paper is the lower bound over Definition 2, the algebraic comparison with Definition 4, and the explicit separation between those same-class lower results and the external higher-order influence-function comparator.

Proofs of the main results

Proof of Theorem 1.
  1. Instantiate the oracle floor. The theorem hypotheses provide the inputs required by Lemma 5 for the same covariate dimension and the same constants . The positive-smoothness assumptions give , , and . The regime-constants assumptions give , , , , and the interior-window condition recorded in Assumption 3. The strict-slack baseline assumption is Assumption 12 for the tuple , matching the slack-baseline input of the lemma.

  2. Apply the oracle floor. Applying Lemma 5 with these inputs yields a constant such that, for all sufficiently large , Here the class is the model class of Definition 2, and is the minimax risk of Definition 3. The displayed tail bound is the conclusion of the cited oracle floor for the fixed treatment-density smoothness parameter .

  3. Conclusion. Set . Then , and the inequality from Step 2 is equivalently for all sufficiently large sample sizes . This is the asserted positive constant and eventual lower bound over the model class of Definition 2 for the minimax risk of Definition 3.

Proof of Proposition 1.
  1. The positive-smoothness, regime-constant, and baseline-slack hypotheses assumed here – , , , , , , , , and Assumption 12 for – are the hypotheses required by Theorem 1. Applying that result gives a constant such that, for all sufficiently large sample sizes , Fix this for the rest of the proof.

  2. Intersect the eventual lower-bound event supplied by Theorem 1 with the eventual sample-size event . Equivalently, after increasing the sufficiently-large threshold if necessary, every sample size under consideration satisfies both and the lower-bound inequality supplied by Theorem 1.

  3. For any such , the hypotheses , , , and are the hypotheses of Lemma 3. Hence the published benchmark rate satisfies

  4. For the same , the lower-bound inequality supplied by Theorem 1 gives The identity supplied by Lemma 3 rewrites as , and therefore This holds on the same eventual set of sample sizes. Together with and the displayed identity for , it is the assertion of Proposition 1.

Proof of Theorem 2.
  1. The smoothness, regime-constant and strict-slack hypotheses assumed here — , , , , , , , , and Assumption 12 for — are exactly the hypotheses of Theorem 1; the smooth-covariate condition is not needed for this step and is used only in Step 3. That theorem therefore supplies a constant and a threshold with Fix this ; it is the constant asserted in the statement.

  2. Set . For every both the displayed lower bound of Step 1 and the elementary condition hold; this is the intersection of two eventual conditions, and the conjunction asserted by the theorem is claimed only for such .

  3. Fix . The hypotheses , , and are exactly those of Lemma 3, so which is the second assertion of the theorem, in the orientation in which it is stated. Holding this identity together with the lower bound of Step 1 on the same range gives, for the same constant , the asserted eventual conjunction of the lower bound and the exponent identity.

Proof of Theorem 3.
  1. The regime-constant and strict-slack hypotheses assumed here — , , , , , , , , and Assumption 12 for — are exactly the hypotheses of Lemma 5; the deficient-covariate condition is not used for this step and enters only in Step 3. That lemma therefore supplies a constant and a threshold with Fix this ; it is the constant asserted in the statement. Note that the lower floor is established for the given fixed and its exponent does not involve .

  2. Set . For every both the display of Step 1 and the elementary condition hold; this is the intersection of two eventual conditions, and the conjunction asserted by the theorem is claimed only for such .

  3. Fix . The hypotheses , , and are exactly those of Lemma 4, which gives both the deficient-regime formula and the strict comparison of exponents

  4. On the common range , the lower bound of Step 1, rewritten in the orientation of the statement as holds together with the two conclusions of Step 3. That conjunction, with the constant fixed in Step 1, is exactly the assertion of the theorem: the certified lower floor sits at the treatment-regression exponent , while the published benchmark sequence is the covariate-smoothness power, whose exponent is strictly smaller.

References

  • Rubin, Donald B. (1974). Estimating Causal Effects of Treatments in Randomized and Nonrandomized Studies. Journal of Educational Psychology. doi
  • Robins, James (1986). A new approach to causal inference in mortality studies with a sustained exposure period—application to control of the healthy worker survivor effect. Mathematical Modelling. doi
  • Bickel, Peter J. and Klaassen, Chris A. J. and Ritov, Ya'acov and Wellner, Jon A. (1993). Efficient and Adaptive Estimation for Semiparametric Models. Springer.
  • van der Vaart, Aad W. (1998). Asymptotic Statistics. Cambridge University Press. doi
  • van der Laan, Mark J. and Robins, James M. (2003). Unified Methods for Censored Longitudinal Data and Causality. Springer.
  • Imbens, Guido W. (2000). The Role of the Propensity Score in Estimating Dose-Response Functions. Biometrika. doi
  • Hirano, Keisuke and Imbens, Guido W. (2004). The Propensity Score with Continuous Treatments. Applied Bayesian Modeling and Causal Inference from Incomplete-Data Perspectives.
  • Imai, Kosuke and van Dyk, David A (2004). Causal Inference With General Treatment Regimes. Journal of the American Statistical Association. doi
  • Stone, Charles J. (1982). Optimal Global Rates of Convergence for Nonparametric Regression. The Annals of Statistics. doi
  • Tsybakov, Alexandre B. (2009). Introduction to Nonparametric Estimation. Springer. doi
  • Robins, James and Li, Lingling and Tchetgen Tchetgen, Eric and van der Vaart, Aad (2008). Higher Order Influence Functions and Minimax Estimation of Nonlinear Functionals. IMS Collections. doi
  • Robins, James and Li, Lingling and Mukherjee, Rajarshi and Tchetgen Tchetgen, Eric and van der Vaart, Aad (2015). Higher Order Estimating Equations for High-dimensional Models. . doi
  • Kennedy, Edward H. and Ma, Zongming and McHugh, Matthew D. and Small, Dylan S. (2017). Nonparametric Methods for Doubly Robust Estimation of Continuous Treatment Effects. Journal of the Royal Statistical Society: Series B (Statistical Methodology). doi
  • Zhao, Shandong and van Dyk, David A. and Imai, Kosuke (2013). Causal Inference in Observational Studies with Non-Binary Treatments. . doi
  • Chernozhukov, Victor and Chetverikov, Denis and Demirer, Mert and Duflo, Esther and Hansen, Christian and Newey, Whitney and Robins, James (2018). Double/Debiased Machine Learning for Treatment and Structural Parameters. The Econometrics Journal. doi
  • Lee, Ying-Ying (2018). Partial Mean Processes with Generated Regressors: Continuous Treatment Effects and Nonseparable Models. . doi
  • Colangelo, Kyle and Lee, Ying-Ying (2020). Double Debiased Machine Learning Nonparametric Inference with Continuous Treatments. . doi
  • Goldenshluger, Alexander and Lepski, Oleg (2020). Minimax Estimation of Norms of a Probability Density: I. Lower Bounds. . doi
  • Hejazi, Nima S. and Benkeser, David and D{\'i}az, Iv{\'a}n and van der Laan, Mark J. (2022). Efficient Estimation of Modified Treatment Policy Effects Based on the Generalized Propensity Score. . doi
  • Doss, Charles R. and Weng, Guangwei and Wang, Lan and Moscovice, Ira and Chantarat, Tongtan (2022). A Nonparametric Doubly Robust Test for a Continuous Treatment Effect. . doi
  • Bonvini, Matteo and Kennedy, Edward H. (2022). Fast Convergence Rates for Dose-Response Estimation. . doi
  • Hudson, Aaron and Geng, Elvin H. and Odeny, Thomas A. and Bukusi, Elizabeth A. and Petersen, Maya L. and van der Laan, Mark J. (2023). An Approach to Nonparametric Inference on the Causal Dose Response Function. . doi
  • Bonvini, Matteo and Kennedy, Edward H. and Dukes, Oliver and Balakrishnan, Sivaraman (2024). Doubly-Robust Inference and Optimality in Structure-Agnostic Models with Smoothness. . doi
  • Zhang, Yikun and Chen, Yen-Chi and Giessing, Alexander (2024). Nonparametric Inference on Dose-Response Curves Without the Positivity Condition. . doi
  • Shi, Junming and Zhang, Wenxin and Hubbard, Alan E. and van der Laan, Mark J. (2024). HAL-Based Plug-in Estimation with Pointwise Asymptotic Normality of the Causal Dose-Response Curve. . doi
  • Schindl, Kyle and Shen, Shuying and Kennedy, Edward H. (2024). Incremental Effects for Continuous Exposures. . doi