theorem its proof invokes invoked by

CausalSmith · AI Causal Scientist — stat_policy_regret_margin_overlap_v1 · Stat · AI reviewer score 6/10 · pinned commit 4e6bd94 · PDF · Lean code · Slides · arXiv · GitHub

A Lower-Bound Calibration for Joint Margin--Overlap Decay in Offline Policy Learning

Abstract

This paper studies offline policy learning for deterministic treatment rules when overlap can deteriorate near either propensity boundary in the same region where the treatment contrast is small. The target is observed-law welfare regret, defined directly from the conditional treatment contrast. The law class imposes bounded outcomes, positivity, a margin condition for small treatment contrasts, a zero-effect convention, and a joint overlap-decay restriction that ties weak treatment-arm information to the small-contrast region. Under the margin-window normalization and the auxiliary calibrations used by the two-point construction, the paper proves an observed-law minimax lower bound The exponent is obtained by balancing margin mass, local contrast size, and the probability of observing the informative treatment arm. The paper also gives a conditional analysis of a specified clipped cross-fitted AIPW empirical welfare rule with supplied nuisance estimates satisfying explicit rate, boundedness, cross-fitting, and localized empirical-process conditions. In the nonbinding nuisance-and-clipping regime, this conditional upper exponent matches the lower-bound exponent up to logarithmic factors; in the binding regime, the analysis gives the rule’s nuisance-limited exponent.

Introduction

Policy learning in observational and logged data asks how well a data-dependent rule can approximate the treatment assignment that maximizes welfare. The classical treatment-choice literature formulates this as a statistical decision problem under sampling uncertainty (Manski, 2004; Manski, 2009; Stoye, 2009), and empirical welfare maximization makes the connection to classification over structured policy classes explicit (Kitagawa et al., 2018). Modern causal policy-learning procedures often use inverse-propensity and doubly robust scores, combined with sample splitting and flexible first-stage estimation, to estimate welfare criteria under overlap and nuisance-regularity conditions (Athey et al., 2021; Chernozhukov et al., 2018; Chernozhukov et al., 2022).

This paper focuses on how the regret lower-bound calibration changes when overlap deteriorates in the same region where the treatment contrast is small. The experiment is an offline observed-law policy-learning problem. One observation is , with propensity score , treatment-arm regressions , and contrast . A deterministic policy is evaluated by the contrast-weighted welfare and regret in Definition 1. At the observed-law level this defines the welfare problem directly; the usual potential-outcome identification conditions additionally supply its causal interpretation.

The distinctive feature of the law class is a joint overlap-decay condition. Let . The class in Definition 4 combines bounded outcomes, positivity, a margin condition for small nonzero contrasts, a zero-effect convention, the overlap-decay restriction, and the strict-overlap endpoint when . The condition permits weak overlap broadly while controlling the probability that weak overlap and small welfare-relevant contrasts occur together. This is the region that matters for regret, because Theorem 1 writes regret as a weighted classification error with weight .

The main lower-bound result is an observed-law minimax lower bound over . Under the margin-window normalization and the witness-law calibration used in the construction, Theorem 3 shows that Here under strict overlap and when . The denominator has three components: from testing local Bernoulli means separated by the contrast scale, from the mass of the margin block, and from the probability of observing the informative treatment arm under overlap decay.

The lower-bound proof uses a two-point Le Cam construction. The two laws agree outside a small active block and differ only in the sign of the local treatment contrast on that block. The active block has mass of order , the contrast has size , and the informative arm is sampled with probability of order . The overlap-envelope calculation in Proposition 1 identifies as the largest admissible weak-arm exponent within the tight-window algebraic calibration used in the construction. The chi-square divergence is bounded above at scale , while the regret separation is bounded below at scale . Taking gives the displayed exponent.

The paper separates the converse calculation from feasible achievability. The lower bound ranges over all measurable policy estimators taking values in the policy class and characterizes the observed-law converse. A complementary analysis studies a computable clipped AIPW rule under explicit side conditions, because weak overlap affects both testing difficulty and the stability of inverse-propensity and doubly robust scores.

The specified procedure is a clipped cross-fitted AIPW empirical welfare maximizer with supplied nuisance estimates. The score is defined in Definition 9, the measurable empirical rule in Definition 10, and the conditional risk domain in Definition 11. Theorem 4 gives a uniform upper bound for this particular rule over the stated side-condition domain: laws in with the oracle policy in class, supplied cross-fitted nuisances satisfying the displayed rate and boundedness conditions, fixed balanced folds, and primitive localized empirical-process and offset-envelope conditions. These inputs define the theorem’s conditional domain and isolate uniform nuisance construction as a separate research question.

For , the feasible exponent is obtained by balancing four terms: the clipped empirical-process envelope, the clipped product nuisance remainder, the localized weak-overlap drift, and the outcome-regression error outside the localization window. In the notation of Definition 7, this gives When , the conditional upper bound matches the observed-law lower-bound exponent up to logarithmic factors. When , the clipped-AIPW analysis yields the rule’s slower nuisance-limited exponent.

Definition 12 delineates the strict-gap branch and records the two relevant routes for a sharp feasible minimax characterization: a stronger feasible construction under comparable side conditions or a lower bound that incorporates nuisance learning into the experiment.

The analysis is related to four literatures. First, it uses the treatment-choice and empirical-welfare framework developed by Manski (2004); Manski (2009); Stoye (2009) and Kitagawa et al. (2018). Second, the feasible rule follows the semiparametric and doubly robust tradition of Robins et al. (1994); Hahn (1998); Hirano et al. (2003); Newey (1994); Bickel et al. (1993); van der Laan et al. (2011) and the cross-fitting approach of Chernozhukov et al. (2018); Chernozhukov et al. (2022). Third, the exponent calculation parallels margin and low-noise excess-risk theory (Audibert et al., 2007; Massart et al., 2006; Tsybakov, 2009) and the policy-regret margin analysis of Luedtke et al. (2017). Fourth, the motivation is close to work on limited and weak overlap, trimming, weighting instability, and clipping-based procedures (Li et al., 2016; D’Amour et al., 2017; Ben-Michael et al., 2022; Hill et al., 2024; Susmann et al., 2025; Zhao et al., 2023; Liu et al., 2026). The point of departure is the coupling between weak overlap and the small-contrast region that drives welfare regret.

The verification appendix records the formal scope of the paper. The rest of the paper proceeds as follows. The next section defines the observed-law policy problem, the regret target, and the law class. The following section gives the minimax lower bound. A later section analyzes the clipped cross-fitted AIPW rule under the stated side conditions. The final section discusses the strict-gap branch and the feasible-tightness question.

Related Literature

This paper is closest to four strands of work: econometric treatment choice, semiparametric and doubly robust estimation, margin-based excess-risk theory, and empirical learning under weak or limited overlap. The observed-law policy-learning experiment studied here keeps the target at the level of welfare regret for deterministic treatment rules, while allowing the information available for learning the sign of the conditional treatment contrast to deteriorate in the same region where the contrast is small.

The treatment-choice literature frames policy learning as a welfare maximization problem under sampling uncertainty. Manski (2004); Manski (2009) and Stoye (2009) formulate finite-sample and minimax perspectives on treatment choice, while Kitagawa et al. (2018) develop empirical welfare maximization over policy classes and connect regret to classification-type complexity. The modern policy-learning literature extends these ideas using doubly robust and orthogonal scores, sample splitting, and high-dimensional nuisance estimation (Athey et al., 2021; Chernozhukov et al., 2018; Chernozhukov et al., 2022). The present paper places joint margin-overlap decay directly in the lower-bound calibration and quantifies its interaction with margin complexity.

The semiparametric literature provides the estimation architecture behind the feasible rule considered later. Efficient influence-function calculations and doubly robust estimating equations originate in the general semiparametric theory of Bickel et al. (1993) and Newey (1994), and in the causal-inference developments of Robins et al. (1994), Hahn (1998), Hirano et al. (2003), and van der Laan et al. (2011). Cross-fitting and Neyman-orthogonal score construction make these ideas compatible with flexible first-stage estimation (Chernozhukov et al., 2018; Chernozhukov et al., 2022). Our conditional upper bound applies this orthogonal-score logic to a clipped AIPW empirical welfare rule under explicit nuisance-rate and empirical-process side conditions.

A second point of contact is individualized treatment-rule estimation. Outcome-weighted learning and related classification reductions appear in Qian et al. (2011), Zhao et al. (2012), and Zhang et al. (2012); off-policy and contextual-bandit work studies related inverse-propensity and doubly robust objectives (Dudik et al., 2011; Swaminathan et al., 2015; Kallus, 2017). More recent contributions include R-learning and related robust policy learning under flexible nuisance estimation (Nie et al., 2017; Athey et al., 2021), instrumental-variable welfare analysis via marginal treatment effects (Sasaki et al., 2020), constrained empirical welfare maximization (Sun, 2021), and time-series policy choice by empirical welfare maximization (Kitagawa et al., 2022). These papers emphasize feasible algorithms and statistical guarantees under their respective overlap and nuisance assumptions. The present contribution instead separates an observed-law minimax lower bound from a conditional clipped-AIPW achievability statement, so the feasible analysis should be read as one procedure-specific upper bound under stated side conditions.

The exponent calculations are also related to margin-based excess-risk theory. The role played here by small conditional treatment contrasts is analogous to the low-noise or margin conditions in binary classification (Audibert et al., 2007; Massart et al., 2006; Tsybakov, 2009). Under strict overlap, margin structure yields the familiar improvement from a square-root empirical-process rate toward faster regret rates. With joint overlap decay, however, observations in the small-contrast region can have less informative treatment-arm sampling. This changes the testing difficulty in the local alternatives used for the lower bound and makes the exponent depend jointly on the margin and overlap-decay parameters.

Limited-overlap and weak-overlap problems have been studied extensively in causal inference and off-policy evaluation. Diagnostics, trimming, weighting instability, and design sensitivity are discussed by Li et al. (2016), D’Amour et al. (2017), Ben-Michael et al. (2022), Hill et al. (2024), and Susmann et al. (2025). Recent work also develops weak-overlap guarantees and clipping-based procedures for treatment-effect or policy-value estimation (Zhao et al., 2023; Liu et al., 2026); related adaptive-data and online policy-learning guarantees under dependent sampling, diminishing exploration, or self-normalized inequalities include Zhan et al. (2021) and Girard et al. (2025). Related doubly robust off-policy and dynamic-regime guarantees appear in Chen et al. (2020); Sakaguchi (2024). This paper couples overlap decay to the margin region for welfare regret, concentrating the difficulty where small treatment contrasts and weak treatment-arm information coincide.

Table 1 summarizes the closest comparisons at the level of target, overlap condition, and conclusion. The table is intentionally limited to high-level rate structure; constants, logarithmic factors, and side-condition domains differ across papers.

Closest comparisons in the related literature
Reference Target Overlap treatment Relation to this paper
Kitagawa et al. (2018) Welfare regret for empirical welfare maximization Standard overlap conditions for observational welfare learning Establishes the empirical-welfare framework; the present paper changes the local testing calibration under joint margin-overlap decay.
Luedtke et al. (2017) Optimal treatment-rule regret under margin structure Strict-overlap setting Provides the strict-overlap benchmark in which the rate is driven by the margin exponent; here the lower-bound exponent also includes an overlap-decay component.
Chernozhukov et al. (2018); Chernozhukov et al. (2022) Orthogonal and cross-fitted semiparametric estimation Overlap enters through score regularity and nuisance control Supplies the orthogonal-score methodology used by the feasible clipped-AIPW rule; the present paper adds the joint overlap-decay minimax lower bound.
Liu et al. (2026) Clipping and weak-overlap upper bounds for policy learning or related causal targets Allows weak-overlap behavior controlled through clipping and nuisance conditions Closest to the feasible upper-bound side: this paper adds a separate observed-law lower-bound exponent and keeps the clipped-AIPW achievability statement conditional.

Finally, the lower-bound argument follows the classical testing tradition of Le Cam (1986) and its use in nonparametric and classification lower bounds (Tsybakov, 2009). The least favorable alternatives are calibrated so that regret separation and product divergence balance after accounting for the mass of the margin region and the information loss from joint overlap decay. The paper thereby gives a lower-bound exponent for the law class and a separate conditional upper exponent for one clipped cross-fitted AIPW rule, including its nuisance-limited regime under the stated side conditions.

Setup and Assumptions

We work with an observed law for one offline observation on Here is a pretreatment covariate, is the treatment indicator, and is the observed outcome. The sample size is , and denote the offline observations. We write for the marginal law of . The observed-law propensity score is , and the treatment-arm outcome regressions are , . The treatment contrast is . The overlap score used throughout the joint overlap condition is , the distance from the propensity score to the nearest endpoint. Potential-outcome notation is used for the boundedness assumption below; the regret and minimax statements themselves are stated in terms of the observed law , , , and .

Assumption 1 [ass:iid] (Independent sampling).

The observed data units are independent and identically distributed draws from .

⊢ Lean

This is the standard i.i.d. sampling condition for offline policy learning, fixing the experiment in which welfare regret is evaluated (Athey et al., 2021).

Assumption 2 [ass:bounded-outcome] (Bounded potential outcomes).

Under the observed law , the observed outcome is bounded almost surely, and, for every covariate value ,

⊢ Lean

This standard bounded-outcome condition controls the scale of welfare, regret, and inverse-propensity scores; it also implies bounded observed-law conditional means in the treatment arms (Athey et al., 2021).

Assumption 3 [ass:positivity] (Positivity).

The propensity score satisfies -almost surely.

⊢ Lean

This is the standard positivity condition, requiring both treatment arms to occur with positive probability at almost every covariate value relevant under (Imbens et al., 2015).

A deterministic policy is a measurable map . The welfare criterion uses the contrast-weighted value of a policy, so policies are compared only through how they assign treatment on covariate regions with positive or negative .

Definition 1 [def:welfare-regret] (Welfare and Regret ).

For a policy , its welfare under is The oracle policy is The regret of under is

⊢ Lean

Thus is normalized by dropping the treatment-invariant baseline component, is the threshold rule induced by the sign of the treatment contrast, and is the welfare loss from deviating from that rule.

Assumption 4 [ass:policy-class] (Finite-VC policy class).

The policy class is a pointwise measurable class of measurable deterministic policies with finite VC dimension . There exists a countable subclass such that every is the pointwise limit of a sequence of policies from .

⊢ Lean

This is the standard pointwise measurable finite-VC policy-class condition, used to make empirical welfare maximization over statistically and measurably well behaved (Kitagawa et al., 2018).

Assumption 5 [ass:optimal-in-class] (Oracle policy inclusion).

For every law in the law class, the oracle policy belongs to .

⊢ Lean

This is the standard optimum-in-class condition for welfare learning over restricted policy classes, ensuring that the oracle benchmark is attainable within when an upper-bound argument requires such an in-class comparator (Kitagawa et al., 2018). It is kept separate from the observed-law class defined below: the minimax lower-bound class is an observed-law class, whereas oracle inclusion is imposed only where the relevant statement explicitly requires comparison to an element of .

The next assumptions restrict the distribution of small treatment contrasts. The margin condition is the analogue of the low-noise condition in classification: it limits how much covariate mass can lie near the decision boundary .

Assumption 6 [ass:margin] (Margin condition).

The parameters satisfy , , and . For every with ,

⊢ Lean

This is the standard Tsybakov decision margin condition, controlling the probability of covariate values at which the welfare consequences of the treatment decision are small (Tsybakov, 2004).

Assumption 7 [ass:zero-effect] (Zero-effect agreement).

Either or every policy agrees with the oracle policy -almost everywhere on the set .

⊢ Lean

This convention removes welfare-irrelevant disagreements on covariate points where both treatment choices have the same value for the regret target.

Assumption 8 [ass:margin-window] (Margin window).

The margin window satisfies

⊢ Lean

This margin-window normalization fixes the range on which the small-contrast condition is imposed. Since the bounded-outcome condition restricts treatment-arm means to , treatment contrasts lie in , and keeps the window inside that scale.

Definition 2 [def:disagreement] (Disagreement Set ).

For a policy , its disagreement set is

⊢ Lean

The disagreement set records exactly where a policy makes a different treatment assignment from the oracle threshold rule. Regret can therefore be read as a weighted classification error, with weights given by the absolute treatment contrast.

Theorem 1 [thm:welfare-identity] (Welfare regret identity).

Under Assumption 2, for every deterministic policy ,

⊢ Lean

The identity makes the welfare-regret target explicit: mistakes on high-contrast covariate regions are costly, while mistakes near contribute little to regret.

Theorem 2 [thm:margin-localization] (Margin Localization).

Suppose that the following conditions hold:

  • Margin condition. Assumption 6 holds.

  • Zero-effect convention. Assumption 7 holds.

  • Bounded outcomes. Assumption 2 holds.

  • Disagreement sets. is defined as in Definition 2.

Then there exists a constant such that, for every ,

⊢ Lean

The localization result converts small regret into a small disagreement region under the margin condition, a step that parallels the standard excess-risk localization used in margin-based classification analysis (Audibert et al., 2007; Massart et al., 2006; Tsybakov, 2009).

The distinctive restriction in this paper governs the interaction between weak overlap and small contrasts. For , the law class controls the probability that weak overlap and small welfare-relevant contrasts occur together.

Assumption 9 [ass:overlap-decay] (Overlap decay).

For all and with and satisfying the joint small-contrast and weak-overlap probability obeys with the convention that when .

⊢ Lean

This condition is specific to this analysis: it permits overlap to deteriorate near either propensity boundary, but only at a rate tied to the small-contrast region that drives welfare regret. The lower-bound witness itself uses one boundary.

Assumption 10 [ass:strict-overlap-endpoint] (Strict-overlap endpoint).

When , there exists a constant such that holds -almost surely.

⊢ Lean

This is the standard strict-overlap, or uniform-positivity, endpoint condition for the case (Imbens et al., 2015).

Definition 3 [def:exponents] (Exponents , , and ).

The overlap calibration exponent, denominator exponent, and minimax exponent are

⊢ Lean

The term is zero under strict overlap and, for and , positive when weak overlap is allowed to concentrate near small contrasts. The denominator is the exponent balance that will enter the two-point lower-bound construction, and is the corresponding lower-bound regret exponent.

Definition 4 [def:law-class] (Law class ).

Fix the policy class and the constants . The class consists of all well-formed observed laws on for which the following properties hold: -a.s. and for every ; -a.s.; , , , and for every ; either , or every agrees with on up to -null sets; for every and , and, when , and -a.s.

⊢ Lean

Equivalently, is the observed-law class satisfying bounded outcomes, positivity, the margin and zero-effect restrictions, the joint overlap-decay condition, and the strict-overlap endpoint convention when . The sampling condition in Assumption 1, the finite-VC structure in Assumption 4, the margin-window normalization in Assumption 8, and the oracle-in-class condition in Assumption 5 are separate structural conditions invoked by the statements that need them.

Minimax Lower Bound

This section gives the observed-law minimax lower bound over the class in Definition 4. The argument uses a two-point testing construction in the sense of Le Cam (1986): two laws agree away from a small active block, differ only in the sign of a local treatment contrast on that block, and are calibrated so that their product distributions remain statistically close. The regret separation is then read through the welfare identity in Theorem 1. The result characterizes the observed-law converse exponent for this law class.

Definition 5 [def:minimax-regret] (Minimax regret ).

The minimax regret over the law class is where the infimum ranges over all measurable data-dependent estimators taking values in .

⊢ Lean

The risk in Definition 5 is evaluated under the observed law , and the supremum is taken only over laws satisfying the boundedness, positivity, margin, zero-effect, overlap-decay, and endpoint restrictions collected in Definition 4. The estimator is required to output an element of the policy class , while regret remains measured relative to the oracle rule in Definition 1.

The lower bound is driven by the balance between three quantities: the block mass allowed by the margin condition, the local contrast size, and the treatment-arm sampling probability on the weak-overlap cell. The construction below fixes the contrast scale , makes the active block have mass of order , and assigns the weak treatment arm probability of order when overlap decay permits it.

Definition 6 [def:two-point-witness] (Two-point witness laws ).

Let and let be Lebesgue measure restricted to . Set and Set . For each , the law has the same covariate law , the same logging propensity and the same off-block outcome law supported on , with Thus the common off-block optimal label is , and the common off-block contrast is . On , set almost surely when . On the active treated cell , let be supported on with Consequently, The two laws agree everywhere except on the treated active cell. The constants satisfy

⊢ Lean

The construction in Definition 6 is localized in the relevant sense: the two laws differ only in the treated observations on , and the probability of entering that informative cell is . Thus weak overlap reduces information exactly where the sign of the contrast is hardest to learn.

The exponent in Definition 3 is the largest weak-arm exponent within the tight-window algebraic calibration used for a block of margin mass . The next proposition records this calibration.

Proposition 1 [prop:overlap-envelope] (Overlap envelope calibration).

For , , , and , consider the tight window Then and with equality in the displayed power comparison whenever . Moreover , and for every , the corresponding tight window , satisfies Thus, within this algebraic tight-window envelope, is the largest admissible weak-arm exponent.

⊢ Lean

Proposition 1 is the source of the denominator . The term comes from distinguishing Bernoulli means separated by , the term from the block mass allowed by the margin condition, and the term from the probability of observing the informative treatment arm.

The witness laws must also satisfy the restrictions defining . The next statement verifies membership and records the two oracle rules induced by the sign change on .

Lemma 1 [lem:witness-membership] (Witness-law membership).

Assume , , the margin-window normalization in Assumption 8, , , , and . Also assume the witness constant satisfies the overlap-decay smallness conditions Then, for each witness sign , for all sufficiently large , the witness law of Definition 6 satisfies the observed-law membership conditions defining the class in Definition 4. Hence, for all sufficiently large , Moreover, the two corresponding witness-optimal policies are where . If these two policies belong to , then they are available as the two policy actions in the associated two-point lower-bound reduction. The construction uses the off-block contrast , so under the normalization .

⊢ Lean

The proof is deferred to Section E.

Lemma 1 has two roles. First, it places both local alternatives inside the same observed-law class used in the minimax risk. Second, it identifies the two witness-optimal rules. When those two rules lie in , they are the two admissible policy actions in the associated two-action testing reduction; the regret lower bound itself is still stated for the minimax risk in Definition 5, with regret evaluated relative to the oracle rule.

It remains to check that the two laws are hard to distinguish at sample size . Because the laws differ only on a set with covariate mass , treatment probability , and conditional mean separation of order , the one-observation divergence has order .

Lemma 2 [lem:two-point-divergence] (Two-point divergence).

For the witness laws in Definition 6, under Assumption 8, assume in addition that Then there exists a constant such that, for all sufficiently large , Moreover the one-observation chi-square divergence satisfies For the same constant , the product divergence is bounded:

⊢ Lean

The proof is deferred to Section E.

The bandwidth in Definition 6 is chosen so that is bounded. Hence no estimator can reliably determine which sign generated the active block merely from the observed sample.

The same active block creates the welfare separation. Under one law the oracle treats , and under the other law the oracle does not. Any single policy must disagree with at least one of these two assignments on a nontrivial part of the block.

Lemma 3 [lem:regret-separation] (Regret separation).

For the witness laws in Definition 6, suppose Assumption 8 holds, , , , , , every is measurable, and the block-size constant satisfies Then there exists a constant such that, for all sufficiently large , For such , the two optimal policies disagree on the active block. Hence every misclassifies the active block under at least one of the two laws, and the sum of the two regrets is bounded below by the contrast scale times a covariate subblock of mass . Passing from the sum to the maximum gives the stated separation of order .

⊢ Lean

The proof is deferred to Section E.

Lemma 3 converts the testing ambiguity into welfare regret. The separation is per misclassified point on a set of mass , giving .

The testing step uses the standard chi-square form of Le Cam’s two-point method (Le Cam, 1986; Tsybakov, 2009).

Lemma 4 [lem:le-cam-two-point-chisq] (Chi-square testing bound).

Let and be probability laws on a common measurable space, and assume the two product laws satisfy the regularity conditions If, in addition, then every measurable test taking values in satisfies where the floor depends only on the budget , and not on the pair of laws or on the test. In particular, for and , this conclusion holds whenever the one-observation pair obeys the same two regularity conditions, and the one-observation chi-square divergence is bounded by : the two displayed product conditions are then inherited from the one-observation pair, and

The proof is deferred to Section E.

Combining Lemma 2, Lemma 3, and Lemma 4 yields the minimax lower bound. Since , the lower-bound exponent is .

Theorem 3 [thm:minimax-lower] (Minimax lower bound).

Under Assumption 8, suppose that , , , , , and . In addition, when and , assume , and when and , assume . If the policy class is nonempty and every is measurable, then for the minimax regret of Definition 5 over the law class of Definition 4, there is a constant , independent of , such that for all sufficiently large ,

⊢ Lean

Theorem 3 is an observed-law lower bound stated entirely in terms of the law class in Definition 4, the regret functional in Definition 1, and the minimax risk in Definition 5. Theorem 4 develops the complementary procedure-specific achievability analysis.

Writing the exponent of Definition 3 explicitly, . Under strict overlap, , and the lower-bound exponent reduces to the usual margin-driven benchmark. When and , the additional positive term reflects the loss of treatment-arm information in the small-contrast region; at the boundary , it is zero. The section that follows considers a clipped cross-fitted AIPW empirical welfare rule and gives a separate conditional upper bound under explicit nuisance-rate and empirical-process side conditions.

Conditional Analysis of a Clipped Cross-Fitted AIPW Rule

We now turn from the observed-law lower bound to one specified empirical rule. The rule uses a clipped augmented inverse-propensity-weighted score for the treatment contrast and maximizes the resulting empirical welfare criterion over the same pointwise measurable finite-VC policy class used above. This section establishes a conditional rate for the clipped cross-fitted AIPW ERM with supplied nuisance estimates under the nuisance-rate, cross-fitting, boundedness, and localized empirical-process side conditions stated below. These explicit inputs define the procedure-specific achievability domain.

The AIPW score follows the semiparametric and doubly robust construction of Robins et al. (1994); Hahn (1998); Hirano et al. (2003); van der Laan et al. (2011), while cross-fitting is used to separate first-stage nuisance estimation from foldwise policy evaluation as in double/debiased machine learning (Chernozhukov et al., 2018; Chernozhukov et al., 2022). Clipping is needed because the overlap-decay class permits small propensity-boundary distance in the margin region.

Assumption 11 [ass:vc-localized-envelope] (Localized VC envelope).

For the pointwise measurable finite-VC policy class , there exist constants and such that, for every sample size , every , every localized radius , and every policy-compatible policy-indexed increment class satisfying for all and all , and the expected fixed-radius localized centered supremum obeys

⊢ Lean

This is a primitive high-level localized empirical-process condition for the policy-compatible increment classes used below. It is motivated by finite-VC localization results (Bartlett, 2005); the displayed uniform bound enters the analysis as an explicit assumption beyond finite VC dimension.

Assumption 12 [ass:nuisance-rate] (Cross-fitted nuisance rates).

The cross-fitted nuisance estimators satisfy, for each , and their product rate satisfies

⊢ Lean

This is the standard cross-fitted nuisance-rate condition for orthogonal score analysis; the product requirement is the usual second-order remainder control in double machine learning (Chernozhukov et al., 2018).

Assumption 13 [ass:bounded-crossfit-nuisances] (Bounded Cross-Fit Nuisances).

The cross-fitted outcome-regression estimators and take values in .

⊢ Lean

This is the standard bounded nuisance truncation condition; it matches the bounded-outcome scale and keeps the foldwise clipped scores uniformly controlled (Chernozhukov et al., 2018).

Assumption 14 [ass:polynomial-nuisance-exponents] (Polynomial Nuisance Rates).

There exist exponents and , and constants and , such that, for all sufficiently large ,

⊢ Lean

This is the standard polynomial first-stage nuisance-rate condition; the exponents and make explicit which nuisance terms can limit the final regret exponent (Chernozhukov et al., 2018).

Assumption 15 [ass:fixed-crossfit-fold-count] (Fixed Cross-Fit Fold Count).

The number of cross-fitting folds is a fixed finite integer independent of , and the deterministic partition of the observation indices is balanced.

⊢ Lean

This is the fixed- cross-fitting convention used by the conditional upper bound. The balanced deterministic folds are a side condition for the foldwise pooling step in the present empirical-process argument; the general use of cross-fitting follows the double/debiased machine-learning approach (Chernozhukov et al., 2018).

Assumption 16 [ass:vc-localized-offset-envelope] (Localized Offset Envelope).

For the pointwise measurable finite-VC policy class , there exist constants and such that, for every integer , every , and every policy-compatible increment class (that is, factors through and the decision ), the centered empirical process built from an i.i.d. sample from , with envelope for all and , and second moment bounded by satisfies

⊢ Lean

This is a second primitive high-level condition, now for an offset empirical process. It converts stochastic fluctuation into an expected regret contribution at the margin-dependent exponent; the exact statement enters as an explicit assumption beyond finite VC dimension (Bartlett, 2005).

The tuning schedule balances four terms: the clipped empirical-process envelope, the product nuisance remainder after clipping, the localized weak-overlap bias, and the outcome-regression error on the complement of the localization window. The next definition records this balance for a fixed nuisance regime.

Definition 7 [def:feasible-rate] (Feasible upper exponent ).

Fix one nuisance regime , and set

If , fix constants For and , define Let be any maximizer of over this compact feasible set, and set Then for all sufficiently large , and the solved conditional feasible upper exponent is

If , take a fixed clipping sequence with , and set

⊢ Lean

For , the constraint is the exponent form of the admissibility condition in the overlap-decay region. The exponent in Definition 7 is therefore the conditional exponent delivered by this balance, truncated by the lower-bound exponent from Definition 3.

Definition 8 [def:clipped-propensity] (Clipped propensity ).

For a nuisance triple and a clipping level , the clipped propensity score is

⊢ Lean

The nuisance triple collects the two outcome regressions and the propensity score used in the plug-in score. Clipping replaces the propensity component by a value in , controlling inverse-propensity weights while preserving the target welfare functional in Definition 1.

Definition 9 [def:clipped-aipw-score] (Clipped AIPW score ).

For , nuisance triple , and clipping level , the clipped augmented inverse-propensity-weighted score for the treatment contrast is

⊢ Lean

When the nuisance functions are equal to the true observed-law regressions and propensity and no clipping binds, is the usual AIPW score for the contrast . Under overlap decay, clipping introduces a deterministic drift that requires explicit control beyond the strict-overlap simplification.

Definition 10 [def:feasible-erm] (Feasible clipped-AIPW ERM ).

Fix the countable pointwise dense subclass of , and fix a deterministic balanced -fold partition of , with each fold having size or . Let denote the evaluation fold containing observation . Given foldwise cross-fitted nuisance estimators and a clipping level , define the empirical welfare criterion Let be the smallest index such that and set This measurable -valued -empirical risk maximizer over is the feasible clipped-AIPW ERM. The clipping level is the tuning parameter chosen by the bias-variance balance.

⊢ Lean

Definition 10 fixes the enumeration, the balanced fold partition, and the foldwise supplied nuisance estimates. The superscript indicates that the nuisance estimate used on fold is trained away from that fold; its construction is otherwise abstracted into the side conditions above.

Definition 11 [def:upper-risk] (Conditional upper risk ).

Fix a nuisance regime , the associated deterministic schedules , and the supplied policy class, enumeration, nuisance estimators, and fold assignment. The conditional upper risk is where the supremum is over all laws for which the oracle policy belongs to , the supplied cross-fitted nuisance estimators satisfy the stated nuisance-rate bounds, boundedness conditions, and polynomial nuisance-rate bounds with exponents and constants , the supplied policy class satisfies the stated finite-VC, localized-envelope, and localized-offset-envelope conditions, the supplied cross-fitting scheme has fixed fold count , and the supplied enumeration is a countable pointwise-dense skeleton of .

⊢ Lean

The risk in Definition 11 is evaluated over the stated side-condition domain: laws in , oracle-in-class policy problems, supplied cross-fitted nuisances with the required rates and boundedness, policy classes satisfying the localized empirical-process conditions, and fixed balanced cross-fitting schemes. The definition precisely describes the conditional risk controlled by this clipped-AIPW ERM.

The basic inequality is the usual consequence of near-maximizing the empirical welfare criterion over the dense countable subclass.

Lemma 5 [lem:feasible-erm-basic-inequality] (Feasible ERM Inequality).

Under Assumption 4, Definition 10 defines a measurable -valued estimator satisfying Hence, for every comparator , Under Assumption 5, this conclusion applies with .

⊢ Lean

The proof is deferred to Section E.

The comparator is available only when Assumption 5 is imposed. This is why the upper-bound analysis carries oracle inclusion as an explicit side condition even though the lower bound in Theorem 3 is stated as an observed-law result over .

Lemma 6 [lem:crude-clipped-score-envelope] (Clipped-Score Envelope).

Under the bounded-outcome condition of Assumption 2, for every and every cross-fitted nuisance triple whose outcome-regression components satisfy for every , there exists a constant such that the clipped AIPW score of Definition 9 satisfies Consequently, with the same constant,

⊢ Lean

The proof is deferred to Section E.

Lemma 6 identifies the stochastic price of clipping at level . Smaller reduces clipping bias but increases the empirical-process envelope, producing the term in the master bound below.

Lemma 7 [lem:localized-clipped-drift-bound] (Localized Clipped Drift Bound).

Under Assumption 12, Assumption 14, Assumption 9, Assumption 7, Assumption 2, and Assumption 10, and Lemma 13, let be a well-formed observed law, let every be measurable, and fix a clipping level with . With the cross-fitted nuisance estimators fixed on the evaluation fold, write , and suppose , and that these three nuisance-error functions belong to . Define

If , set . Then, for every and every localization window satisfying , writing , whenever .

If and , set . Then, for every ,

⊢ Lean

The proof is deferred to Section E.

The drift is the conditional mean bias of the clipped and estimated score relative to the true contrast. In the weak-overlap case, Lemma 7 separates the drift into a clipped product remainder, a small-contrast weak-overlap term, and a localization-complement term; under strict overlap, only the product nuisance remainder remains.

Lemma 8 [lem:crude-localized-master-bound] (Crude Localized Master Bound).

Let and be deterministic clipping and contrast-window schedules, and let be the cross-fitted clipped-AIPW -ERM of Definition 10, computed with clip over the enumeration .

Assume:

  • (Policy class, skeleton, and oracle.) The policy class satisfies Assumption 4; welfare and regret are as in Definition 1. The enumeration used by the ERM is the pointwise-dense skeleton of : every , and every is the pointwise limit of a subsequence from . For each law to which the bound is applied, the oracle satisfies Assumption 5.

  • (Sampling.) For each law to which the bound is applied, the observations are i.i.d. as in Assumption 1.

  • (Nuisance rates.) The polynomial nuisance exponent condition in Assumption 14 holds. For each law to which the bound is applied, each foldwise nuisance triple satisfies Assumption 12. In addition, and for all sufficiently large .

  • (Cross-fitting and boundedness.) The number of folds is fixed and the fold assignment is balanced as in Assumption 15; for each law to which the bound is applied, the foldwise cross-fitted outcome nuisances satisfy Assumption 13.

  • (Clipping schedule.) For all sufficiently large , . If , then there is a fixed such that for all sufficiently large .

  • (Measurability and regularity.) For every and fold , the nuisance functions , , and are measurable. For all sufficiently large , every , and every fold , the errors , , and belong to .

  • (Localized empirical processes.) The uniform class-level localized envelope and offset envelope conditions in Assumption 11 and Assumption 16 hold.

Then there exist constants , with and , such that for all sufficiently large and every observed law , if and the law-specific oracle, i.i.d. sampling, foldwise nuisance-rate, and bounded cross-fit nuisance conditions above hold, then the following two implications hold.

If , , and then

If and , then

⊢ Lean

The proof is deferred to Section E.

Lemma 8 is the organizing inequality for the upper-bound calculation. Its assumptions collect the policy, sampling, nuisance, cross-fitting, boundedness, empirical-process, and drift inputs needed by the clipped-AIPW rule.

For reference, Table 2 maps the terms in the master bound to the exponent components optimized in Definition 7, after substituting , , , and .

Terms entering the conditional feasible exponent when
Master-bound term Source Exponent component
clipped empirical-process envelope
clipped product nuisance remainder
localized weak-overlap drift
localization-complement outcome error

The balance in Definition 7 chooses the clipping and localization exponents to maximize the minimum of the four entries in the last column. The next lemma records the resulting algebraic rate.

Lemma 9 [lem:clip-balance-exponent] (Clipping Balance Exponent).

Under the polynomial nuisance-rate exponents of Assumption 14, assume also that the nuisance constants satisfy and , that for all sufficiently large , and that . If , assume in addition that and that the deterministic schedules below are eventually admissible, . Then the deterministic clipping and localization schedules satisfy the following exponent balance.

For , let and There exist constants and such that, for all sufficiently large , where Thus the optimized upper-bound exponent in the positive-overlap-decay regime is

For , with the fixed clipping schedule , the same constants may be chosen so that, for all sufficiently large ,

⊢ Lean

The proof is deferred to Section E.

Lemma 9 shows exactly when this clipped-AIPW rule reaches the lower-bound exponent from Theorem 3: this happens in the nonbinding branch for . In the binding branch, the result gives the conditional upper exponent for the specified rule.

We will refer to the side-condition domain of Definition 11 as the collection of laws, policy classes, nuisance estimates, and fold schemes satisfying the assumptions explicitly listed in that definition and in the theorem below. This terminology abbreviates exactly those displayed hypotheses.

Theorem 4 [oeq:feasible-upper] (Feasible Upper Exponent).

Fix parameters with and a nuisance regime satisfying Assumption 14, with and . Consider the law class of Definition 4, the feasible cross-fitted clipped-AIPW -ERM of Definition 10, formed using an enumeration of a pointwise-dense skeleton of , and the conditional upper risk of Definition 11. Let the feasible clipping and localization constants lie in the input domain of Definition 7: if , then and if , then

Assume:

  • (Policy class and skeleton.) The policy class satisfies Assumption 4, and the enumeration used by is a pointwise-dense skeleton of .

  • (Cross-fitting.) The number of folds satisfies Assumption 15.

  • (Nuisance boundedness.) The cross-fitted outcome nuisance estimators satisfy Assumption 13.

  • (Nuisance measurability and integrability.) For every and fold , , , and are measurable. The rate sequences satisfy and for all sufficiently large . For all sufficiently large , every , and every fold , the errors , , and belong to .

  • (Localized envelopes.) The uniform localized empirical-process and offset envelope conditions in Assumption 11 and Assumption 16 hold.

Then there exist constants with and such that, for all sufficiently large , the conditional upper risk of , uniformly over the side-condition domain in Definition 11, satisfies where is the feasible upper exponent supplied by Definition 7 for the above and .

⊢ Lean

Theorem 4 is the promised conditional achievability statement for the plain clipped cross-fitted AIPW empirical welfare rule. Its uniformity is over the domain specified in Definition 11 and the theorem’s displayed hypotheses. Assumption 11–Assumption 16 are explicit inputs to this domain. The discussion section compares the strict-gap branch with the observed-law converse.

Discussion, Limitations, and Open Question

The preceding sections separate two complementary objects. Theorem 3 gives an observed-law minimax lower bound over the class in Definition 4. Theorem 4 gives a conditional upper bound for the clipped cross-fitted AIPW empirical welfare rule over the side-condition domain in Definition 11. Thus the exponent is a converse benchmark for the law class, whereas is the exponent delivered by one feasible procedure under explicit nuisance-rate, cross-fitting, boundedness, and localized empirical-process conditions.

This distinction affects the displayed exponent in regimes where the nuisance and clipping balance binds. In the nonbinding branch of Definition 7, the optimized nuisance exponent is at least the lower-bound exponent, so the conditional upper bound has the same polynomial exponent as the observed-law lower bound, up to logarithmic factors. In the strict-gap branch, the feasible exponent is smaller because one of the terms in the master bound of Lemma 8 controls the rate. The two conclusions are consistent: the lower bound ranges over all measurable policy estimators, while the upper bound analyzes a particular clipped AIPW ERM with estimated cross-fitted nuisance functions.

A statistical source represented in the bound is the weak-overlap mechanism emphasized in the limited-overlap literature (Li et al., 2016; D’Amour et al., 2017; Ben-Michael et al., 2022; Hill et al., 2024; Susmann et al., 2025). The lower-bound alternatives make the informative treatment arm rare precisely on the small-contrast block that determines regret. The observed-law minimax lower bound shows that the law class has a testing difficulty encoded in . A feasible clipped-score procedure must also estimate the nuisance functions and choose a clipping level that controls inverse-propensity variability. When nuisance learning is slow in the weak arm, the product remainder and localized drift terms can dominate the purely testing-driven benchmark. This connects to recent weak-overlap and clipping analyses of how overlap, nuisance estimation, and localization interact (Zhao et al., 2023; Liu et al., 2026).

Definition 12 [oeq:feasible-tight] (Feasible Tightness Question).

In the strict-gap branch of the conditional feasible achievability bound for the regime-indexed risk , it is open whether a genuinely feasible estimator using estimated cross-fitted nuisance functions can attain the converse exponent , or whether weak-arm nuisance learning imposes the slower conditional exponent derived by the feasible upper bound.

⊢ Lean

Definition 12 is intentionally stated as an open question. It asks whether the gap in the strict-gap branch is an artifact of the plain clipped AIPW ERM and its analysis, or whether it reflects an intrinsic cost of learning the nuisance functions in the same weak-overlap region that drives the lower-bound construction. Resolving this would require either a feasible procedure whose risk reaches under comparable nuisance-side conditions, or a converse argument that incorporates nuisance learning into the lower-bound experiment.

Theorem 4 controls for the specified clipped cross-fitted rule and the stated side-condition domain. The broader feasible minimax problem also includes other estimators, adaptive clipping schemes, and procedures that exploit additional structure in the nuisance functions. Theorem 3 supplies the observed-law converse free of nuisance-estimation restrictions. The strict-gap branch therefore poses a genuine feasible minimax question beyond the completed observed-law lower-bound calculation.

Appendices

Algebra, Testing, and Witness-Law Details

This appendix collects the algebraic details behind the lower-bound construction in Theorem 3. The purpose is only to connect the calibration, membership, divergence, separation, and testing ingredients used by Theorem 3. Throughout, the rate statement is a lower-bound statement over the observed-law class in Definition 4.

The exponent calculation begins with the overlap envelope in Proposition 1. For a local contrast scale , an active block of mass proportional to , and weak-arm probability proportional to , the overlap-decay condition can be tested on the tight admissible window On that window, Since , the envelope is at least of order precisely when This is the calibration recorded in Proposition 1. Within the tight-window algebra used by the witness construction, it identifies the weakest admissible treatment-arm probability on the active block, and therefore explains why the denominator in Definition 3 contains the additional term when .

The witness laws in Definition 6 are designed so that this calibration is binding. The covariate law is Lebesgue measure on , the active block has mass , and the logging propensity on that block is , equal to in the strict-overlap endpoint and to in the weak-overlap case. Off the block, both laws have the same strictly positive contrast . On the block, they differ only by the sign of the treated-arm conditional mean, so that the two oracle rules are and , as stated in Lemma 1. The membership lemma is also where the availability of these two rules in the policy class is recorded for the two-action reduction.

The chi-square calculation in Lemma 2 has only one contributing cell. The laws coincide outside , and they also coincide on the untreated part of . Conditional on and , the two outcome distributions are Bernoulli laws on with success probabilities and . Their one-cell chi-square divergence is of order . Multiplying by the probability of entering the cell gives With the product divergence remains bounded because Equivalently, the product formula is bounded by a constant, as in Lemma 2.

The regret calculation in Lemma 3 uses the fact that the two oracle assignments disagree exactly on . Any single deterministic policy must assign at least half of the block incorrectly for one of the two laws, up to the constant convention in the lemma. On that law, the welfare identity in Theorem 1 gives a pointwise regret weight on the misclassified part of . Thus the loss is bounded below by a constant multiple of Substituting the calibrated bandwidth gives

The final step is the two-point testing reduction. Lemma 4 is the chi-square version of Le Cam’s method (Le Cam, 1986; Tsybakov, 2009): bounded product divergence prevents every test from identifying the sign with vanishing total error. Applying that testing bound to the two witness laws converts the regret separation into the minimax lower bound in Theorem 3. Its exponent characterizes the converse construction, while feasible policy-learning upper rates are developed through a separate achievability analysis.

Empirical-Process and Cross-Fitting Details

This appendix collects the empirical-process inputs used by the conditional feasible upper bound in Theorem 4. The statements translate the localized offset-envelope side condition in Assumption 16 into the cross-fitted and self-bounding forms used by Lemma 8; the localized VC envelope condition in Assumption 11 enters the conditional domain directly. Fixed , balanced folds, and the displayed empirical-process envelopes thereby define the conditional feasible-analysis domain.

The first auxiliary statement is the self-bounding step that converts an offset empirical-process control into an expected regret bound. It is the margin-localized analogue of the standard offset argument in excess-risk analysis for VC classes (Audibert et al., 2007; Massart et al., 2006; Tsybakov, 2009).

Lemma 10 [lem:localized-vc-self-bound] (Localized Self Bound).

Let and let be the constant of the cross-fit localized offset bound, so that the offset positive part of the centered -indexed process is controlled at scale . Assume the sample size is large enough that Then every -indexed estimator satisfying also satisfies If, in addition, , then—using —the term is absorbed into the leading rate and

⊢ Lean
Proof of Lemma 10.

Write for the offset scale supplied to the lemma, so that the hypothesis on the centered -indexed process reads and the large- hypothesis reads . In particular .

Step 1: a samplewise bound. Fix a realization of the sample. Since , the value of the offset functional at is one of the terms in the supremum defining , so Put , , and , so that . Substituting this into the assumed selection inequality gives If , then and multiplying by gives . If , then and , so . In both cases, using the first display of this step,

Step 2: integrate. Both sides are integrable by the stated regularity, so integrating the samplewise bound and using gives Since , we may enlarge the factor to , obtaining which is the asserted bound.

Step 3: absorbing the slack. If in addition , then the assumed large- comparison gives , and the bound of Step 2 becomes so the term is absorbed into the localized rate.

Lemma 10 is the device that absorbs the stochastic term appearing after the ERM basic inequality. The slack in Definition 10 is negligible relative to the localized rate , so the near-maximization error preserves the polynomial exponent.

The final empirical-process input is the offset version needed by the self-bound. It is the cross-fitted counterpart of the localized offset-envelope condition in Assumption 16.

Lemma 11 [lem:crossfit-localized-offset-control] (Cross-Fit Offset Control).

Let be the pooled cross-fit centered process formed from foldwise increments , indexed by .

Assume:

  • (Data and folds.) The observations are i.i.d. as in Assumption 1, the observed law is well formed, and the fixed balanced cross-fitting scheme of Assumption 15 is used.

  • (Policies and localization.) The policy class satisfies Assumption 4, the margin and zero-effect conditions of Assumption 6 and Assumption 7 hold, and is the disagreement set in Definition 2.

  • (Localized offset envelope.) The uniform class-level localized offset-envelope condition of Assumption 16 holds for and exponent .

  • (Outcome.) The bounded-outcome condition of Assumption 2 holds.

  • (Foldwise increments.) For every sample size , every , and every foldwise increment family , each fold map is policy-compatible, each is measurable for , the pooled offset supremum and the corresponding foldwise offset suprema are integrable, and for every fold , every , and every observation .

Then there exist constants and such that

⊢ Lean
Proof of Lemma 11.

Write for the size of fold , for its weight, and The pooled cross-fit centered process is For every fold, use the totalized foldwise centered process With this convention, empty folds are included in fold sums and have zero weight. Grouping the pooled sum by fold gives the exact decomposition the weight identity because partitions .

Apply Assumption 16 in its uniform class-level form: it supplies constants and , not depending on the law, and we set which is positive because Assumption 15 gives . Also, Assumption 6 gives , hence

Step 1: a pointwise foldwise bound on the offset supremum. Fix a sample and a policy , and write . By the triangle inequality and , the second step because and . The right-hand side is a nonnegative sum, so it also dominates , and therefore dominates the positive part on the left: The regret is bounded below by a fixed constant on all of : the setup identity and Assumption 2 give for every , every is measurable by Assumption 4, and the oracle is measurable because is; hence so . Each foldwise offset term is therefore bounded above uniformly over : the envelope hypothesis gives , with the totalized definition giving the same bound when , and . Thus each foldwise supremum is finite, and taking the supremum over in the previous display yields the pointwise-in-sample bound

Step 2: integrate and reduce each nonempty fold to an i.i.d. sample of its own size. The stated regularity hypotheses give integrability of the pooled offset supremum and of the foldwise offset suprema on their fold product spaces. The coordinate projection onto fold is measure preserving, so the corresponding full-sample foldwise supremum is integrable. Integrating the last display over the product law and exchanging the finite sum with the integral gives If , then and that fold contributes nothing. If , the same measure-preserving projection identifies with the offset-supremum expectation for an i.i.d. sample of size . For such a fold, is policy-compatible with envelope and second moment so Assumption 16 applies at sample size and gives

Step 3: compare each weighted fold rate with the target -rate. Put . Then , so because and . Moreover and give

Combining the last three displays, every fold contributes at most

Conclusion. Summing over the folds, which is the asserted bound.

Lemma 11 supplies the offset positive-part bound required by Lemma 10. In the feasible clipped-AIPW analysis, the envelope is later instantiated at the clipped-score scale from Lemma 6; the resulting term is the empirical-process component in Lemma 8. Together, Lemma 10–Lemma 11 provide only the empirical-process part of the conditional upper bound; the drift and clipping bias are handled separately in Appendix C.

Clipped-Score Drift and Bias Localization

This appendix records the deterministic bias calculations used by the clipped-score analysis in Theorem 4. The key point is that clipping a doubly robust score changes its conditional mean unless the clipped plug-in propensity agrees with the observed-law propensity, or the outcome-regression errors vanish at the same covariate value. Thus the usual orthogonality logic for AIPW scores (Robins et al., 1994; Hahn, 1998; Hirano et al., 2003; Chernozhukov et al., 2018) must be combined with an explicit localization argument near the clipped propensity region.

Lemma 12 [lem:clip-bias] (Clipped Score Drift).

Under Assumption 2, Assumption 3, and Assumption 13, and with the clipped propensity and clipped AIPW score of Definition 8 and Definition 9, let be any measurable plug-in nuisance triple. Define Then the clipped score reproduces the contrast up to the drift in the tested (population) sense: for every bounded measurable , equivalently for -almost every . In particular, the drift vanishes whenever or , since either makes a factor of vanish; these conditions are sufficient but not necessary, as the bracketed term may itself vanish. Cancellation does not, however, follow merely from : there exist laws and covariate values with at which the drift is nonzero.

⊢ Lean
Proof of Lemma 12.

Fix a bounded measurable covariate test function , and write for the clipped plug-in propensity of Definition 8, whose clipping level is fixed in . Since , we have , and the definition therefore gives the two-sided bound so both denominators are bounded away from zero: and . Together with Assumption 2 (which gives -a.s. and ) and the fixed plug-in regularity for this drift calculation (measurability of , measurability of , and uniform boundedness of ), every integrand appearing below is measurable and bounded, hence - and -integrable; this is the only role of the boundedness hypotheses.

Step 1: expand the score and pass to covariate integrals. Multiplying the score of Definition 9 by and splitting it into five terms,

Each of the five terms is integrable, so the integral of the sum splits into the sum of the integrals. The defining conditional-mean relations of the observed law, namely and , state that for every bounded measurable ,

Applying these with , , , and – all bounded and measurable by the previous paragraph – converts the five terms into -integrals:

Step 2: collect the integrand pointwise. Fix . Using and , the two treated terms combine as and the two control terms combine as Both rearrangements are legitimate because and are nonzero. Adding them to gives the pointwise identity with exactly the function displayed in the statement.

Step 3: integrate. Both and are bounded measurable, hence -integrable, so the integral of the right-hand side splits and Step 1 becomes Since this holds for every bounded measurable , it is precisely the asserted conditional-mean drift identity for -almost every .

Step 4: the cancellation mechanisms. The closed form is a product of two factors. If , the first factor is ; if , the second factor is . In either case .

Step 5: no cancellation from alone. It remains to show that the stated assumptions do not permit replacing this algebra by a cancellation claim on the region . Take the covariate space , and let be the law that puts covariate mass at , assigns treatment by a fair coin, and returns always; explicitly, is the two-atom law This is a well-formed observed law with bounded outcomes and satisfies positivity, since . Take and the plug-in triple , , , and evaluate at . The overlap score is so lies strictly inside the region where clipping is claimed not to matter. Yet the clipped plug-in propensity is , while and , so the drift formula gives

Hence alone does not force the drift to vanish, and no additive clipped-region cancellation bound follows from the stated boundedness, positivity, and observed-law conditions alone.

Lemma 12 is the exact drift identity behind the bias term in Lemma 7. It separates the propensity discrepancy from the regression errors and shows that correct handling of clipping requires the full product identity. In particular, stronger cancellation requires assumptions beyond the stated boundedness and positivity conditions.

The remaining issue is to control where clipping can matter for regret. Under the joint overlap-decay condition, the relevant clipped region becomes small after intersection with the welfare-relevant disagreement and small-contrast regions. The next lemma gives that localization.

Lemma 13 [lem:clipped-region-localization] (Clipped Region Localization).

Suppose that the following conditions hold:

  • Positive weak-overlap exponent. .

  • Well-formed law. The observed law is well formed, with a probability covariate marginal, , measurable nuisance functionals, and .

  • Overlap decay. Assumption 9 holds with constants .

  • Zero-effect convention. Assumption 7 holds for the policy class .

  • Bounded outcomes. Assumption 2 holds.

  • Measurable policies. Every is measurable.

  • Disagreement sets. is defined as in Definition 2.

Then , and every policy with regret satisfies whenever and .

⊢ Lean
Proof of Lemma 13.

Since is a real constant, , which is the first assertion.

Fix a policy , let and , and fix with , , and . Put the set whose -mass is to be bounded. We cover by the three sets and Indeed, let . If , then, since means , we get . If and , then, since , we get . Otherwise , and since , we get . Thus All three sets are measurable, because , , and are measurable.

The zero-contrast piece is null. Assumption 7 gives In its first alternative , so its subset is null by monotonicity. In its second alternative, applied to the present policy , it says exactly that agrees with -almost everywhere on , that is, .

The small-contrast weak-overlap piece. Assumption 9 applies with the present and , because , , and . Since , the convention is not in force and the assumption reads the last step because and .

The large-contrast piece. Here we use the welfare regret identity of Theorem 1, whose integrand is nonnegative and, on , at least (since and on ). Restricting the integral to and then bounding the integrand below by the constant gives the middle inequality because the integrand is nonnegative. The integral is finite because Assumption 2 gives . Dividing by ,

Conclusion. By monotonicity of along , subadditivity, and , which is the asserted bound.

Lemma 13 is the measure bound used to localize the drift from Lemma 12. The first term is the overlap-decay contribution on the small-contrast window , while the second term is the regret-controlled contribution from disagreements outside that window. The admissibility condition is the same one used in the feasible tuning schedule of Definition 7.

Together, Lemma 12 and Lemma 13 supply the deterministic bias inputs for Lemma 7. They bound the clipped-score drift by combining nuisance-rate control, overlap-decay localization, and the regret radius of the policy under consideration.

Verification Note

This appendix records the scope of the formal layer underlying the statements in the paper. The object of analysis is the observed-law offline policy-learning experiment introduced in Assumption 1: one observation has law on , and the regret target is the contrast-weighted welfare loss in Definition 1. The potential-outcome notation used in Assumption 2 fixes the outcome scale, but the lower-bound and upper-bound statements themselves are expressed through the observed law, the propensity score, the outcome regressions, the contrast, and the induced oracle policy.

The observed-law side of the analysis consists of the welfare identity in Theorem 1, the disagreement notation in Definition 2, the margin localization statement in Theorem 2, and the law class in Definition 4. The exponents used in the lower bound are those in Definition 3; in particular, the denominator and the lower-bound exponent are fixed there before the minimax statement is formed.

The lower-bound argument is represented by the two-point witness construction in Definition 6, the overlap-envelope calibration in Proposition 1, witness membership in Lemma 1, the product-divergence control in Lemma 2, the regret separation in Lemma 3, and the chi-square testing reduction in Lemma 4. These statements support the minimax lower bound in Theorem 3; the conditional upper analysis is represented by the separate clipped-AIPW chain below.

The feasible rule is represented separately. The clipping schedule and feasible exponent are defined in Definition 7; the clipped propensity and clipped AIPW score are defined in Definition 8 and Definition 9; and the cross-fitted empirical welfare rule is defined in Definition 10. The basic ERM inequality in Lemma 5, the clipped-score envelope in Lemma 6, the clipped-bias identity in Lemma 12, the drift localization in Lemma 7, and the master bound in Lemma 8 are the components used for the conditional upper bound in Theorem 4.

The upper-bound result has an explicit side-condition domain in addition to the observed-law class. This domain includes the nuisance-rate assumptions in Assumption 12 and Assumption 14, the bounded cross-fitted nuisance condition in Assumption 13, the fixed balanced fold condition in Assumption 15, and the localized empirical-process and offset-envelope conditions in Assumption 11 and Assumption 16. The process reductions in Lemma 10 and Lemma 11 are invoked within that conditional domain.

Accordingly, Theorem 4 is a procedure-specific conditional achievability statement for the clipped cross-fitted AIPW ERM under Definition 11. Definition 12 records the open strict-gap question: whether another genuinely feasible estimator can attain the converse exponent , or whether the slower nuisance-limited exponent is intrinsic.

The theorem statements and proofs described above are machine-checked in Lean 4. The machine-checked scope covers the observed-law objects, exponent definitions, witness construction, lower-bound skeleton, clipped-score identities, ERM inequality, and conditional upper-bound dependencies. Identification primitives, nuisance-rate conditions, fixed-fold cross-fitting, localized VC envelope bounds, and offset-envelope controls enter as stated assumptions defining the conditional domain.

Proofs of the main results

Proof of Theorem 1.

Fix a well-formed observed law satisfying Assumption 2, and fix . By Assumption 4, is measurable. Well-formedness gives that is a probability measure, that is measurable, and that for every . The regression part of Assumption 2 gives for every , and therefore

The set is measurable because is measurable, and is measurable because is measurable. Hence the functions are measurable. Since is a probability measure and , both functions are integrable with respect to , so the difference of the two welfare integrals may be written as one -integral.

Writing Definition 1 in its covariate-marginal form, with and , the preceding integrability gives

Fix . If , then . When , the two terms in the bracket are both , so the bracket and the disagreement indicator are both zero. When , the bracket is , and the disagreement indicator is one. If , then . When , the bracket is , and the disagreement indicator is one. When , the two terms in the bracket are both zero, and the disagreement indicator is zero. Therefore, for every ,

Substituting this pointwise identity into the preceding integral yields This is the -weighted disagreement representation of regret with respect to the covariate marginal.

Proof of Theorem 2.
  1. The constant. Fix Assumption 6 gives , , and , so every summand is nonnegative and ; moreover This constant depends only on , , and . It remains to prove the displayed inequality for an arbitrary well-formed observed law , policy class whose members are measurable, and every .

  2. Two elementary facts. Let as in Definition 2. By Theorem 1, The integrand is nonnegative, so . The standing well-formed-law conditions give , and Assumption 2 gives and ; hence , which supplies the integrability used in the identity. Also , since well-formedness makes a probability measure.

  3. The basic decomposition. For every , Indeed, write Every lies in one of these three sets, according to whether , , or ; hence Assumption 7 gives : in its first alternative , so its subset is null; in its second alternative, applied to , the set of zero-contrast disagreements is null by hypothesis. Assumption 6, applied at the admissible window , gives Finally, on the integrand in the regret identity is at least , and the integrand is nonnegative, so restricting the integral to gives Monotonicity and subadditivity of along the displayed inclusion now give the asserted decomposition. The required measurability comes from the well-formed-law measurability of together with the policy-measurability part of Assumption 4; boundedness and finiteness are supplied by the well-formed-law identity , Assumption 2, and the probability property of .

  4. The case . Then and , while gives . Hence

  5. The case , , small window. Set and suppose first . Then the two exponent identities hold, so Step 3 gives the last step by and .

  6. The case , , large window. Suppose instead . Since , raising to the power is increasing on nonnegative arguments, so Multiplying by and using , Combining with from Step 2 gives the desired bound in this subcase.

  7. The case , . Let , and put Then , and monotonicity of positive powers gives Step 3 with therefore gives Since was arbitrary and , . Also , so and therefore

Proof of Proposition 1.

Throughout, the tight window is and for a real exponent , so that the envelope value in question is For any real , the power laws (legitimate since and ) give

Since , the map is strictly order-reversing, so As and , multiplying by turns the right-hand inequality into with as in Definition 3. Applying the two displays with gives the asserted identity and the asserted equivalence

If , the exponent evaluates exactly: which is the equality case in the power comparison. Finally, because , , and ; and since above was an arbitrary real, the same equivalence holds for every . Thus the largest, and therefore least informative, weak-arm exponent allowed within this tight-window overlap-envelope calibration is precisely .

Proof of Lemma 1.

Fix a witness sign . Write , , , and . Since by Assumption 8, we have . After discarding finitely many , the construction of Definition 6 has , , and . By construction the contrast, propensity, and overlap score of are the last because makes . The arm regressions are and on , and , off . We check the defining conditions of Definition 4 in turn.

Well-formedness and bounded outcomes. The construction specifies the covariate law, propensity, outcome kernels, and nuisance functions explicitly; these fields are measurable, the contrast is the difference of the two regressions, and the tested propensity and regression identities agree with the kernels, so is a well-formed observed law. Its outcome is supported on or on , so almost surely; and the arm regressions take only the values , , and , which lie in because and .

Positivity. From the display above, , and , so for every .

Zero-effect. The contrast takes only the values and , so and in particular .

Margin. Fix and put . Off the contrast has magnitude , so , and since is Lebesgue measure restricted to , If , then on the contrast has magnitude , so and the bound is trivial. If , then because , and, using ,

Strict-overlap endpoint. Suppose . Then , so Definition 6 gives , and the overlap score takes only the values and . Since , we get for every .

Overlap decay. Fix with , , and , and let As in the margin check, off-block points have contrast magnitude , so Moreover, on the contrast magnitude is . Hence if , then and the required bound holds because its right-hand side is nonnegative. It remains to consider the case , and we split first on .

If , the required bound is by the convention factor in Definition 4. From , , and ,

Now suppose . On the overlap score is , so if , then and the required bound again follows from nonnegativity of the right-hand side. Assume .

If , then , so , and reads . Raising to the nonnegative power gives , so with the assumed and ,

If , then with , and the calibration identity is Using in both factors, and then in the second factor, Multiplying by and using the assumed ,

Thus the overlap-decay condition holds in every case.

Conclusion. All the defining conditions of Definition 4 hold for the fixed sign . Applying the preceding argument for and , and intersecting the two eventual ranges of , gives for all sufficiently large . For the oracle rules: under the contrast equals on and off , so for every . Under the contrast equals on and off , so exactly when , that is . If these two maps belong to , the displayed oracle identities identify the two policy actions used in the associated two-point reduction. Finally, the normalization is exactly what permits the displayed choice , which is both strictly above the margin window and compatible with the outcome bound .

Proof of Lemma 2.

Take and fix in the eventual range where and . Write , , , and .

Step 1: the one-observation divergence. By Definition 6, the two laws have the same covariate law, the same propensity, and the same outcome law except on the treated active cell. Let These three measurable cells partition the full observation space. On , the restrictions of and agree: this includes the common off-cell part of the law and the residual part of the active treated cell outside the two charged outcome atoms, which has zero mass under both laws. On the two charged cells the outcome probabilities in Definition 6 give constant restriction ratios so , and the squared likelihood-ratio deviation is -integrable.

Set Under , the two charged cells have masses The finite-partition chi-square formula for these restriction ratios yields Since , we have , and hence

Because is Lebesgue measure restricted to and , its block mass obeys . Thus The weak-arm level satisfies : if , then , while otherwise , as specified in Definition 6. Therefore which is the asserted one-observation bound.

Step 2: the product experiment. The one-observation absolute continuity and square-integrability just obtained tensorize, so and

By Definition 3, , and Definition 6 gives . Hence The chi-square divergence tensorizes as

Using for and the assumption , Therefore , and so

Proof of Lemma 3.

Take the separation constant and work at all sufficiently large for which Lemma 1 applies to both signs, so that and their oracle rules are

Write , , and shrink the active block to Since and , we have and hence ; as is Lebesgue measure restricted to , the covariate mass of is exactly

Moreover , so

Fix any policy . Let and be its disagreement sets under and , and set Under each of the two witness laws, which are well formed with bounded outcomes, the welfare identity of Theorem 1 writes the regret as with a nonnegative integrand exceeding on the respective set ; restricting the integral to that set therefore gives

On , the contrasts are and , so both have magnitude ; and the two oracle rules take opposite values there, Since , it differs from at least one of these two values at every : if then , and if then . Combined with the contrast bound this gives The two witness laws share the covariate marginal , so by monotonicity and subadditivity, Multiplying by and inserting the two regret bounds gives Since and , the left-hand side is Finally, a sum of two reals is at most twice their maximum, so dividing by , Since was arbitrary, this holds for every policy in , for all such sufficiently large .

Proof of Lemma 4.

Write and , and let be the (measurable) region on which the test chooses , so that Throughout we work under the two standing regularity conditions that make the chi-square divergence a genuine density functional: , and is -integrable. Both are displayed hypotheses of the lemma. Under them the chi-square divergence has the second-moment form obtained by expanding and using . Note since .

Step 1: Cauchy–Schwarz mass transfer. Writing the -mass of as a -integral against the density and applying Cauchy–Schwarz,

Step 2: the testing floor. Let be the combined error. Since and , Step 1 gives Suppose, to get a contradiction, that Since , we have , hence and therefore Also gives , so and hence . This contradicts the previous display. Therefore a floor that depends only on and not on the particular pair of laws or on the test.

Step 3: the one-observation sufficient condition. Suppose now , , and the single-observation divergence satisfies , the two standing regularity conditions above now being imposed on the single-observation pair . The product experiment obeys the tensorization identity

Using with and then , so the product divergence is bounded, . Applying Steps 1–2 to the product experiment with the budget yields the uniform testing floor

Proof of Theorem 3.

1. The three ingredients and the lower-bound constant. By Lemma 3, there are constants and such that, for all and every , By Lemma 2, there are and such that, for all , the product laws satisfy with square-integrable density deviation and Applying Lemma 4 to this fixed budget gives a testing floor . Equivalently, applying the binary test that chooses on a measurable event and on , every measurable in the -sample space satisfies Set

2. Membership and the rate identity. By Lemma 1, enlarging the eventual lower bound on if necessary, both witness laws and belong to for all under consideration. The displays in Definitions 6 and 3 give so The policy class is nonempty, so the collection of measurable -valued estimators is nonempty: a constant estimator at any fixed policy in qualifies. For laws in , the well-formedness and bounded-outcome components of Definition 4 imply that the regret maps are nonnegative and bounded by , hence the displayed expectations below are finite.

3. The threshold event. Fix such an , and let be any measurable -valued estimator. Write put The event is measurable because the map from the sample to is measurable, as required for admissible estimators in Definition 5.

4. Two Markov-type threshold bounds. For a probability measure , a nonnegative -integrable function , and a level , the pointwise inequality integrates to Applying this to the nonnegative integrable regret maps under the two product laws gives

On , the minus-regret is below . The separation bound above applies pointwise to the realized policy , so Hence the plus-regret is at least on , and By monotonicity of and the second threshold bound,

5. Combining with the testing floor. Applying the testing floor to and adding the two threshold bounds yields Since ,

6. Passing to the minimax risk. Because , the two expectations and are values indexed by admissible laws in the supremum of Definition 5. Therefore Combining this with the preceding display, every admissible estimator satisfies where depends only on the fixed theorem parameters. Taking the infimum over all measurable -valued estimators in Definition 5 gives for all sufficiently large .

Proof of Theorem 4.
  1. The schedules. The clipping and localization schedules are those of Definition 7: for , while for the clipping schedule is fixed, for every . The input restrictions assumed in the theorem give in both regimes; when they also give , , and the eventual admissibility recorded in Definition 7, Thus for all sufficiently large .

  2. The empty law-class branch. Suppose and the law class in Definition 4 is empty. Then the image set whose supremum defines Definition 11 is empty, since every law in that side-condition domain must first belong to . We use the real-supremum convention . Hence , and for the candidate bound with and is nonnegative. This proves the asserted eventual bound in this branch; the remainder treats the complementary branch.

  3. The clip lies in the admissible clipping range. If , the complementary branch supplies some law . Its strict-overlap endpoint condition in Definition 4 gives , and the input restriction gives If , the maximizer of Definition 7 lies in the feasible box, so . Hence for , and therefore

    Likewise, when , gives for , so with ,

  4. The two ingredients. Lemma 8, applied with these deterministic schedules and the theorem hypotheses, supplies constants and such that, for all sufficiently large , every law satisfying the side-condition domain of Definition 11 obeys the corresponding master inequality. The schedule hypotheses needed there are the eventual bounds , together with eventual constancy of when , established above. Lemma 9, applied to the same schedules, the polynomial nuisance-rate exponents of Assumption 14, the nonnegative nuisance constants, the eventual nonnegativity of , and the positive- admissibility above, supplies constants and bounding the deterministic bracket of the master bound by by Definition 7. Both constant pairs are chosen before . Set

  5. The case . For every in the domain of Definition 11 and all sufficiently large , the strict-overlap branch of Lemma 8 applies because , and gives The branch of Lemma 9 bounds the brace by . Since for all sufficiently large , the two logarithmic factors multiply as , and therefore

  6. The case . For all sufficiently large we have , , and by the schedule facts above. The positive- branch of Lemma 8 applies and yields, for each in the domain, The positive- branch of Lemma 9 bounds the brace by , and the same logarithm identity gives

  7. Passing to the supremum. Fix a sufficiently large , and set The quantity is nonnegative and uniform over the side-condition domain of Definition 11. The preceding case analysis shows that every value in the image set defining is at most . If that image set is empty, the convention gives the same conclusion because ; otherwise the pointwise bound gives the supremum bound. Therefore for all sufficiently large , with the single pair fixed above. This is the claimed bound.

Proof of Lemma 5.

Work throughout in the positive-sample-size case . Fix a realized sample, the pointwise-dense skeleton used in Definition 10, and measurable foldwise nuisance components , , and for each fold . Abbreviate the observed foldwise scores of Definition 9 and the empirical criterion of Definition 10 by The criterion depends on only through the finitely many values .

Step 1: a -near maximizer over the skeleton exists. Since each indicator is or , Thus the set of skeleton scores is nonempty and bounded above. Write , which is finite. Because , , so . By the defining property of the supremum, there is an index with , and hence The set of such indices is a nonempty set of positive integers and has a least element. This least element is the selector index in Definition 10; consequently

Step 2: the selector is measurable and the selected policy lies in . Every enumerated policy belongs to by the skeleton condition in Assumption 4 and Definition 10, so . For measurability, the event is exactly the event that is a -near maximizer and every smaller index is outside the near-maximizer set: Each score is a measurable function of the sample: is measurable by Assumption 4, and the foldwise components , , and are measurable by the regularity condition fixed above. Therefore each is a countable intersection of measurable score comparisons, and is a finite intersection of such sets and their complements. Hence every fiber of is measurable, so is a measurable integer-valued selector. Since is a function on a countable set, the regret map is measurable.

Step 3: the skeleton inequality reaches every comparator in . Let . By Assumption 4 and the pointwise-dense skeleton fixed in Definition 10, there is a sequence from such that is eventually equal to at each covariate value . Applying this eventual equality to the finite set , choose one index such that Since depends on the policy only through those values, Applying the near-maximality from Step 1 to the skeleton index gives Thus every comparator has empirical score at most , which is the -wide -ERM inequality in pointwise-comparator form.

Step 4: the empirical-process form. For any comparator , the definition of gives The inequality from Step 3 is therefore

Finally, under Assumption 5 the oracle policy satisfies , so the comparator inequality applies with .

Proof of Lemma 6.

Fix and one foldwise nuisance triple, and write Since , the outer minimum and inner maximum give the two-sided clip so both denominators are at least and The treatment indicators satisfy and . By Assumption 2, almost surely, and by the assumed foldwise range of the cross-fitted outcome regressions, and . Therefore, by the triangle inequality,

Multiplying the indicator, inverse-denominator, and residual bounds gives, for the two inverse-propensity terms, Applying the triangle inequality to the three summands of Definition 9, Since , we have , so the constant summand is absorbed and As gives , the choice delivers the asserted absolute envelope

Squaring the sharper bound , whose right-hand side is nonnegative, yields so the same constant serves both conclusions.

Proof of Lemma 9.

Throughout, we use the elementary monotonicity rule for powers of : if and , then

Choice of the constants. The pair is chosen once, before the case distinction on , and the same pair serves both branches. Set and The four auxiliary constants are formed from rather than , and are read under the convention that a quotient with vanishing denominator is and that the exponent is when ; with this convention all four are defined for every value of the parameters, including the fixed-overlap regime, where neither nor is available. They are nonnegative there: is a real power of ; and are quotients of the nonnegative numbers and by the nonnegative numbers and ; and is a product of the nonnegative factors , and . Hence and . Neither branch below alters these constants: the fixed-overlap branch never uses the values of , which enter its estimate only as nonnegative slack inside .

Case . Here , so and the two -dependent constants take their stated values and . Write , , so that the deterministic schedules of Definition 7 are By Definition 7,

The pair is selected in Definition 7 as a maximizer of over the feasible box; in particular it is a point of that box, so the two schedules above are the ones the definition prescribes and is evaluated at an admissible pair. The box is nonempty and compact and is continuous on it, so attains its supremum there, and the selected pair realizes that value:

Since a minimum is at most each of its entries, this yields the four exponent inequalities and therefore, combining with ,

Work on the eventual index set on which , the polynomial nuisance bounds of Assumption 14, hold, and . Write for the target rate, and bound the five master-bound terms one at a time.

For the converse term, gives directly

For the empirical-process term, , so raising to the power is an exact identity, and the exponent inequality then applies:

For the clipped product-nuisance remainder, dividing the product bound by ,

For the localized weak-overlap drift term, the factor is nonnegative, so it may be multiplied into the bound on :

For the localization-complement outcome term, the squaring step uses both and , so that implies ; dividing by ,

Adding the five displays gives the sum bound with coefficient ; since , that coefficient is at most , and , so which is the asserted positive- bound with and .

Case . The constants are those fixed above; nothing is rechosen. Here Definition 7 gives the fixed clipping schedule and On the eventual index set where and the polynomial product bound holds, the same monotonicity rule gives the two termwise bounds, again with , Adding them gives the sum bound with coefficient ; since , that coefficient is at most , whence

Proof of Lemma 8.

Constants. Lemma 11 supplies constants and ; put where is the constant appearing in the branch of Lemma 7. Choose a branchwise auxiliary constant as follows: if , take with eventually, as supplied by the fixed-clip hypothesis; otherwise set . Set , which in the strict-overlap branch is the constant appearing in the branch of Lemma 7. Set so that . The logarithmic power in the conclusion is this same .

Large- restriction. Work on the eventual index set on which , , , , , the foldwise nuisance errors lie in , and, when , . Fix such an and a law satisfying the stated law-class, oracle, sampling, nuisance-rate, and bounded-nuisance conditions. The measurability, integrability, and boundedness-above properties of the selected regret and of the offset suprema follow from the bounded cross-fitted nuisances, the measurability of the foldwise nuisance estimators, the positive clip, and the countable pointwise-dense skeleton of Assumption 4, which reduces each supremum over to a supremum over the countable ; these integrability side conditions are therefore established rather than assumed. The nuisance measurability and the memberships of the foldwise nuisance errors are the hypotheses collected in the statement’s measurability and regularity condition, and the memberships enter through the Cauchy–Schwarz step of the drift bound invoked below.

Objects. Write and set the envelope so because . For each fold , let where is the clipped-score drift of Lemma 12 formed from the fold- nuisance triple . Let , so and , and let be the pooled cross-fit centered process built from these increments.

Step 1: the offset bound at envelope . By Lemma 6, the clipped AIPW score admits an almost-sure envelope proportional to ; the truncation performed below needs one explicit admissible constant, which we now record from Definitions 8 and 9 and the boundedness assumptions. On the bounded-outcome event, the clipped propensity lies between and . Since , the two inverse denominators are at most ; since , , and lie in , and the binary treatment indicators are bounded by one, Thus the three terms in the clipped AIPW score satisfy almost surely. Consequently truncating the score at level changes nothing almost surely, and the pooled and foldwise offset suprema built from and from its -truncation coincide almost surely and therefore have equal expectations. Write for the increment obtained from by truncating the score at level . The increments are policy-compatible, measurable, bounded by , and vanish off , so

Hence Lemma 11 applies at sample size with this envelope and, transporting back to the untruncated increments, gives Moreover, since , , , , and ,

Step 2 (): the samplewise selection inequality. Assume , , and ; write and . For every , applying Lemma 12 foldwise with the bounded measurable test function gives the foldwise drift decomposition. The contrast part is identified from the definitions of welfare, oracle policy, and regret in Definition 1: Combining these foldwise identities with the definition of gives the exact empirical-welfare identity By Assumption 5 the oracle lies in , so Lemma 5 applies with comparator and gives . Evaluating the identity at and bounding each term by its absolute value yields the samplewise selection bound

The weighted drift sum is controlled foldwise. Each fold’s nuisance triple satisfies the same error bounds , so Lemma 7 applies to each with the localization window and the admissible clip ; since and , averaging preserves the bound:

The last summand still involves , and is absorbed by Young’s inequality applied with and :

Substituting the last two displays into the selection bound and moving to the left, that is,

Step 3 (): self-bounding and algebra. The samplewise inequality just proved is exactly the input of Lemma 10, applied with envelope , offset bound from the offset step, and slack . It gives Since , the envelope factor is exactly the clipped empirical-process term:

Now abbreviate the five nonnegative target terms and write . Then , and the large- bound above reads . Expanding , the previous bound becomes Each coefficient , , is at most by the definition of ; using to insert the factor into the last three terms, and adding the nonnegative term , gives This proves the positive- branch.

Step 4 (). Now assume and ; on the eventual set considered, is the fixed clip, so is fixed in . The empirical-welfare identity and Lemma 5 give the same samplewise selection bound as above, and the strict-overlap branch of Lemma 7, averaged over the fold weights exactly as before, gives

Hence, using , and Lemma 10 gives With the fixed clip, the envelope factor is a constant times the parametric-envelope rate,

Writing and , the large- bound above gives , so Since and , and , which is the strict-overlap branch.

Proof of Lemma 7.

Fix the evaluation fold and abbreviate the foldwise estimators by . Write and, for a policy , By the exact drift identity of Lemma 12, the conditional-mean drift defined in the statement has the closed form

Step 1: a pointwise majorization. We bound the two factors of separately.

First, the clipping map is -Lipschitz, being a composition of the -Lipschitz maps and ; applying it to and to therefore gives

Second, clipping the true propensity moves it only inside the weak-overlap region: if , then and the clip is inactive, so the displacement is zero; and in all cases the displacement is at most , because and the clipped value lies in . Hence and by the triangle inequality

Third, since , both denominators are at least , so Multiplying the last two displays, Finally, the policy factor satisfies , and it vanishes off ; consequently

Combining the last two displays and cancelling the factor against in the indicator term gives the pointwise majorization

Step 2: integrate. All four majorant terms are -integrable, being products of functions (the assumed memberships of the three nuisance errors, and the indicator, which is bounded). Integrating and applying Cauchy–Schwarz with the assumed rates gives, for ,

and, again by Cauchy–Schwarz with the indicator in the first slot,

Summing the four contributions yields the basic drift bound

Step 3: the case . Assume , , and , and write and . By Lemma 13, Both summands are nonnegative, so taking square roots and using , together with the exponent identity gives

Substituting into Step 2 and abbreviating the four nonnegative quantities we get , all terms being nonnegative. Since , , and , that is, with exactly as in the statement,

Step 4: the case . Assume and . Assumption 10 gives and -almost surely. Since , the set is -null, hence so is its subset : The second term of the basic drift bound of Step 2 therefore vanishes, leaving the last step because and . This is the asserted strict-overlap bound.

References

  • Bickel, Peter J. and Klaassen, Chris A. J. and Ritov, Ya'acov and Wellner, Jon A. (1993). Efficient and Adaptive Estimation for Semiparametric Models. Johns Hopkins University Press.
  • Newey, Whitney K. (1994). The Asymptotic Variance of Semiparametric Estimators. Econometrica. doi
  • Robins, James M. and Rotnitzky, Andrea and Zhao, Lue Ping (1994). Estimation of Regression Coefficients When Some Regressors Are Not Always Observed. Journal of the American Statistical Association. doi
  • Hahn, Jinyong (1998). On the Role of the Propensity Score in Efficient Semiparametric Estimation of Average Treatment Effects. Econometrica. doi
  • Hirano, Keisuke and Imbens, Guido W. and Ridder, Geert (2003). Efficient Estimation of Average Treatment Effects Using the Estimated Propensity Score. Econometrica. doi
  • van der Laan, Mark J. and Rose, Sherri (2011). Targeted Learning. Springer. doi
  • Imbens, Guido W. and Rubin, Donald B. (2015). Causal Inference for Statistics, Social, and Biomedical Sciences. Cambridge University Press. doi
  • Chernozhukov, Victor and Chetverikov, Denis and Demirer, Mert and Duflo, Esther and Hansen, Christian and Newey, Whitney and Robins, James (2018). Double/Debiased Machine Learning for Treatment and Structural Parameters. The Econometrics Journal. doi
  • Chernozhukov, Victor and Escanciano, Juan Carlos and Ichimura, Hidehiko and Newey, Whitney K. and Robins, James M. (2022). Locally Robust Semiparametric Estimation. Econometrica. doi
  • Manski, Charles F. (2004). Statistical Treatment Rules for Heterogeneous Populations. Econometrica. doi
  • Manski, Charles F. (2009). Identification for Prediction and Decision. Harvard University Press.
  • Stoye, J{\"o}rg (2009). Minimax Regret Treatment Choice with Finite Samples. Econometrica.
  • Kitagawa, Toru and Tetenov, Aleksey (2018). Who Should Be Treated? Empirical Welfare Maximization Methods for Treatment Choice. Econometrica. doi
  • Athey, Susan and Wager, Stefan (2021). Policy Learning with Observational Data. Econometrica. doi
  • Luedtke, Alexander and Chambaz, Antoine (2017). Faster Rates for Policy Learning. . arXiv
  • Qian, Min and Murphy, Susan A. (2011). Performance Guarantees for Individualized Treatment Rules. The Annals of Statistics. doi
  • Zhao, Yingqi and Zeng, Donglin and Rush, A. John and Kosorok, Michael R. (2012). Estimating Individualized Treatment Rules Using Outcome Weighted Learning. Journal of the American Statistical Association. doi
  • Zhang, Baqun and Tsiatis, Anastasios A. and Davidian, Marie and Zhang, Min and Laber, Eric B. (2012). Estimating Optimal Treatment Regimes from a Classification Perspective. Stat. doi
  • Dudik, Miroslav and Langford, John and Li, Lihong (2011). Doubly Robust Policy Evaluation and Learning. . arXiv
  • Swaminathan, Adith and Joachims, Thorsten (2015). Counterfactual Risk Minimization: Learning from Logged Bandit Feedback. . arXiv
  • Kallus, Nathan (2017). Balanced Policy Evaluation and Learning. . arXiv
  • Xinkun Nie and Stefan Wager (2017). Quasi-Oracle Estimation of Heterogeneous Treatment Effects. Biometrika. arXiv
  • Sasaki, Yuya and Ura, Takuya (2020). Welfare Analysis via Marginal Treatment Effects. . arXiv
  • Sun, Liyang (2021). Empirical Welfare Maximization with Constraints. . arXiv
  • Kitagawa, Toru and Wang, Weining and Xu, Mengshan (2022). Policy Choice in Time Series by Empirical Welfare Maximization. . arXiv
  • Audibert, Jean-Yves and Tsybakov, Alexandre B. (2007). Fast Learning Rates for Plug-in Classifiers. The Annals of Statistics. doi
  • Massart, Pascal and N{\'e}d{\'e}lec, {\'E}lodie (2006). Risk Bounds for Statistical Learning. The Annals of Statistics. doi
  • Tsybakov, Alexandre B. (2009). Introduction to Nonparametric Estimation. Springer. doi
  • Le Cam, Lucien (1986). Asymptotic Methods in Statistical Decision Theory. Springer-Verlag.
  • Li, Fan and Morgan, Kari Lock and Zaslavsky, Alan M. (2016). Balancing Covariates via Propensity Score Weighting. . arXiv
  • D'Amour, Alexander and Ding, Peng and Feller, Avi and Lei, Lihua and Sekhon, Jasjeet (2017). Overlap in Observational Studies with High-Dimensional Covariates. . arXiv
  • Ben-Michael, Eli and Keele, Luke (2022). Using Balancing Weights to Target the Treatment Effect on the Treated when Overlap is Poor. . arXiv
  • Hill, Jonathan B. and Chaudhuri, Saraswata (2024). Heavy Tail Robust Estimation and Inference for Average Treatment Effects. . arXiv
  • Susmann, Herbert P. and McClean, Alec and D{\'i}az, Iv{\'a}n (2025). Non-overlap Average Treatment Effect Bounds. . arXiv
  • Zhan, Ruohan and Ren, Zhimei and Athey, Susan and Zhou, Zhengyuan (2021). Policy Learning with Adaptively Collected Data. . arXiv
  • Chen, Minshuo and Liu, Hao and Liao, Wenjing and Zhao, Tuo (2020). Doubly Robust Off-Policy Learning on Low-Dimensional Manifolds by Deep Neural Networks. . arXiv
  • Zhao, Pan and Chambaz, Antoine and Josse, Julie and Yang, Shu (2023). Positivity-free Policy Learning with Observational Data. . arXiv
  • Sakaguchi, Shosei (2024). Policy Learning for Optimal Dynamic Treatment Regimes with Observational Data. . arXiv
  • Samuel Girard and Aurelien Bibaut and Arthur Gretton and Nathan Kallus and Houssam Zenati (2025). Fast Best-in-Class Regret for Contextual Bandits. . arXiv
  • Liu, Jingren and Qin, Hanzhang and Liu, Junyi and Chou, Mabel C. and Pang, Jong-Shi (2026). Offline Policy Learning with Weight Clipping and Heaviside Composite Optimization. . arXiv
  • Tsybakov (2004). Optimal aggregation of classifiers in statistical learning. .
  • Bartlett (2005). Local Rademacher complexities. .