CausalSmith · seminar slides
Calibrating Regret under Weak Overlap
Offline policy learning becomes statistically harder when the treatment arm needed to learn small welfare-relevant contrasts is also rarely observed.
slides for A Lower-Bound Calibration for Joint Margin--Overlap Decay in Offline Policy Learning
Overview
- We study deterministic treatment rules learned from offline observational data.
- Welfare regret is measured from the observed-law treatment contrast τP(x), the conditional gain from treatment.
- The law class links two local difficulties: small contrasts and weak overlap.
- The minimax lower bound gives the benchmark exponent r⋆(α,γ).
- A clipped cross-fitted AIPW (augmented inverse propensity weighting) empirical welfare rule attains the benchmark exponent when nuisance and clipping terms are nonbinding.
informal · Theorem T-3 Over the joint margin-overlap law class, worst-case regret is at least order n−r⋆(α,γ).
Motivation
- A policymaker has logged treatment data and wants a rule for who should receive treatment.
- Think of a job-training program assigned with observational discretion.
- For some workers, the earnings gain is close to zero.
- Among those workers, the historical assignment rule may place almost everyone in one arm.
- The target remains the welfare loss from assigning the wrong treatment.
- The statistical difficulty is learning the sign of a small contrast with few observations from the informative arm.
- The question is how this joint scarcity changes the best possible regret rate.
Setup
- One observation is O=(X,A,Y): covariates, binary treatment, bounded outcome.
- The propensity eP(x) is the treatment probability under the observed law.
- The overlap score pP(x) is the distance of eP(x) to the nearest propensity boundary.
- The treatment contrast τP(x) is the conditional mean outcome under treatment minus control.
- A deterministic policy π assigns treatment as a function of X.
For a policy π, its welfare under P is VP(π)=EP[π(X)τP(X)]. The oracle policy is πP⋆(x)=1{τP(x)≥0}. The regret of π under P is RP(π)=VP(πP⋆)−VP(π).
Regret as Weighted Classification
informal · Theorem T-1 Welfare regret is exactly the contrast-weighted probability of disagreeing with the oracle treatment rule.
Under Assumption A-4, for every deterministic policy π∈Π, RP(π)=EP[∣τP(X)∣1{π(X)=πP⋆(X)}]=∫X∣τP(x)∣1{π(x)=πP⋆(x)}dPX(x).
- A wrong decision where ∣τP(X)∣ is large is costly.
- A wrong decision near τP(X)=0 has little welfare cost.
- This identity lets us use margin logic from classification.
Margin and Overlap
- The margin exponent α controls how much covariate mass has a small nonzero treatment contrast.
- The overlap-decay exponent γ controls how weak overlap can concentrate inside that small-contrast region.
- Under strict overlap, γ=0 and the overlap score stays bounded away from zero.
- Under positive overlap decay, weak treatment-arm information is allowed near the decision boundary.
The parameters satisfy 0≤α, 0<Cm, and 0<u0. For every u with 0<u≤u0, P{0<∣τP(X)∣≤u}≤Cmuα.
For all u and v with 0<u≤u0 and v>0 satisfying v≤couγ, the joint small-contrast and weak-overlap probability obeys P{pP(X)≤v, 0<∣τP(X)∣≤u}≤Couαv1/γ, with the convention that v1/γ=1 when γ=0.
Localizing Mistakes
informal · Theorem T-2 Under the margin condition, low-regret policies disagree with the oracle only on a small covariate region.
Suppose that the following conditions hold:
- Margin condition. Assumption A-6 holds.
- Zero-effect convention. Assumption A-7 holds.
- Bounded outcomes. Assumption A-4 holds.
- Disagreement sets. Dπ is defined as in Definition P-2.
Then there exists a constant C=C(Cm,u0,α) such that, for every π∈Π, PX(Dπ)≤CRP(π)α/(1+α).
- This is the regret version of a low-noise classification localization.
- In the job-training example, a good rule can differ from the oracle mainly among workers with nearly zero gains.
Related Literature
- Manski (2004), Manski (2009), Stoye (2009), and Kitagawa and Tetenov (2018) frame treatment choice through welfare regret.
- Athey and Wager (2021) and Chernozhukov et al. (2018, 2022) motivate doubly robust and cross-fitted policy learning.
- Audibert and Tsybakov (2007), Massart and Nedelec (2006), and Tsybakov (2009) explain how margins accelerate excess-risk rates.
- Li et al. (2016), D'Amour et al. (2017), Ben-Michael and Keele (2022), Hill and Chaudhuri (2024), and Susmann et al. (2025) analyze weak-overlap behavior.
- Liu et al. (2026) is closest on clipping-based upper-bound analysis.
Key Idea
- Build two observed laws that agree almost everywhere.
- On a small active block Bn, flip the sign of a tiny treatment contrast.
- Make the informative treatment arm rare on that same block.
- Any learner must choose one sign, and one of the two laws charges regret for that choice.
Calibration
informal · Lemma L-1 The overlap envelope makes βα,γ the largest weak-arm exponent compatible with a block of margin mass hα.
For α≥0, γ>0, h∈(0,1), and β≥0, consider the tight window v=hβ,u=hβ/γ. Then uαv1/γ=h(α+1)β/γ, and uαv1/γ≥hα⟺β≤βα,γ=α+1αγ, with equality in the displayed power comparison whenever β=βα,γ. Moreover βα,γ≥0, and for every β′≥0, the corresponding tight window v=hβ′, u=hβ′/γ satisfies uαv1/γ≥hα⟺β′≤βα,γ. Thus, within this algebraic tight-window envelope, βα,γ is the largest admissible weak-arm exponent.
- The denominator 2+α+βα,γ has three sources.
- The 2 is the cost of distinguishing two close conditional means.
- The α is the margin mass of the active block.
- The βα,γ is the loss from rare informative-arm sampling.
Lower-bound Witness
informal · Lemma L-2 The two local alternatives satisfy the observed-law class restrictions and induce opposite oracle choices on the active block.
informal · Lemma L-3 The product distributions of the two alternatives remain statistically close at sample size n.
informal · Lemma L-4 Every policy incurs at least one of the two witness regrets at order hn1+α.
informal · Lemma L-5 A bounded chi-square product divergence gives a positive lower bound on testing error.
Main Result
informal · Theorem T-3 Under the stated margin-window, overlap-decay, and witness-calibration conditions, minimax regret is at least cn−r⋆(α,γ).
Under Assumption A-12, suppose that α,γ≥0, Cm,Co,co,cB,p>0, cB≤Cm, cB≤Co, 0<p≤1/4, and 8cB<log5. In addition, when γ>0 and α>0, assume cB≤Coco−α/γ, and when γ>0 and α=0, assume cB≤Co4−1/γ. If the policy class Π is nonempty and every π∈Π is measurable, then for the minimax regret Mn(α,γ) of Definition P-8 over the law class Pα,γ of Definition P-10, there is a constant c>0, independent of n, such that for all sufficiently large n, Mn(α,γ)=πinfP∈Pα,γsupEPRP(π)≥cn−r⋆(α,γ).
- This is the observed-law converse benchmark for the class Pα,γ.
- Strict overlap gives the usual margin-driven denominator.
- Joint margin-overlap decay adds the weak-arm exponent to the denominator.
Feasible Rule
- We also analyze one implementable empirical welfare rule.
- Estimate nuisance functions on folds held away from the evaluation fold.
- Clip the estimated propensity into [qn,1−qn].
- Score each observation with the clipped AIPW contrast score.
- Choose the policy with nearly maximal clipped empirical welfare — the empirical risk minimization (ERM) step.
Upper-bound Mechanics
informal · Lemma L-6 The clipped AIPW score equals the treatment contrast plus an explicit drift term.
informal · Lemma L-7 The feasible ERM satisfies the empirical welfare comparison inequality against any policy comparator in the class.
informal · Lemma L-14 Clipping at level q bounds the score envelope at scale 1/q.
informal · Lemma L-13 The clipped-score drift is controlled by product nuisance error, weak-overlap localization, and outcome-regression localization.
Rate Balance
informal · Lemma L-9 The regret of the clipped cross-fitted AIPW ERM is bounded by the lower-bound benchmark, the clipped empirical-process term, and three nuisance-driven terms.
informal · Lemma L-16 Optimizing the clipping and localization schedules yields the feasible exponent rup=rfeas.
Fix one nuisance regime (a,c,Cμ,Cprod), and set Aα=2+α1+α. If γ>0, fix constants uˉ∈(0,u0],q0∈(0,min{1/2,couˉγ}]. For 0≤s≤1/2 and 0≤t≤s/γ, define ϕ(s,t)=min{Aα(1−2s),c−s,a+2γs+2αt,2a−t}. Let (sfeas,tfeas) be any maximizer of ϕ over this compact feasible set, and set gjoint(α,γ,a,c)=ϕ(sfeas,tfeas),qn=q0n−sfeas,un=uˉn−tfeas. Then qn≤counγ for all sufficiently large n, and the solved conditional feasible upper exponent is rup=rfeas=min{r⋆(α,γ),gjoint(α,γ,a,c)}. If γ=0, take a fixed clipping sequence qn=q0 with q0∈(0,p/2], and set rup=rfeas=min{Aα,c}.
- The tuning chooses how aggressively to clip and how tightly to localize near small contrasts.
- In the nonbinding branch, the feasible exponent reaches r⋆(α,γ).
- In the binding branch, the displayed balance gives the procedure’s nuisance-limited exponent.
Conditional Upper Bound
informal · Theorem T-4 Under the stated policy, cross-fitting, nuisance, boundedness, and localized empirical-process conditions, the clipped AIPW ERM has regret at most Cn−rup(logn)p.
Fix parameters with 0≤γ and a nuisance regime (a,c,Cμ,Cprod) satisfying Assumption A-16, with Cμ≥0 and Cprod≥0. Consider the law class Pα,γ of Definition P-10, the feasible cross-fitted clipped-AIPW 1/n-ERM πn of Definition P-7, formed using an enumeration of a pointwise-dense skeleton of Π, and the conditional upper risk Un(α,γ,a,c;η) of Definition P-9. Let the feasible clipping and localization constants q0,uˉ lie in the input domain of Definition P-4: if γ>0, then 0<uˉ≤u0,0<q0≤min{1/2,couˉγ}, and if γ=0, then 0<q0≤p/2. Assume:
- (Policy class and skeleton.) The policy class satisfies Assumption A-9, and the enumeration used by πn is a pointwise-dense skeleton of Π.
- (Cross-fitting.) The number of folds satisfies Assumption A-17.
- (Nuisance boundedness.) The cross-fitted outcome nuisance estimators satisfy Assumption A-15.
- (Nuisance measurability and integrability.) For every n and fold k, μ0,n(−k), μ1,n(−k), and en(−k) are measurable. The rate sequences satisfy rμ,n≥0 and re,n≥0 for all sufficiently large n. For all sufficiently large n, every P∈Pα,γ, and every fold k, the errors μ0,n(−k)−μ0,P, μ1,n(−k)−μ1,P, and en(−k)−eP belong to L2(PX).
- (Localized envelopes.) The uniform localized empirical-process and offset envelope conditions in Assumption A-11 and Assumption A-18 hold.
Then there exist constants C,p with 0<C and 0≤p such that, for all sufficiently large n, the conditional upper risk of πn, uniformly over the side-condition domain in Definition P-9, satisfies Un(α,γ,a,c;η)≤Cn−rup(logn)p, where rup is the feasible upper exponent supplied by Definition P-4 for the above q0 and uˉ.
- The result is conditional on supplied nuisance estimates satisfying the stated rates.
- For γ=0, fixed clipping recovers the strict-overlap margin exponent subject to the product nuisance rate.
- For γ>0, clipping and localization trade off variance, drift, and nuisance learning.
Open Questions
In the strict-gap branch of the conditional feasible achievability bound for the regime-indexed risk Un, it is open whether a genuinely feasible estimator using estimated cross-fitted nuisance functions can attain the converse exponent r⋆(α,γ), or whether weak-arm nuisance learning imposes the slower conditional exponent derived by the feasible upper bound.
- The lower bound supplies the observed-law converse exponent r⋆(α,γ).
- The clipped AIPW analysis supplies the conditional exponent rup for one feasible rule.
- The strict-gap branch identifies where weak-arm nuisance learning is the remaining statistical issue.
Conclusion
- We calibrate offline policy-learning regret when weak overlap and small contrasts occur together.
- The lower-bound exponent comes from balancing contrast size, margin mass, and informative-arm probability.
- The clipped cross-fitted AIPW ERM matches that exponent up to logarithms in the nonbinding nuisance-and-clipping regime.
- In the binding regime, the analysis gives the rule’s nuisance-limited exponent and isolates the feasible-tightness question.