CausalSmith · seminar slides

Calibrating Regret under Weak Overlap

Offline policy learning becomes statistically harder when the treatment arm needed to learn small welfare-relevant contrasts is also rarely observed.

Overview

  • We study deterministic treatment rules learned from offline observational data.
  • Welfare regret is measured from the observed-law treatment contrast τP(x)\tau_P(x), the conditional gain from treatment.
  • The law class links two local difficulties: small contrasts and weak overlap.
  • The minimax lower bound gives the benchmark exponent r(α,γ)r_\star(\alpha,\gamma).
  • A clipped cross-fitted AIPW (augmented inverse propensity weighting) empirical welfare rule attains the benchmark exponent when nuisance and clipping terms are nonbinding.

informal · Theorem T-3 Over the joint margin-overlap law class, worst-case regret is at least order nr(α,γ)n^{-r_\star(\alpha,\gamma)}.

Motivation

  • A policymaker has logged treatment data and wants a rule for who should receive treatment.
  • Think of a job-training program assigned with observational discretion.
  • For some workers, the earnings gain is close to zero.
  • Among those workers, the historical assignment rule may place almost everyone in one arm.
  • The target remains the welfare loss from assigning the wrong treatment.
  • The statistical difficulty is learning the sign of a small contrast with few observations from the informative arm.
  • The question is how this joint scarcity changes the best possible regret rate.

Setup

  • One observation is O=(X,A,Y)O=(X,A,Y): covariates, binary treatment, bounded outcome.
  • The propensity eP(x)e_P(x) is the treatment probability under the observed law.
  • The overlap score pP(x)p_P(x) is the distance of eP(x)e_P(x) to the nearest propensity boundary.
  • The treatment contrast τP(x)\tau_P(x) is the conditional mean outcome under treatment minus control.
  • A deterministic policy π\pi assigns treatment as a function of XX.
Definition P-1 (Welfare \(V_P(\pi)\) and Regret \(R_P(\pi)\))

For a policy π\pi, its welfare under PP is VP(π)=EP ⁣[π(X)τP(X)]. V_P(\pi)=\mathbb E_P\!\left[\pi(X)\tau_P(X)\right]. The oracle policy is πP(x)=1{τP(x)0}. \pi^\star_P(x)=\mathbf 1\{\tau_P(x)\ge 0\}. The regret of π\pi under PP is RP(π)=VP(πP)VP(π). R_P(\pi)=V_P(\pi^\star_P)-V_P(\pi).

Regret as Weighted Classification

informal · Theorem T-1 Welfare regret is exactly the contrast-weighted probability of disagreeing with the oracle treatment rule.

Theorem T-1 (Welfare regret identity)

Under Assumption A-4, for every deterministic policy πΠ\pi\in\Pi, RP(π)=EP ⁣[τP(X)1{π(X)πP(X)}]=XτP(x)1{π(x)πP(x)}dPX(x). R_P(\pi) = E_P\!\left[ |\tau_P(X)|\,\mathbf 1\{\pi(X)\ne\pi^\star_P(X)\} \right] = \int_{\mathcal X} |\tau_P(x)|\,\mathbf 1\{\pi(x)\ne\pi^\star_P(x)\}\,dP_X(x).

  • A wrong decision where τP(X)|\tau_P(X)| is large is costly.
  • A wrong decision near τP(X)=0\tau_P(X)=0 has little welfare cost.
  • This identity lets us use margin logic from classification.

Margin and Overlap

  • The margin exponent α\alpha controls how much covariate mass has a small nonzero treatment contrast.
  • The overlap-decay exponent γ\gamma controls how weak overlap can concentrate inside that small-contrast region.
  • Under strict overlap, γ=0\gamma=0 and the overlap score stays bounded away from zero.
  • Under positive overlap decay, weak treatment-arm information is allowed near the decision boundary.
Assumption A-6 (Margin condition)

The parameters satisfy 0α0\le \alpha, 0<Cm0<C_m, and 0<u00<u_0. For every uu with 0<uu00<u\le u_0, P{0<τP(X)u}Cmuα. P\{0<|\tau_P(X)|\le u\}\le C_m u^\alpha .

Assumption A-8 (Overlap decay)

For all uu and vv with 0<uu00<u\le u_0 and v>0v>0 satisfying vcouγ, v\le c_o u^\gamma, the joint small-contrast and weak-overlap probability obeys P{pP(X)v, 0<τP(X)u}Couαv1/γ, P\{p_P(X)\le v,\ 0<|\tau_P(X)|\le u\} \le C_o u^\alpha v^{1/\gamma}, with the convention that v1/γ=1v^{1/\gamma}=1 when γ=0\gamma=0.

Localizing Mistakes

informal · Theorem T-2 Under the margin condition, low-regret policies disagree with the oracle only on a small covariate region.

Theorem T-2 (Margin Localization)

Suppose that the following conditions hold:

Then there exists a constant C=C(Cm,u0,α) C=C(C_m,u_0,\alpha) such that, for every πΠ\pi\in\Pi, PX(Dπ)CRP(π)α/(1+α). P_X(D_\pi) \le C\,R_P(\pi)^{\alpha/(1+\alpha)} .

  • This is the regret version of a low-noise classification localization.
  • In the job-training example, a good rule can differ from the oracle mainly among workers with nearly zero gains.

Related Literature

  • Manski (2004), Manski (2009), Stoye (2009), and Kitagawa and Tetenov (2018) frame treatment choice through welfare regret.
  • Athey and Wager (2021) and Chernozhukov et al. (2018, 2022) motivate doubly robust and cross-fitted policy learning.
  • Audibert and Tsybakov (2007), Massart and Nedelec (2006), and Tsybakov (2009) explain how margins accelerate excess-risk rates.
  • Li et al. (2016), D'Amour et al. (2017), Ben-Michael and Keele (2022), Hill and Chaudhuri (2024), and Susmann et al. (2025) analyze weak-overlap behavior.
  • Liu et al. (2026) is closest on clipping-based upper-bound analysis.

Key Idea

  • Build two observed laws that agree almost everywhere.
  • On a small active block BnB_n, flip the sign of a tiny treatment contrast.
  • Make the informative treatment arm rare on that same block.
  • Any learner must choose one sign, and one of the two laws charges regret for that choice.
Covariate line pretreatment X Active block Bₙ small covariate mass Local contrast size hₙ, two signs Joint decay caps weak-arm mass Weak arm probability qₙ on Bₙ Signs hard to tell little arm information Rate question how small may hₙ be
illustrative Box-and-arrow schematic showing covariates leading to an active block, local contrast, joint decay, weak arm information, hard sign learning, and the resulting rate question.

Calibration

informal · Lemma L-1 The overlap envelope makes βα,γ\beta_{\alpha,\gamma} the largest weak-arm exponent compatible with a block of margin mass hαh^\alpha.

Lemma L-1 (Overlap envelope calibration)

For α0\alpha\ge0, γ>0\gamma>0, h(0,1)h\in(0,1), and β0\beta\ge0, consider the tight window v=hβ,u=hβ/γ. v=h^\beta, \qquad u=h^{\beta/\gamma}. Then uαv1/γ=h(α+1)β/γ, u^\alpha v^{1/\gamma} = h^{(\alpha+1)\beta/\gamma}, and uαv1/γhαββα,γ=αγα+1, u^\alpha v^{1/\gamma}\ge h^\alpha \quad\Longleftrightarrow\quad \beta\le \beta_{\alpha,\gamma}=\frac{\alpha\gamma}{\alpha+1}, with equality in the displayed power comparison whenever β=βα,γ\beta=\beta_{\alpha,\gamma}. Moreover βα,γ0\beta_{\alpha,\gamma}\ge0, and for every β0\beta'\ge0, the corresponding tight window v=hβv=h^{\beta'}, u=hβ/γu=h^{\beta'/\gamma} satisfies uαv1/γhαββα,γ. u^\alpha v^{1/\gamma}\ge h^\alpha \quad\Longleftrightarrow\quad \beta'\le\beta_{\alpha,\gamma}. Thus, within this algebraic tight-window envelope, βα,γ\beta_{\alpha,\gamma} is the largest admissible weak-arm exponent.

  • The denominator 2+α+βα,γ2+\alpha+\beta_{\alpha,\gamma} has three sources.
  • The 22 is the cost of distinguishing two close conditional means.
  • The α\alpha is the margin mass of the active block.
  • The βα,γ\beta_{\alpha,\gamma} is the loss from rare informative-arm sampling.

Lower-bound Witness

informal · Lemma L-2 The two local alternatives satisfy the observed-law class restrictions and induce opposite oracle choices on the active block.

informal · Lemma L-3 The product distributions of the two alternatives remain statistically close at sample size nn.

informal · Lemma L-4 Every policy incurs at least one of the two witness regrets at order hn1+αh_n^{1+\alpha}.

informal · Lemma L-5 A bounded chi-square product divergence gives a positive lower bound on testing error.

Common covariate law P_X on [0,1] Active block B_n small contrast region Weak treatment arm A=1 on B_n sampled q_n Off block same positive contrast common label 1 Sign-positive law τ=+h_n on B_n Sign-negative law τ=−h_n on B_n Oracle policy choice infer correct sign rare low-signal obs
illustrative Box-and-arrow schematic showing a common covariate law, active block, weak treatment arm, two sign laws, and the oracle policy choice.

Main Result

informal · Theorem T-3 Under the stated margin-window, overlap-decay, and witness-calibration conditions, minimax regret is at least cnr(α,γ)c n^{-r_\star(\alpha,\gamma)}.

Theorem T-3 (Minimax lower bound)

Under Assumption A-12, suppose that α,γ0\alpha,\gamma\ge0, Cm,Co,co,cB,p>0C_m,C_o,c_o,c_B,\underline p>0, cBCmc_B\le C_m, cBCoc_B\le C_o, 0<p1/40<\underline p\le 1/4, and 8cB<log58c_B<\log 5. In addition, when γ>0\gamma>0 and α>0\alpha>0, assume cBCocoα/γc_B\le C_o c_o^{-\alpha/\gamma}, and when γ>0\gamma>0 and α=0\alpha=0, assume cBCo41/γc_B\le C_o 4^{-1/\gamma}. If the policy class Π\Pi is nonempty and every πΠ\pi\in\Pi is measurable, then for the minimax regret Mn(α,γ)M_n(\alpha,\gamma) of Definition P-8 over the law class Pα,γ\mathcal P_{\alpha,\gamma} of Definition P-10, there is a constant c>0c>0, independent of nn, such that for all sufficiently large nn, Mn(α,γ)=infπ^supPPα,γEPRP(π^)cnr(α,γ). M_n(\alpha,\gamma) = \inf_{\widehat\pi} \sup_{P\in\mathcal P_{\alpha,\gamma}} \mathbb E_P R_P(\widehat\pi) \ge c\,n^{-r_\star(\alpha,\gamma)}.

  • This is the observed-law converse benchmark for the class Pα,γ\mathcal P_{\alpha,\gamma}.
  • Strict overlap gives the usual margin-driven denominator.
  • Joint margin-overlap decay adds the weak-arm exponent to the denominator.

Feasible Rule

  • We also analyze one implementable empirical welfare rule.
  • Estimate nuisance functions on folds held away from the evaluation fold.
  • Clip the estimated propensity into [qn,1qn][q_n,1-q_n].
  • Score each observation with the clipped AIPW contrast score.
  • Choose the policy with nearly maximal clipped empirical welfare — the empirical risk minimization (ERM) step.
Offline observations X, A, bounded Y Nuisance estimates Clipped AIPW scores Empirical welfare maximization Learned policy deterministic Welfare regret
illustrative Box-and-arrow schematic showing offline observations, nuisance estimates, clipped AIPW scores, empirical welfare maximization, learned policy, and welfare regret.

Upper-bound Mechanics

informal · Lemma L-6 The clipped AIPW score equals the treatment contrast plus an explicit drift term.

informal · Lemma L-7 The feasible ERM satisfies the empirical welfare comparison inequality against any policy comparator in the class.

informal · Lemma L-14 Clipping at level qq bounds the score envelope at scale 1/q1/q.

informal · Lemma L-13 The clipped-score drift is controlled by product nuisance error, weak-overlap localization, and outcome-regression localization.

Rate Balance

informal · Lemma L-9 The regret of the clipped cross-fitted AIPW ERM is bounded by the lower-bound benchmark, the clipped empirical-process term, and three nuisance-driven terms.

informal · Lemma L-16 Optimizing the clipping and localization schedules yields the feasible exponent rup=rfeasr_{\mathrm{up}}=r_{\mathrm{feas}}.

Definition P-4 (Feasible upper exponent $r_{\mathrm{up}}=r_{\mathrm{feas}}$)

Fix one nuisance regime (a,c,Cμ,Cprod)(a,c,C_\mu,C_{\mathrm{prod}}), and set Aα=1+α2+α. A_\alpha=\frac{1+\alpha}{2+\alpha}. If γ>0\gamma>0, fix constants uˉ(0,u0],q0(0,min{1/2,couˉγ}]. \bar u\in(0,u_0], \qquad q_0\in\bigl(0,\min\{1/2,c_o\bar u^\gamma\}\bigr]. For 0s1/20\le s\le 1/2 and 0ts/γ0\le t\le s/\gamma, define ϕ(s,t)=min{Aα(12s),cs,a+s2γ+αt2,2at}. \phi(s,t) = \min\left\{ A_\alpha(1-2s),\, c-s,\, a+\frac{s}{2\gamma}+\frac{\alpha t}{2},\, 2a-t \right\}. Let (sfeas,tfeas)(s_{\mathrm{feas}},t_{\mathrm{feas}}) be any maximizer of ϕ\phi over this compact feasible set, and set gjoint(α,γ,a,c)=ϕ(sfeas,tfeas),qn=q0nsfeas,un=uˉntfeas. g_{\mathrm{joint}}(\alpha,\gamma,a,c) = \phi(s_{\mathrm{feas}},t_{\mathrm{feas}}), \qquad q_n=q_0 n^{-s_{\mathrm{feas}}}, \qquad u_n=\bar u n^{-t_{\mathrm{feas}}}. Then qncounγq_n\le c_o u_n^\gamma for all sufficiently large nn, and the solved conditional feasible upper exponent is rup=rfeas=min{r(α,γ),gjoint(α,γ,a,c)}. r_{\mathrm{up}} = r_{\mathrm{feas}} = \min\left\{r_\star(\alpha,\gamma),\,g_{\mathrm{joint}}(\alpha,\gamma,a,c)\right\}. If γ=0\gamma=0, take a fixed clipping sequence qn=q0q_n=q_0 with q0(0,p/2]q_0\in(0,\underline p/2], and set rup=rfeas=min{Aα,c}. r_{\mathrm{up}} = r_{\mathrm{feas}} = \min\{A_\alpha,c\}.

  • The tuning chooses how aggressively to clip and how tightly to localize near small contrasts.
  • In the nonbinding branch, the feasible exponent reaches r(α,γ)r_\star(\alpha,\gamma).
  • In the binding branch, the displayed balance gives the procedure’s nuisance-limited exponent.

Conditional Upper Bound

informal · Theorem T-4 Under the stated policy, cross-fitting, nuisance, boundedness, and localized empirical-process conditions, the clipped AIPW ERM has regret at most Cnrup(logn)pC n^{-r_{\mathrm{up}}}(\log n)^p.

Theorem T-4 (Feasible Upper Exponent)

Fix parameters with 0γ0\le \gamma and a nuisance regime (a,c,Cμ,Cprod)(a,c,C_\mu,C_{\mathrm{prod}}) satisfying Assumption A-16, with Cμ0C_\mu\ge0 and Cprod0C_{\mathrm{prod}}\ge0. Consider the law class Pα,γ\mathcal P_{\alpha,\gamma} of Definition P-10, the feasible cross-fitted clipped-AIPW 1/n1/n-ERM π^n\widehat\pi_n of Definition P-7, formed using an enumeration of a pointwise-dense skeleton of Π\Pi, and the conditional upper risk Un(α,γ,a,c;η^)U_n(\alpha,\gamma,a,c;\widehat\eta) of Definition P-9. Let the feasible clipping and localization constants q0,uˉq_0,\bar u lie in the input domain of Definition P-4: if γ>0\gamma>0, then 0<uˉu0,0<q0min{1/2,couˉγ}, 0<\bar u\le u_0,\qquad 0<q_0\le \min\{1/2,c_o\bar u^\gamma\}, and if γ=0\gamma=0, then 0<q0p/2. 0<q_0\le \underline p/2. Assume:

  • (Policy class and skeleton.) The policy class satisfies Assumption A-9, and the enumeration used by π^n\widehat\pi_n is a pointwise-dense skeleton of Π\Pi.
  • (Cross-fitting.) The number of folds satisfies Assumption A-17.
  • (Nuisance boundedness.) The cross-fitted outcome nuisance estimators satisfy Assumption A-15.
  • (Nuisance measurability and integrability.) For every nn and fold kk, μ^0,n(k)\widehat\mu_{0,n}^{(-k)}, μ^1,n(k)\widehat\mu_{1,n}^{(-k)}, and e^n(k)\widehat e_n^{(-k)} are measurable. The rate sequences satisfy rμ,n0r_{\mu,n}\ge0 and re,n0r_{e,n}\ge0 for all sufficiently large nn. For all sufficiently large nn, every PPα,γP\in\mathcal P_{\alpha,\gamma}, and every fold kk, the errors μ^0,n(k)μ0,P\widehat\mu_{0,n}^{(-k)}-\mu_{0,P}, μ^1,n(k)μ1,P\widehat\mu_{1,n}^{(-k)}-\mu_{1,P}, and e^n(k)eP\widehat e_n^{(-k)}-e_P belong to L2(PX)L^2(P_X).
  • (Localized envelopes.) The uniform localized empirical-process and offset envelope conditions in Assumption A-11 and Assumption A-18 hold.

Then there exist constants C,pC,p with 0<C0<C and 0p0\le p such that, for all sufficiently large nn, the conditional upper risk of π^n\widehat\pi_n, uniformly over the side-condition domain in Definition P-9, satisfies Un(α,γ,a,c;η^)Cnrup(logn)p, U_n(\alpha,\gamma,a,c;\widehat\eta)\le C n^{-r_{\mathrm{up}}}(\log n)^p, where rupr_{\mathrm{up}} is the feasible upper exponent supplied by Definition P-4 for the above q0q_0 and uˉ\bar u.

  • The result is conditional on supplied nuisance estimates satisfying the stated rates.
  • For γ=0\gamma=0, fixed clipping recovers the strict-overlap margin exponent subject to the product nuisance rate.
  • For γ>0\gamma>0, clipping and localization trade off variance, drift, and nuisance learning.

Open Questions

Definition P-12 (Feasible Tightness Question)

In the strict-gap branch of the conditional feasible achievability bound for the regime-indexed risk UnU_n, it is open whether a genuinely feasible estimator using estimated cross-fitted nuisance functions can attain the converse exponent r(α,γ)r_\star(\alpha,\gamma), or whether weak-arm nuisance learning imposes the slower conditional exponent derived by the feasible upper bound.

  • The lower bound supplies the observed-law converse exponent r(α,γ)r_\star(\alpha,\gamma).
  • The clipped AIPW analysis supplies the conditional exponent rupr_{\mathrm{up}} for one feasible rule.
  • The strict-gap branch identifies where weak-arm nuisance learning is the remaining statistical issue.

Conclusion

  • We calibrate offline policy-learning regret when weak overlap and small contrasts occur together.
  • The lower-bound exponent comes from balancing contrast size, margin mass, and informative-arm probability.
  • The clipped cross-fitted AIPW ERM matches that exponent up to logarithms in the nonbinding nuisance-and-clipping regime.
  • In the binding regime, the analysis gives the rule’s nuisance-limited exponent and isolates the feasible-tightness question.