CausalSmith · seminar slides

Contour Instruments for Partially Linear Models

We use zeros of the treatment-innovation transform to build contour instruments that identify and estimate the partially linear treatment coefficient at fixed-separation root-nn scale.

Overview

  • We study a partially linear model with a supplied treatment code.
  • The target is θ0\theta_0, the treatment coefficient after residualizing treatment on covariates.
  • The key resource is non-Gaussian treatment noise with a fixed nonzero cumulant.
  • A zero of the treatment-innovation moment-generating function creates an instrument that survives treatment-code error.
  • A finite contour bank turns that population identity into a fixed-code estimator with minimax mean-squared error of order n1n^{-1}.

informal · Theorem T-3 Under fixed cumulant separation and the stated fixed-code stability and boundedness conditions, the contour statistic has minimax mean-squared error between c/nc/n and C/nC/n, and its generalized quantile error is at most C/(γn)\sqrt{C/(\gamma n)}.

Motivation

  • Partially linear models estimate a low-dimensional treatment effect while allowing flexible covariate adjustment.
  • Robinson (1988) gives the classical root-nn residualized construction.
  • Modern double/debiased machine learning keeps the same target stable under learned nuisance functions; see Chernozhukov et al. (2018).
  • Here the supplied treatment code may have residual error.
  • Think of TT as an endogenous exposure whose predictable part is learned from covariates.
  • We observe Zn=Tgˉn(X)Z_n=T-\bar g_n(X), the residualized treatment formed with the supplied clipped treatment code.
  • The code error shifts the true treatment innovation by Dn(X)=g0(X)gˉn(X)D_n(X)=g_0(X)-\bar g_n(X).
  • Non-Gaussian treatment noise gives complex transform zeros that can protect estimation from these real shifts.

Setup

  • We observe O=(X,T,Y)O=(X,T,Y).
  • The conditional treatment and outcome regressions are g0(X)=E[TX]g_0(X)=E[T\mid X] and q0(X)=E[YX]q_0(X)=E[Y\mid X].
  • The treatment innovation is η=Tg0(X)\eta=T-g_0(X).
  • The outcome innovation is ξ=Yq0(X)θ0η\xi=Y-q_0(X)-\theta_0\eta.
  • The supplied treatment and outcome codes are clipped to fixed ranges before use.
  • The sample is i.i.d. and split into a pilot fold and an evaluation fold.
Definition P-19 (Observed-data partially linear model)

An observation is the triple O=(X,T,Y)O=(X,T,Y), where XX takes values in a measurable covariate space X\mathcal X and T,YRT,Y\in\mathbb R. A model over a parameter record consists of a probability law PP for OO under which TT and YY are integrable, a real coefficient θ0\theta_0, measurable regression functions g0,q0g_0,q_0 on the covariate space with g0(X)=EP[TX],q0(X)=EP[YX] g_0(X)=E_P[T\mid X],\qquad q_0(X)=E_P[Y\mid X] PP-almost surely, and deterministic measurable supplied code sequences g^=(g^n)nN\widehat g=(\widehat g_n)_{n\in\mathbb N} and q^=(q^n)nN\widehat q=(\widehat q_n)_{n\in\mathbb N} for g0g_0 and q0q_0. The treatment and outcome innovations are η=Tg0(X),ξ=Yq0(X)θ0η, \eta=T-g_0(X),\qquad \xi=Y-q_0(X)-\theta_0\eta, and the clipped supplied codes at index nn are gˉn(x)=min{max{g^n(x),Cg},Cg},qˉn(x)=min{max{q^n(x),Cq},Cq}. \bar g_n(x)=\min\{\max\{\widehat g_n(x),-C_g\},C_g\},\qquad \bar q_n(x)=\min\{\max\{\widehat q_n(x),-C_q\},C_q\}. The covariate marginal of PP is written PXP_X. A model is written mm, and when the law has to be displayed the coefficient is written θ0(P)\theta_0(P).

Assumptions

  • The treatment innovation is independent of covariates.
  • The outcome innovation has conditionally mean zero given X,TX,T.
  • The coefficient and regression functions are bounded.
  • Treatment and outcome innovations are sub-Gaussian under our fixed envelope convention.
  • The kkth cumulant of the treatment innovation is separated from zero by δ\delta, the cumulant-separation level.
  • The treatment-code radius ε1,n\varepsilon_{1,n} controls gˉng0L1(PX)\|\bar g_n-g_0\|_{L^1(P_X)}.
Assumption A-10 (Cumulant separation)

In the setting of Definition P-19, Definition P-18, the kkth cumulant of the treatment innovation η\eta satisfies κk(η)δ. |\kappa_k(\eta)|\ge\delta.

Key Idea

  • Let M(z)=E[ezη]M(z)=E[e^{z\eta}] be the treatment-innovation moment-generating function.
  • If MM has a zero z0z_0, a polynomial-exponential weight built from z0z_0 annihilates every real shift of η\eta.
  • Conditioning on XX, that same weight becomes an instrument for the code residual ZnZ_n.
  • The resulting ratio identifies θ0\theta_0 whenever the denominator is separated from zero.

informal · Theorem T-1 A zero of the treatment-innovation transform yields a valid shifted-residual instrument and identifies θ0\theta_0 through a population ratio when the denominator is nonzero.

Contour Identification

  • A known zero gives one instrument; a contour handles the zeros without naming them one by one.
  • FnF_n is the observable transform of the residualized treatment.
  • GnG_n is the observable outcome-weighted transform.
  • A contour enclosing at least one zero averages Gn/FnG_n/F_n over the enclosed zero structure.
  • Nuisance zero-freeness keeps the treatment-code error from creating spurious denominator zeros.
Code and data supplied inputs Residual Z_n treatment code F_n transform carries zero locations G_n transform numerator information Nuisance factor zero-free inside Contour C_j zero-free boundary enclosed zeros Contour average counts zeros averages G_n/F_n Estimate θ₀ identified coefficient
illustrative Box-and-arrow schematic showing supplied code and data forming residuals, observable treatment and outcome transforms, a zero-free nuisance factor, a selected contour, a contour average, and the identified coefficient.

informal · Theorem T-2 On a contour where FnF_n is zero-free at the boundary, the nuisance factor is zero-free inside, and at least one zero is enclosed, the normalized contour ratio equals θ0(P)\theta_0(P).

Estimator

  • The estimator uses a finite translated-dyadic bank of circles between the zero-localization radius and an outer radius.
  • The pilot fold selects a circle with a positive winding count and a certified denominator margin.
  • The evaluation fold computes the normalized contour average on the selected circle.
  • The real midpoint is clipped to the known coefficient range.
  • Every empirical branch returns a value, so the estimator is always well defined.
Observations pilot fold evaluation fold Residualize treatments Z_n Empirical transforms F̂ and Ĝ Zero radius R₀ from separation R₁=R₀+1 Contour bank circles R₀ to R₁ Pilot selection winding count lower modulus Contour average eval Ĝ/F̂ selected circle Clipped midpoint real midpoint [−Cθ,Cθ]
illustrative Box-and-arrow schematic showing observations split into pilot and evaluation folds, residualization, empirical transforms, zero-radius bank construction, pilot contour selection, evaluation contour averaging, and clipped output.
Definition P-13 (Adaptive contour estimator \(\widehat\theta_{\mathrm{spec},n}\))

The estimator θ^spec,n\widehat\theta_{\mathrm{spec},n} is the packaged adaptive contour estimator on the public domain where the current supplied treatment-regression code underlying gˉn\bar g_n is measurable. It takes as inputs O1,,OnO_1,\ldots,O_n, the supplied treatment-regression code sequence through its current clipped values gˉn\bar g_n, the parameter record pp, and fixed certified records p\mathfrak p_{\star} and cθ,\mathfrak c_{\theta,\star} representing k,δ,ψη,R0,R1k,\delta,\psi_{\eta},R_0,R_1, and CθC_{\theta}, with R0=Ak(ψη2δ)(k2)1,R1=R0+1. R_0=A_k\left(\frac{\psi_{\eta}^2}{\delta}\right)^{(k-2)^{-1}}, \qquad R_1=R_0+1. Here pp contains n1n\ge1, r,kNr,k\in\mathbb N with k=r+1k=r+1 and r2r\ge2, s[0,]s\in[0,\infty] with rsr\le s, γ(1/2,1)\gamma\in(1/2,1), positive constants Cθ,Cg,Cq,ψη,ψξ,δ,σC_{\theta},C_g,C_q,\psi_{\eta},\psi_{\xi},\delta,\sigma, and nonnegative nonincreasing sequences ε1n,ε2n:NR\varepsilon_{1n},\varepsilon_{2n}:\mathbb N\to\mathbb R on positive indices.

  1. Define the deterministic folds I0={i:i<n/2},I1={i:in/2}. I_0=\{i:i<\lfloor n/2\rfloor\}, \qquad I_1=\{i:i\ge\lfloor n/2\rfloor\}. For each real observation coordinate xx and integer m0m\ge0, define the floor-dyadic interval Im(x)=[2(m+1)2m+1x,2(m+1){2m+1x+1}]. \mathcal I_m(x)= \left[ 2^{-(m+1)}\lfloor2^{m+1} x\rfloor, 2^{-(m+1)}\{\lfloor2^{m+1} x\rfloor+1\} \right].
  2. Form the residualized treatment values Zn,i=Tigˉn(Xi),i=1,,n. Z_{n,i}=T_i-\bar g_n(X_i), \qquad i=1,\ldots,n. For each fold a{0,1}a\in\{0,1\}, derivative order pp, and complex zz, define F^a,n(p)(z)=(max{Ia,1})1iIaZn,ipezZn,i,p{0,1,2}, \widehat F_{a,n}^{(p)}(z) = \bigl(\max\{|I_a|,1\}\bigr)^{-1}\sum_{i\in I_a}Z_{n,i}^{p}e^{zZ_{n,i}}, \qquad p\in\{0,1,2\}, and G^a,n(p)(z)=(max{Ia,1})1iIaYiZn,ipezZn,i,p{0,1}. \widehat G_{a,n}^{(p)}(z) = \bigl(\max\{|I_a|,1\}\bigr)^{-1}\sum_{i\in I_a}Y_iZ_{n,i}^{p}e^{zZ_{n,i}}, \qquad p\in\{0,1\}. Thus an empty fold contributes the empty sum and is normalized by one.
  3. For represented observation data, define s~=((ti,yi,gi)i=1n), \widetilde{\mathfrak s} = \bigl((\mathfrak t_i,\mathfrak y_i,\mathfrak g_i)_{i=1}^n\bigr), where ti\mathfrak t_i, yi\mathfrak y_i, and gi\mathfrak g_i are certified-real records naming TiT_i, YiY_i, and gˉn(Xi)\bar g_n(X_i). The full represented-data input is s=((ti,yi,gi)i=1n,p,cθ,). \mathfrak s_{\star} = \bigl((\mathfrak t_i,\mathfrak y_i,\mathfrak g_i)_{i=1}^n, \mathfrak p_{\star},\mathfrak c_{\theta,\star}\bigr). For i=1,,ni=1,\ldots,n, put zi=tigi. \mathfrak z_i=\mathfrak t_i-\mathfrak g_i.
  4. Invoke the contour bank determined by p\mathfrak p_{\star}. For the jjth circle, define Cj={zj(t):0t1},zj(t)=ρje2πit. C_j=\{z_j(t):0\le t\le1\}, \qquad z_j(t)=\rho_je^{2\pi it}. At error tolerance one, refine zi\mathfrak z_i, yi\mathfrak y_i, and ρj\mathfrak\rho_j using their own executable moduli, and let UZ,iU_{Z,i}, UY,iU_{Y,i}, and Uρ,jU_{\rho,j} be the largest absolute endpoint of the corresponding rational intervals. For a canonical observation name this is the interval Iμ(1)\mathcal I_{\mu(1)} with the convention above; for zi\mathfrak z_i it is the certified subtraction interval obtained from ti\mathfrak t_i and gi\mathfrak g_i, whose modulus uses the two half-error input moduli. Using ex3xe^x\le3^{\lceil x\rceil} for x0x\ge0 and π<4\pi<4, form rational upper bounds for the empirical derivative sums on each CjC_j.
  5. For each jj, compute a rational bracket of width at most a/64a_{\star}/64 enclosing infzCjF^0,n(z). \inf_{z\in C_j}|\widehat F_{0,n}(z)|. Let mjm_j be the lower endpoint of this bracket. When mj>0m_j>0, define W0j(t)=ρje2πitF^0,n(zj(t))F^0,n(zj(t)). W_{0j}(t)= \rho_je^{2\pi it} \frac{\widehat F'_{0,n}(z_j(t))} {\widehat F_{0,n}(z_j(t))}. Use the rational Lipschitz bound W0j(t)64[ρjsupF^0,nmj+ρj2{supF^0,nmj+supF^0,n2mj2}]. |W'_{0j}(t)|\le 64\left[ \rho_j\frac{\sup|\widehat F'_{0,n}|}{m_j} +\rho_j^2\left\{ \frac{\sup|\widehat F''_{0,n}|}{m_j} +\frac{\sup|\widehat F'_{0,n}|^2}{m_j^2} \right\}\right].
  6. Compute a rectangle enclosure, with coordinate widths strictly below 1/41/4, for 01W0j(t)dt=12πiCjF^0,n(z)F^0,n(z)dz. \int_0^1W_{0j}(t)\,dt = \frac{1}{2\pi i}\oint_{C_j} \frac{\widehat F'_{0,n}(z)} {\widehat F_{0,n}(z)}\,dz. Define Nj=12πiCjF^0,n(z)F^0,n(z)dz N_j= \frac{1}{2\pi i}\oint_{C_j} \frac{\widehat F'_{0,n}(z)} {\widehat F_{0,n}(z)}\,dz when the real coordinate contains a unique nonnegative integer and the imaginary coordinate contains zero. A circle is admissible when mja/2andNj1. m_j\ge a_{\star}/2 \qquad\text{and}\qquad N_j\ge1. Let j=min{j:jargmax:ma/2, N1m}. j_{\star} = \min\left\{ j: j\in\arg\max_{\ell:\,m_{\ell}\ge a_{\star}/2,\ N_{\ell}\ge1}m_{\ell} \right\}. If the admissible set is empty, set θ^spec,n=0. \widehat\theta_{\mathrm{spec},n}=0.
  7. On the evaluation fold for the selected circle CjC_{j_{\star}}, compute a lower endpoint mj(1)infzCjF^1,n(z). m_{j_{\star}}^{(1)} \le \inf_{z\in C_{j_{\star}}}|\widehat F_{1,n}(z)|. If mj(1)<a/4, m_{j_{\star}}^{(1)}<a_{\star}/4, set θ^spec,n=0. \widehat\theta_{\mathrm{spec},n}=0.
  8. Otherwise define the evaluation-fold contour integrand Vj(t)=ρje2πitG^1,n(zj(t))F^1,n(zj(t)). V_{j_{\star}}(t)= \rho_{j_{\star}}e^{2\pi it} \frac{\widehat G_{1,n}(z_{j_{\star}}(t))} {\widehat F_{1,n}(z_{j_{\star}}(t))}. Use the rational Lipschitz bound Vj(t)64[ρjsupG^1,nmj(1)+ρj2{supG^1,nmj(1)+supG^1,nsupF^1,n(mj(1))2}]. |V'_{j_{\star}}(t)|\le 64\left[ \rho_{j_{\star}}\frac{\sup|\widehat G_{1,n}|}{m_{j_{\star}}^{(1)}} +\rho_{j_{\star}}^2\left\{ \frac{\sup|\widehat G'_{1,n}|}{m_{j_{\star}}^{(1)}} +\frac{\sup|\widehat G_{1,n}|\sup|\widehat F'_{1,n}|}{(m_{j_{\star}}^{(1)})^2} \right\}\right]. For each rational tolerance e>0e>0 and rational Lipschitz bound LL, the schedule spends one third of the tolerance on discretization and uses Nmesh=3Le+1, N_{\mathrm{mesh}}=\left\lceil\frac{3L}{e}\right\rceil+1, so that L/Nmeshe/3L/N_{\mathrm{mesh}}\le e/3.
  9. Enclose the unnormalized contour integral 01Vj(t)dt=12πiCjG^1,n(z)F^1,n(z)dz. \int_0^1V_{j_{\star}}(t)\,dt = \frac{1}{2\pi i}\oint_{C_{j_{\star}}} \frac{\widehat G_{1,n}(z)} {\widehat F_{1,n}(z)}\,dz. Divide the resulting rectangle by NjN_{j_{\star}} to enclose 1Nj2πiCjG^1,n(z)F^1,n(z)dz \frac{1}{N_{j_{\star}}2\pi i} \oint_{C_{j_{\star}}} \frac{\widehat G_{1,n}(z)} {\widehat F_{1,n}(z)}\,dz with coordinate radius at most 1/n1/n. Take yQy\in\mathbb Q to be the real midpoint of that enclosure, so that yy approximates the normalized contour integral to within 1/n1/n: yRe(1Nj2πiCjG^1,n(z)F^1,n(z)dz)1n. \left|\,y-\operatorname{Re}\left(\frac{1}{N_{j_{\star}}2\pi i} \oint_{C_{j_{\star}}}\frac{\widehat G_{1,n}(z)}{\widehat F_{1,n}(z)}\,dz\right)\right|\le\frac1n. Only rational quantities are ever formed, so the estimator returns this certified rational approximation rather than the integral itself; the 1/n1/n tolerance is negligible beside the statistical error.
  10. Output θ^spec,n=clipCθ(y)=y+CθyCθ2=Π[Cθ,Cθ](y). \widehat\theta_{\mathrm{spec},n} = \operatorname{clip}_{C_{\theta}}(y) = \frac{|y+C_{\theta}|-|y-C_{\theta}|}{2} = \Pi_{[-C_{\theta},C_{\theta}]}(y).

Main Result

  • Fixed cumulant separation localizes a transform zero in a fixed search region.
  • The L1(PX)L^1(P_X) treatment-code gate preserves the zero geometry of the residual transform.
  • Uniform empirical transform control gives root-nn accuracy on the selected contour.
  • The lower bound comes from a one-dimensional submodel with fixed treatment law and 1/n1/\sqrt n target separation.
Theorem T-3 (Fixed-code root-\(n\) minimaxity under fixed cumulant separation)

Fix r2r\ge 2 and positive constants δ,Cθ,Cg,Cq,ψη,ψξ\delta,C_{\theta},C_g,C_q,\psi_{\eta},\psi_{\xi}. Then there are constants c>0c>0 and C>0C>0, depending only on these fixed primitive constants, such that the following statements hold. For every measurable covariate space, let p0p_0 be a parameter record whose primitive entries r,δ,Cθ,Cg,Cq,ψη,ψξr,\delta,C_{\theta},C_g,C_q,\psi_{\eta},\psi_{\xi} equal the fixed constants above, and fix one experiment-wide record consisting of the certified contour-bank input and the represented range input for p0p_0. For every parameter record pp with the same primitive entries as p0p_0, and for every base model satisfying the usual measurability, conditional-mean, and integrability requirements, write nn for the sample size of pp. Let θ^spec,n\widehat\theta_{\mathrm{spec},n} be the total Borel statistic of Definition P-13, constructed from the transported fixed records and the base treatment-code sequence. Define the fixed-code non-Gaussian class PNG,n(p;base)={P: PPNG,n of def:non-gaussian-class, gˉn(P)=gˉn(base), qˉn(P)=qˉn(base)}. \mathcal P_{\mathrm{NG},n}(p;\mathrm{base}) = \{P:\ P\in\mathcal P_{\mathrm{NG},n}\text{ of }\text{def:non-gaussian-class},\ \bar g_n(P)=\bar g_n(\mathrm{base}),\ \bar q_n(P)=\bar q_n(\mathrm{base})\}. Define the fixed-code JMS ACE class PACE,nJMS(p;base)={P: PPACE,nJMS of def:jms-ace-class, gˉn(P)=gˉn(base), qˉn(P)=qˉn(base)}. \mathcal P_{\mathrm{ACE},n}^{\mathrm{JMS}}(p;\mathrm{base}) = \{P:\ P\in\mathcal P_{\mathrm{ACE},n}^{\mathrm{JMS}}\text{ of }\text{def:jms-ace-class},\ \bar g_n(P)=\bar g_n(\mathrm{base}),\ \bar q_n(P)=\bar q_n(\mathrm{base})\}.

  • (Measurable statistic.) The statistic θ^spec,n\widehat\theta_{\mathrm{spec},n} is measurable.
  • (Represented execution.) For every compiled bounded spectral adapter satisfying the full canonical build-and-compilation specification for pp, the transported fixed records, and the base treatment-code sequence, the represented-data execution realizes the same θ^spec,n\widehat\theta_{\mathrm{spec},n} as in Definition P-13.
  • (Class relations.) Every law in PACE,nJMS\mathcal P_{\mathrm{ACE},n}^{\mathrm{JMS}} belongs to PNG,n\mathcal P_{\mathrm{NG},n}. If s=rs=r, then PACE,nJMS\mathcal P_{\mathrm{ACE},n}^{\mathrm{JMS}} coincides with the Ls(PX)L^s(P_X)-restricted subclass PNG,nACE\mathcal P_{\mathrm{NG},n}^{\mathrm{ACE}} of Definition P-10.
  • (Non-Gaussian fixed-code risk.) If the sample law is the i.i.d. product law of Assumption A-1, PNG,n(p;base)\mathcal P_{\mathrm{NG},n}(p;\mathrm{base}) is nonempty, n2n\ge 2, and ε1,n{4R1exp(2CgR1)}1, \varepsilon_{1,n}\le \{4R_1\exp(2C_gR_1)\}^{-1}, where R1R_1 is the search radius from the contour-bank construction in Definition P-13, then cnRnNGsupPPNG,n(p;base)EP ⁣[(θ^spec,nθ0(P))2]Cn, \frac{c}{n} \le \mathfrak R_n^{\mathrm{NG}} \le \sup_{P\in\mathcal P_{\mathrm{NG},n}(p;\mathrm{base})} E_P\!\left[(\widehat\theta_{\mathrm{spec},n}-\theta_0(P))^2\right] \le \frac{C}{n}, and supPPNG,n(p;base)QP,1γ ⁣(θ^spec,nθ0(P))Cγn, \sup_{P\in\mathcal P_{\mathrm{NG},n}(p;\mathrm{base})} Q_{P,1-\gamma}\!\left(|\widehat\theta_{\mathrm{spec},n}-\theta_0(P)|\right) \le \sqrt{\frac{C}{\gamma n}}, with generalized lower quantiles as in Definition P-15.
  • (ACE fixed-code risk.) If PACE,nJMS(p;base)\mathcal P_{\mathrm{ACE},n}^{\mathrm{JMS}}(p;\mathrm{base}) is nonempty, n2n\ge 2, and ε1,n{4R1exp(2CgR1)}1, \varepsilon_{1,n}\le \{4R_1\exp(2C_gR_1)\}^{-1}, then cninfθ~supPPACE,nJMS(p;base)EP ⁣[(θ~θ0(P))2]supPPACE,nJMS(p;base)EP ⁣[(θ^spec,nθ0(P))2]Cn, \frac{c}{n} \le \inf_{\widetilde\theta} \sup_{P\in\mathcal P_{\mathrm{ACE},n}^{\mathrm{JMS}}(p;\mathrm{base})} E_P\!\left[(\widetilde\theta-\theta_0(P))^2\right] \le \sup_{P\in\mathcal P_{\mathrm{ACE},n}^{\mathrm{JMS}}(p;\mathrm{base})} E_P\!\left[(\widehat\theta_{\mathrm{spec},n}-\theta_0(P))^2\right] \le \frac{C}{n}, where the infimum is over measurable real-valued estimators of O1,,OnO_1,\ldots,O_n. Moreover, supPPACE,nJMS(p;base)QP,1γ ⁣(θ^spec,nθ0(P))Cγn. \sup_{P\in\mathcal P_{\mathrm{ACE},n}^{\mathrm{JMS}}(p;\mathrm{base})} Q_{P,1-\gamma}\!\left(|\widehat\theta_{\mathrm{spec},n}-\theta_0(P)|\right) \le \sqrt{\frac{C}{\gamma n}}.

Sequence Result

  • Along a sequence of experiments with the same primitive constants, one contour bank serves the sequence.
  • On the cumulant-separated non-Gaussian fixed-code class, the minimax mean-squared risk stays between constants times nj1n_j^{-1}.
  • On the bounded-outcome Gaussian comparison class of Jin, Mackey, and Syrgkanis (JMS), the target degenerates to zero under the stated simultaneous restrictions.

informal · Theorem T-5 Along fixed-constant non-Gaussian experiment sequences, the same contour bank gives minimax mean-squared risk of order nj1n_j^{-1}, while the bounded-outcome Gaussian JMS fixed-code risks are zero.

Related Literature

  • Partially linear models trace to Engle et al. (1986), Robinson (1988), Speckman (1988), and Härdle et al. (2000).
  • Orthogonal and debiased learning build on semiparametric stability; see Chernozhukov et al. (2018) and Newey and Robins (2018).
  • Higher-order orthogonality is the closest methodological comparison, including Mackey et al. (2018) and Jin et al. (2025).
  • Transform-based estimation and zero arguments connect to Feuerverger and McDunnough (1981), Carrasco (2017), Kagan et al. (1973), Mattner (1992), and Hu and Shiu (2022).
  • Non-Gaussian identification also appears in ICA and structural-equation work, including Comon (1994), Lanne et al. (2017), Lee and Mesters (2024), and Reizinger et al. (2025).

ACE Comparison

  • Jin et al. (2025) provide the published comparison through their finite-order ACE estimator for cumulant-separated partially linear models.
  • The ACE guarantee uses Lr(PX)L^r(P_X) treatment and outcome code radii and the JMS eligibility condition.
  • The contour guarantee uses the common clipped-code experiment and the L1(PX)L^1(P_X) treatment-code gate.
  • On the aligned class, ACE carries finite-order nuisance terms, while the contour upper guarantee is governed by (γn)1/2(\gamma n)^{-1/2}.
  • When the ACE nuisance terms dominate n1/2n^{-1/2}, the ratio of the contour upper guarantee to the ACE upper guarantee tends to zero.
Theorem T-4 (ACE and spectral alignment)

Fix an ACE order r2r\ge2 and positive constants δ,Cθ,Cg,Cq,ψη,ψξ\delta,C_{\theta},C_g,C_q,\psi_{\eta},\psi_{\xi}. There is a constant Cspec>0C_{\mathrm{spec}}>0, depending only on these fixed constants, such that the following statements hold. Let γ(1/2,1)\gamma\in(1/2,1), let Cγ>0C_{\gamma}>0 be the constant in the published order-rr ACE generalized-quantile guarantee, and write θ^ACE,r,n\widehat\theta_{\mathrm{ACE},r,n} for the estimator supplied by that published ACE handle. Let p0p_0 be a parameter record whose r,δ,Cθ,Cg,Cq,ψη,ψξr,\delta,C_{\theta},C_g,C_q,\psi_{\eta},\psi_{\xi}, and γ\gamma coordinates equal the displayed constants, and fix the certified primitive bank and range records associated with p0p_0. For any parameter record pp sharing the same fixed experiment constants as p0p_0 and the same probability level γ\gamma, and for any supplied code sequences whose clipped versions are gˉn\bar g_n and qˉn\bar q_n, the following hold, under the usual measurability and integrability conditions.

  • Class relations. At the sample size nn of pp, PACE,nJMSPNG,n,PNG,nACEPACE,nJMS. \mathcal P_{\mathrm{ACE},n}^{\mathrm{JMS}} \subseteq \mathcal P_{\mathrm{NG},n}, \qquad \mathcal P_{\mathrm{NG},n}^{\mathrm{ACE}} \subseteq \mathcal P_{\mathrm{ACE},n}^{\mathrm{JMS}}. If the comparison exponent of pp satisfies s=rs=r, then the common clipped-code convention gives PACE,nJMS=PNG,nACE. \mathcal P_{\mathrm{ACE},n}^{\mathrm{JMS}} =\mathcal P_{\mathrm{NG},n}^{\mathrm{ACE}}. Here the classes and the comparison subclass are those of Definition P-3, Definition P-1, Definition P-10.
  • ACE upper guarantee. If a law PPACE,nJMSP\in\mathcal P_{\mathrm{ACE},n}^{\mathrm{JMS}} uses the clipped supplied treatment and outcome codes gˉn,qˉn\bar g_n,\bar q_n, if the JMS eligibility condition EnJMS\mathsf E_n^{\mathrm{JMS}} of Definition P-7, Definition P-17 holds, and if ε1,n>0\varepsilon_{1,n}>0 and ε2,n0\varepsilon_{2,n}\ge0, then QP,1γ(θ^ACE,r,nθ0(P))BACE,n,γ, Q_{P,1-\gamma} \bigl(|\widehat\theta_{\mathrm{ACE},r,n}-\theta_0(P)|\bigr) \le B_{\mathrm{ACE},n,\gamma}, where QP,1γQ_{P,1-\gamma} is the generalized quantile in Definition P-15 and BACE,n,γ=Cγr!16rδ1[ε1,nrε2,n+Cθε1,nr+1+64(Cg+ψη)r{r2(Cg+ψη)+ψξ+Cθψη}(γn)1/2]. \begin{aligned} B_{\mathrm{ACE},n,\gamma} &= C_{\gamma} r!16^r\delta^{-1} \Big[ \varepsilon_{1,n}^r\varepsilon_{2,n} + C_{\theta}\varepsilon_{1,n}^{r+1} \\ &\qquad\qquad +64(C_g+\psi_{\eta})^r \{r^2(C_g+\psi_{\eta})+\psi_{\xi}+C_{\theta}\psi_{\eta}\} (\gamma n)^{-1/2} \Big]. \end{aligned}
  • Spectral upper guarantee. If PACE,nJMS\mathcal P_{\mathrm{ACE},n}^{\mathrm{JMS}} contains at least one law using the clipped supplied codes, n2n\ge2, EnJMS\mathsf E_n^{\mathrm{JMS}} holds, ε1,n>0\varepsilon_{1,n}>0, and ε1,n{4R1exp(2CgR1)}1, \varepsilon_{1,n} \le \{4R_1\exp(2C_gR_1)\}^{-1}, where R1R_1 is the outer contour radius of Definition P-12 determined by the parameter record in force, then the certified contour statistic θ^spec,n\widehat\theta_{\mathrm{spec},n} of Definition P-13, built from the fixed primitive records, satisfies supPPACE,nJMSQP,1γ(θ^spec,nθ0(P))Cspecγn, \sup_{P\in\mathcal P_{\mathrm{ACE},n}^{\mathrm{JMS}}} Q_{P,1-\gamma} \bigl(|\widehat\theta_{\mathrm{spec},n}-\theta_0(P)|\bigr) \le \sqrt{\frac{C_{\mathrm{spec}}}{\gamma n}}, where the supremum ranges over laws in the published ACE class whose clipped supplied treatment and outcome codes are gˉn\bar g_n and qˉn\bar q_n.
  • Upper-guarantee separation along sequences. For any sequence (pn)(p_n) of parameter records whose sample size at index nn is nn, with the same fixed experiment constants as p0p_0, the same radius sequences ε1,,ε2,\varepsilon_{1,\cdot},\varepsilon_{2,\cdot}, and pn.γ=γp_n.\gamma=\gamma, suppose that eventually n2n\ge2, EnJMS\mathsf E_n^{\mathrm{JMS}} holds, ε1,n>0\varepsilon_{1,n}>0, ε2,n0\varepsilon_{2,n}\ge0, ε1,n{4R1exp(2CgR1)}1, \varepsilon_{1,n} \le \{4R_1\exp(2C_gR_1)\}^{-1}, and the published ACE class contains a law using the clipped supplied codes. If ε1,nrε2,n+Cθε1,nr+1n1/2, \frac{ \varepsilon_{1,n}^r\varepsilon_{2,n} + C_{\theta}\varepsilon_{1,n}^{r+1} }{n^{-1/2}} \longrightarrow \infty, then Cspec/(γn)BACE,n,γ0. \frac{ \sqrt{C_{\mathrm{spec}}/(\gamma n)} }{ B_{\mathrm{ACE},n,\gamma} } \longrightarrow 0 .

Gaussian Benchmark

  • The JMS Gaussian comparison imposes Gaussian treatment noise and bounded observed outcomes.
  • Under those simultaneous restrictions, the target coefficient is forced to zero.
  • The fixed-code Gaussian minimax mean-squared risk and generalized-quantile risk are therefore zero whenever the clipped-code intersection is nonempty.

informal · Theorem T-8 In the bounded-outcome Gaussian JMS comparison class, every admissible law has θ0(P)=0\theta_0(P)=0, and the fixed-code Gaussian risks are zero when the class intersection is nonempty.

Explicit Mixtures

  • A symmetric two-component Gaussian mixture has a known imaginary transform zero.
  • The contour construction collapses to a sine-ratio estimator using Ysin(πZn/2)Y\sin(\pi Z_n/2) and Znsin(πZn/2)Z_n\sin(\pi Z_n/2).
  • The L1(PX)L^1(P_X) treatment-code bound keeps the sine denominator bounded away from zero.
  • The resulting clipped ratio attains mean-squared error at most C/nC/n.

informal · Theorem T-6 For the symmetric Gaussian mixture treatment innovation with the stated independence, range, tail, sampling, and treatment-code conditions, the clipped sine-ratio estimator has mean-squared error at most C/nC/n.

Local Benchmarks

  • The Gaussian--Rademacher path makes weak non-Gaussianity explicit through the amplitude aa.
  • Its fourth cumulant magnitude is Δa=2a4\Delta_a=2a^4, and the first positive transform zero determines the sine frequency.
  • The mean-squared error bound grows with the reciprocal denominator scale as the path approaches Gaussian noise.
  • The local ACE oracle envelope records the published ACE upper bound with shrinking cumulant threshold δn\delta_n.
Theorem T-7 (Local Gaussian benchmarks)

Let (δn)nN(\delta_n)_{n\in\mathbb N} be a real sequence with δn>0\delta_n>0 for every nn, nonincreasing in nn, and δn0\delta_n\to0. Fix Cθ,Cg,Cq,ψξRC_{\theta},C_g,C_q,\psi_{\xi}\in\mathbb R. The following two assertions hold.

  • (Gaussian--Rademacher path.) There is a constant C>0C>0, depending only on Cθ,Cg,Cq,ψξC_{\theta},C_g,C_q,\psi_{\xi}, such that the following holds. Let pp be any parameter record whose constants Cθ,Cg,Cq,ψξC_{\theta},C_g,C_q,\psi_{\xi} are the ones just fixed and whose orders are k=4k=4 and r=3r=3. For any a(0,1]a\in(0,1] and any law PP in a model mm, assume that the treatment noise satisfies η1a2G+aS, \eta\sim \sqrt{1-a^2}\,G_{\circ}+aS, where GN(0,1)G_{\circ}\sim N(0,1), SS is symmetric Rademacher, and GSG_{\circ}\perp S. Assume also Assumption A-1, Assumption A-2, Assumption A-3, Assumption A-4, Assumption A-5, Assumption A-6, Assumption A-9. At the sample size nn of pp, assume the supplied clipped treatment code gˉn\bar g_n satisfies xgˉn(x)g0(x)L1(PX),gˉn(x)g0(x)dPX(x)ε1,n. x\mapsto |\bar g_n(x)-g_0(x)|\in L^1(P_X), \qquad \int |\bar g_n(x)-g_0(x)|\,dP_X(x)\le \varepsilon_{1,n}. Then the conclusion of Lemma L-5 — the explicit mean-squared-error bound along the Gaussian--Rademacher path — holds with constant CC for p,m,ap,m,a.
  • (Local ACE oracle envelope.) For every published ACE handle, every γR\gamma\in\mathbb R, and every CγRC_{\gamma}\in\mathbb R, if the handle satisfies the conclusion of Theorem~5.4 of \citet{JinMackeySyrgkanis2025} at (γ,Cγ)(\gamma,C_{\gamma}), then Cγ>0C_{\gamma}>0 and the following local ACE statement holds. For every parameter record pp whose probability level is γ\gamma, writing nn for its sample size, if ε1,n>0\varepsilon_{1,n}>0 and ε2,n>0\varepsilon_{2,n}>0, then for every deterministic treatment code g^\widehat g, deterministic outcome code q^\widehat q, and every model, whenever gˉn=Π[Cg,Cg]g^n,qˉn=Π[Cq,Cq]q^n, \bar g_n=\Pi_{[-C_g,C_g]}\widehat g_n, \qquad \bar q_n=\Pi_{[-C_q,C_q]}\widehat q_n, and, writing rr and kk for the orders of pp and Δ=δn\Delta=\delta_n, the local ACE class conditions 1n,κk(η)Δ, 1\le n, \qquad |\kappa_k(\eta)|\ge \Delta, hold together with Assumption A-2, Assumption A-3, Assumption A-4, Assumption A-5, Assumption A-6, Assumption A-9, the treatment-noise exponential envelope oexp{η(o)2/ψη2}L1(P),exp{η(o)2/ψη2}dP(o)2, o\mapsto \exp\{\eta(o)^2/\psi_{\eta}^2\}\in L^1(P), \qquad \int \exp\{\eta(o)^2/\psi_{\eta}^2\}\,dP(o)\le 2, and the nuisance-radius bounds gˉng0Lr(PX)ε1,n,qˉnq0Lr(PX)ε2,n, \|\bar g_n-g_0\|_{L^r(P_X)}\le \varepsilon_{1,n}, \qquad \|\bar q_n-q_0\|_{L^r(P_X)}\le \varepsilon_{2,n}, and whenever the local JMS eligibility quantities a1,n=2log ⁣(6(Cg+ψη)ε1,n1),b1,n=log(γn/9), a_{1,n}=2\log\!\left(6(C_g+\psi_{\eta})\varepsilon_{1,n}^{-1}\right), \qquad b_{1,n}=\log(\gamma n/9), a2=4(Cg+ψη),b2,n(Δ)=200min{1,Cθ}Δmax{ε1,n,ε2,n,(γn)1/2(ψξ+Cθψη)} a_2=4(C_g+\psi_{\eta}), \qquad b_{2,n}(\Delta)= \frac{200\min\{1,C_{\theta}\}\Delta}{\max\{\varepsilon_{1,n},\varepsilon_{2,n},(\gamma n)^{-1/2}(\psi_{\xi}+C_{\theta}\psi_{\eta})\}} satisfy a1,n0,0<b1,na1,n,0<a2b2,n(Δ),a2log{a2b2,n(Δ)}0, a_{1,n}\ne0, \qquad 0<\frac{b_{1,n}}{a_{1,n}}, \qquad 0<a_2b_{2,n}(\Delta), \qquad a_2\log\{a_2b_{2,n}(\Delta)\}\ne0, and rmin{b1,na1,na1,n1log ⁣(b1,na1,n),b2,n(Δ)a2log{a2b2,n(Δ)}}, r\le \min\left\{ \frac{b_{1,n}}{a_{1,n}}-a_{1,n}^{-1}\log\!\left(\frac{b_{1,n}}{a_{1,n}}\right), \frac{b_{2,n}(\Delta)}{a_2\log\{a_2b_{2,n}(\Delta)\}} \right\}, the published order-rr ACE estimator obeys QP,1γ ⁣(θ^ACE,r,nθ0(P))BACE,n,γ(δn;Cγ), Q_{P,1-\gamma} \!\left(\left|\widehat\theta_{\mathrm{ACE},r,n}-\theta_0(P)\right|\right) \le B_{\mathrm{ACE},n,\gamma}(\delta_n;C_{\gamma}), where BACE,n,γ(Δ;Cγ)=Cγr!16rΔ1(ε1,nrε2,n+Cθε1,nr+1+64(Cg+ψη)r(r2(Cg+ψη)+ψξ+Cθψη)(γn)1/2). B_{\mathrm{ACE},n,\gamma}(\Delta;C_{\gamma}) = C_{\gamma}\,r!\,16^r\,\Delta^{-1} \left( \varepsilon_{1,n}^r\varepsilon_{2,n} +C_{\theta}\varepsilon_{1,n}^{r+1} +64(C_g+\psi_{\eta})^r \bigl(r^2(C_g+\psi_{\eta})+\psi_{\xi}+C_{\theta}\psi_{\eta}\bigr) (\gamma n)^{-1/2} \right).

Also in the Paper

informal · Lemma L-8 The JMS ACE class is contained in the non-Gaussian contour class, the Ls(PX)L^s(P_X)-restricted contour subclass is contained in the JMS ACE class, and the two coincide when s=rs=r.

informal · Lemma L-4 A one-dimensional fixed-code submodel gives an n1n^{-1} minimax lower bound on the non-Gaussian and ACE classes.

informal · Lemma L-5 Along the Gaussian--Rademacher path, the sine denominator and cumulant scale are explicit, and the clipped sine estimator satisfies the displayed mean-squared-error bound.

Future Work

  • The fixed-separation theorem holds with δ\delta, the cumulant-separation level, fixed across the experiment.
  • A local-to-Gaussian frontier lets the cumulant threshold δn\delta_n shrink with nn.
  • The open statistical object is the sharp minimax mean-squared-error rate as a function of (n,ε1,n,ε2,n,δn)(n,\varepsilon_{1,n},\varepsilon_{2,n},\delta_n).
  • A full selector would compare ordinary debiased machine learning, the finite-order ACE estimator, and global-contour procedures under a uniform inference criterion.
Definition P-14 (Local-to-Gaussian frontier)

A natural next question concerns triangular treatment-noise laws satisfying κk(ηn)δnwithδn0 |\kappa_k(\eta_n)|\ge\delta_n \qquad\text{with}\qquad \delta_n\downarrow0 under the same supplied code sequences. The target is a data-driven selector among ordinary DML, finite-order ACE, and global-contour procedures that attains the sharp minimax mean-squared-error rate as a function of (n,ε1,n,ε2,n,δn), (n,\varepsilon_{1,n},\varepsilon_{2,n},\delta_n), supports uniformly valid inference across these regimes, and admits a matching local minimax lower bound. Future work could determine the sharp-rate functional, the corresponding uniform-inference criterion, and the local lower-bound construction for this frontier.

Conclusion

  • We construct contour instruments from zeros of the treatment-innovation transform.
  • The population contour ratio identifies the partially linear treatment coefficient under zero-free boundary, zero-free nuisance, and positive-count conditions.
  • The finite contour statistic attains fixed-code minimax mean-squared error of order n1n^{-1} under fixed cumulant separation and the stated stability and boundedness conditions.
  • The ACE alignment places the contour guarantee and the published finite-order guarantee on a common clipped-code comparison class.
  • The mixture benchmarks show the same mechanism in explicit sine-ratio form.