CausalSmith · seminar slides
Contour Instruments for Partially Linear Models
We use zeros of the treatment-innovation transform to build contour instruments that identify and estimate the partially linear treatment coefficient at fixed-separation root-n scale.
slides for Contour Instruments for Partially Linear Models with Cumulant-separated Treatment Noise
Overview
- We study a partially linear model with a supplied treatment code.
- The target is θ0, the treatment coefficient after residualizing treatment on covariates.
- The key resource is non-Gaussian treatment noise with a fixed nonzero cumulant.
- A zero of the treatment-innovation moment-generating function creates an instrument that survives treatment-code error.
- A finite contour bank turns that population identity into a fixed-code estimator with minimax mean-squared error of order n−1.
informal · Theorem T-3 Under fixed cumulant separation and the stated fixed-code stability and boundedness conditions, the contour statistic has minimax mean-squared error between c/n and C/n, and its generalized quantile error is at most C/(γn).
Motivation
- Partially linear models estimate a low-dimensional treatment effect while allowing flexible covariate adjustment.
- Robinson (1988) gives the classical root-n residualized construction.
- Modern double/debiased machine learning keeps the same target stable under learned nuisance functions; see Chernozhukov et al. (2018).
- Here the supplied treatment code may have residual error.
- Think of T as an endogenous exposure whose predictable part is learned from covariates.
- We observe Zn=T−gˉn(X), the residualized treatment formed with the supplied clipped treatment code.
- The code error shifts the true treatment innovation by Dn(X)=g0(X)−gˉn(X).
- Non-Gaussian treatment noise gives complex transform zeros that can protect estimation from these real shifts.
Setup
- We observe O=(X,T,Y).
- The conditional treatment and outcome regressions are g0(X)=E[T∣X] and q0(X)=E[Y∣X].
- The treatment innovation is η=T−g0(X).
- The outcome innovation is ξ=Y−q0(X)−θ0η.
- The supplied treatment and outcome codes are clipped to fixed ranges before use.
- The sample is i.i.d. and split into a pilot fold and an evaluation fold.
An observation is the triple O=(X,T,Y), where X takes values in a measurable covariate space X and T,Y∈R. A model over a parameter record consists of a probability law P for O under which T and Y are integrable, a real coefficient θ0, measurable regression functions g0,q0 on the covariate space with g0(X)=EP[T∣X],q0(X)=EP[Y∣X] P-almost surely, and deterministic measurable supplied code sequences g=(gn)n∈N and q=(qn)n∈N for g0 and q0. The treatment and outcome innovations are η=T−g0(X),ξ=Y−q0(X)−θ0η, and the clipped supplied codes at index n are gˉn(x)=min{max{gn(x),−Cg},Cg},qˉn(x)=min{max{qn(x),−Cq},Cq}. The covariate marginal of P is written PX. A model is written m, and when the law has to be displayed the coefficient is written θ0(P).
Assumptions
- The treatment innovation is independent of covariates.
- The outcome innovation has conditionally mean zero given X,T.
- The coefficient and regression functions are bounded.
- Treatment and outcome innovations are sub-Gaussian under our fixed envelope convention.
- The kth cumulant of the treatment innovation is separated from zero by δ, the cumulant-separation level.
- The treatment-code radius ε1,n controls ∥gˉn−g0∥L1(PX).
In the setting of Definition P-19, Definition P-18, the kth cumulant of the treatment innovation η satisfies ∣κk(η)∣≥δ.
Key Idea
- Let M(z)=E[ezη] be the treatment-innovation moment-generating function.
- If M has a zero z0, a polynomial-exponential weight built from z0 annihilates every real shift of η.
- Conditioning on X, that same weight becomes an instrument for the code residual Zn.
- The resulting ratio identifies θ0 whenever the denominator is separated from zero.
informal · Theorem T-1 A zero of the treatment-innovation transform yields a valid shifted-residual instrument and identifies θ0 through a population ratio when the denominator is nonzero.
Contour Identification
- A known zero gives one instrument; a contour handles the zeros without naming them one by one.
- Fn is the observable transform of the residualized treatment.
- Gn is the observable outcome-weighted transform.
- A contour enclosing at least one zero averages Gn/Fn over the enclosed zero structure.
- Nuisance zero-freeness keeps the treatment-code error from creating spurious denominator zeros.
informal · Theorem T-2 On a contour where Fn is zero-free at the boundary, the nuisance factor is zero-free inside, and at least one zero is enclosed, the normalized contour ratio equals θ0(P).
Estimator
- The estimator uses a finite translated-dyadic bank of circles between the zero-localization radius and an outer radius.
- The pilot fold selects a circle with a positive winding count and a certified denominator margin.
- The evaluation fold computes the normalized contour average on the selected circle.
- The real midpoint is clipped to the known coefficient range.
- Every empirical branch returns a value, so the estimator is always well defined.
The estimator θspec,n is the packaged adaptive contour estimator on the public domain where the current supplied treatment-regression code underlying gˉn is measurable. It takes as inputs O1,…,On, the supplied treatment-regression code sequence through its current clipped values gˉn, the parameter record p, and fixed certified records p⋆ and cθ,⋆ representing k,δ,ψη,R0,R1, and Cθ, with R0=Ak(δψη2)(k−2)−1,R1=R0+1. Here p contains n≥1, r,k∈N with k=r+1 and r≥2, s∈[0,∞] with r≤s, γ∈(1/2,1), positive constants Cθ,Cg,Cq,ψη,ψξ,δ,σ, and nonnegative nonincreasing sequences ε1n,ε2n:N→R on positive indices.
- Define the deterministic folds I0={i:i<⌊n/2⌋},I1={i:i≥⌊n/2⌋}. For each real observation coordinate x and integer m≥0, define the floor-dyadic interval Im(x)=[2−(m+1)⌊2m+1x⌋,2−(m+1){⌊2m+1x⌋+1}].
- Form the residualized treatment values Zn,i=Ti−gˉn(Xi),i=1,…,n. For each fold a∈{0,1}, derivative order p, and complex z, define Fa,n(p)(z)=(max{∣Ia∣,1})−1i∈Ia∑Zn,ipezZn,i,p∈{0,1,2}, and Ga,n(p)(z)=(max{∣Ia∣,1})−1i∈Ia∑YiZn,ipezZn,i,p∈{0,1}. Thus an empty fold contributes the empty sum and is normalized by one.
- For represented observation data, define s=((ti,yi,gi)i=1n), where ti, yi, and gi are certified-real records naming Ti, Yi, and gˉn(Xi). The full represented-data input is s⋆=((ti,yi,gi)i=1n,p⋆,cθ,⋆). For i=1,…,n, put zi=ti−gi.
- Invoke the contour bank determined by p⋆. For the jth circle, define Cj={zj(t):0≤t≤1},zj(t)=ρje2πit. At error tolerance one, refine zi, yi, and ρj using their own executable moduli, and let UZ,i, UY,i, and Uρ,j be the largest absolute endpoint of the corresponding rational intervals. For a canonical observation name this is the interval Iμ(1) with the convention above; for zi it is the certified subtraction interval obtained from ti and gi, whose modulus uses the two half-error input moduli. Using ex≤3⌈x⌉ for x≥0 and π<4, form rational upper bounds for the empirical derivative sums on each Cj.
- For each j, compute a rational bracket of width at most a⋆/64 enclosing z∈Cjinf∣F0,n(z)∣. Let mj be the lower endpoint of this bracket. When mj>0, define W0j(t)=ρje2πitF0,n(zj(t))F0,n′(zj(t)). Use the rational Lipschitz bound ∣W0j′(t)∣≤64[ρjmjsup∣F0,n′∣+ρj2{mjsup∣F0,n′′∣+mj2sup∣F0,n′∣2}].
- Compute a rectangle enclosure, with coordinate widths strictly below 1/4, for ∫01W0j(t)dt=2πi1∮CjF0,n(z)F0,n′(z)dz. Define Nj=2πi1∮CjF0,n(z)F0,n′(z)dz when the real coordinate contains a unique nonnegative integer and the imaginary coordinate contains zero. A circle is admissible when mj≥a⋆/2andNj≥1. Let j⋆=min{j:j∈argℓ:mℓ≥a⋆/2, Nℓ≥1maxmℓ}. If the admissible set is empty, set θspec,n=0.
- On the evaluation fold for the selected circle Cj⋆, compute a lower endpoint mj⋆(1)≤z∈Cj⋆inf∣F1,n(z)∣. If mj⋆(1)<a⋆/4, set θspec,n=0.
- Otherwise define the evaluation-fold contour integrand Vj⋆(t)=ρj⋆e2πitF1,n(zj⋆(t))G1,n(zj⋆(t)). Use the rational Lipschitz bound ∣Vj⋆′(t)∣≤64[ρj⋆mj⋆(1)sup∣G1,n∣+ρj⋆2{mj⋆(1)sup∣G1,n′∣+(mj⋆(1))2sup∣G1,n∣sup∣F1,n′∣}]. For each rational tolerance e>0 and rational Lipschitz bound L, the schedule spends one third of the tolerance on discretization and uses Nmesh=⌈e3L⌉+1, so that L/Nmesh≤e/3.
- Enclose the unnormalized contour integral ∫01Vj⋆(t)dt=2πi1∮Cj⋆F1,n(z)G1,n(z)dz. Divide the resulting rectangle by Nj⋆ to enclose Nj⋆2πi1∮Cj⋆F1,n(z)G1,n(z)dz with coordinate radius at most 1/n. Take y∈Q to be the real midpoint of that enclosure, so that y approximates the normalized contour integral to within 1/n: y−Re(Nj⋆2πi1∮Cj⋆F1,n(z)G1,n(z)dz)≤n1. Only rational quantities are ever formed, so the estimator returns this certified rational approximation rather than the integral itself; the 1/n tolerance is negligible beside the statistical error.
- Output θspec,n=clipCθ(y)=2∣y+Cθ∣−∣y−Cθ∣=Π[−Cθ,Cθ](y).
Main Result
- Fixed cumulant separation localizes a transform zero in a fixed search region.
- The L1(PX) treatment-code gate preserves the zero geometry of the residual transform.
- Uniform empirical transform control gives root-n accuracy on the selected contour.
- The lower bound comes from a one-dimensional submodel with fixed treatment law and 1/n target separation.
Fix r≥2 and positive constants δ,Cθ,Cg,Cq,ψη,ψξ. Then there are constants c>0 and C>0, depending only on these fixed primitive constants, such that the following statements hold. For every measurable covariate space, let p0 be a parameter record whose primitive entries r,δ,Cθ,Cg,Cq,ψη,ψξ equal the fixed constants above, and fix one experiment-wide record consisting of the certified contour-bank input and the represented range input for p0. For every parameter record p with the same primitive entries as p0, and for every base model satisfying the usual measurability, conditional-mean, and integrability requirements, write n for the sample size of p. Let θspec,n be the total Borel statistic of Definition P-13, constructed from the transported fixed records and the base treatment-code sequence. Define the fixed-code non-Gaussian class PNG,n(p;base)={P: P∈PNG,n of def:non-gaussian-class, gˉn(P)=gˉn(base), qˉn(P)=qˉn(base)}. Define the fixed-code JMS ACE class PACE,nJMS(p;base)={P: P∈PACE,nJMS of def:jms-ace-class, gˉn(P)=gˉn(base), qˉn(P)=qˉn(base)}.
- (Measurable statistic.) The statistic θspec,n is measurable.
- (Represented execution.) For every compiled bounded spectral adapter satisfying the full canonical build-and-compilation specification for p, the transported fixed records, and the base treatment-code sequence, the represented-data execution realizes the same θspec,n as in Definition P-13.
- (Class relations.) Every law in PACE,nJMS belongs to PNG,n. If s=r, then PACE,nJMS coincides with the Ls(PX)-restricted subclass PNG,nACE of Definition P-10.
- (Non-Gaussian fixed-code risk.) If the sample law is the i.i.d. product law of Assumption A-1, PNG,n(p;base) is nonempty, n≥2, and ε1,n≤{4R1exp(2CgR1)}−1, where R1 is the search radius from the contour-bank construction in Definition P-13, then nc≤RnNG≤P∈PNG,n(p;base)supEP[(θspec,n−θ0(P))2]≤nC, and P∈PNG,n(p;base)supQP,1−γ(∣θspec,n−θ0(P)∣)≤γnC, with generalized lower quantiles as in Definition P-15.
- (ACE fixed-code risk.) If PACE,nJMS(p;base) is nonempty, n≥2, and ε1,n≤{4R1exp(2CgR1)}−1, then nc≤θinfP∈PACE,nJMS(p;base)supEP[(θ−θ0(P))2]≤P∈PACE,nJMS(p;base)supEP[(θspec,n−θ0(P))2]≤nC, where the infimum is over measurable real-valued estimators of O1,…,On. Moreover, P∈PACE,nJMS(p;base)supQP,1−γ(∣θspec,n−θ0(P)∣)≤γnC.
Sequence Result
- Along a sequence of experiments with the same primitive constants, one contour bank serves the sequence.
- On the cumulant-separated non-Gaussian fixed-code class, the minimax mean-squared risk stays between constants times nj−1.
- On the bounded-outcome Gaussian comparison class of Jin, Mackey, and Syrgkanis (JMS), the target degenerates to zero under the stated simultaneous restrictions.
informal · Theorem T-5 Along fixed-constant non-Gaussian experiment sequences, the same contour bank gives minimax mean-squared risk of order nj−1, while the bounded-outcome Gaussian JMS fixed-code risks are zero.
Related Literature
- Partially linear models trace to Engle et al. (1986), Robinson (1988), Speckman (1988), and Härdle et al. (2000).
- Orthogonal and debiased learning build on semiparametric stability; see Chernozhukov et al. (2018) and Newey and Robins (2018).
- Higher-order orthogonality is the closest methodological comparison, including Mackey et al. (2018) and Jin et al. (2025).
- Transform-based estimation and zero arguments connect to Feuerverger and McDunnough (1981), Carrasco (2017), Kagan et al. (1973), Mattner (1992), and Hu and Shiu (2022).
- Non-Gaussian identification also appears in ICA and structural-equation work, including Comon (1994), Lanne et al. (2017), Lee and Mesters (2024), and Reizinger et al. (2025).
ACE Comparison
- Jin et al. (2025) provide the published comparison through their finite-order ACE estimator for cumulant-separated partially linear models.
- The ACE guarantee uses Lr(PX) treatment and outcome code radii and the JMS eligibility condition.
- The contour guarantee uses the common clipped-code experiment and the L1(PX) treatment-code gate.
- On the aligned class, ACE carries finite-order nuisance terms, while the contour upper guarantee is governed by (γn)−1/2.
- When the ACE nuisance terms dominate n−1/2, the ratio of the contour upper guarantee to the ACE upper guarantee tends to zero.
Fix an ACE order r≥2 and positive constants δ,Cθ,Cg,Cq,ψη,ψξ. There is a constant Cspec>0, depending only on these fixed constants, such that the following statements hold. Let γ∈(1/2,1), let Cγ>0 be the constant in the published order-r ACE generalized-quantile guarantee, and write θACE,r,n for the estimator supplied by that published ACE handle. Let p0 be a parameter record whose r,δ,Cθ,Cg,Cq,ψη,ψξ, and γ coordinates equal the displayed constants, and fix the certified primitive bank and range records associated with p0. For any parameter record p sharing the same fixed experiment constants as p0 and the same probability level γ, and for any supplied code sequences whose clipped versions are gˉn and qˉn, the following hold, under the usual measurability and integrability conditions.
- Class relations. At the sample size n of p, PACE,nJMS⊆PNG,n,PNG,nACE⊆PACE,nJMS. If the comparison exponent of p satisfies s=r, then the common clipped-code convention gives PACE,nJMS=PNG,nACE. Here the classes and the comparison subclass are those of Definition P-3, Definition P-1, Definition P-10.
- ACE upper guarantee. If a law P∈PACE,nJMS uses the clipped supplied treatment and outcome codes gˉn,qˉn, if the JMS eligibility condition EnJMS of Definition P-7, Definition P-17 holds, and if ε1,n>0 and ε2,n≥0, then QP,1−γ(∣θACE,r,n−θ0(P)∣)≤BACE,n,γ, where QP,1−γ is the generalized quantile in Definition P-15 and BACE,n,γ=Cγr!16rδ−1[ε1,nrε2,n+Cθε1,nr+1+64(Cg+ψη)r{r2(Cg+ψη)+ψξ+Cθψη}(γn)−1/2].
- Spectral upper guarantee. If PACE,nJMS contains at least one law using the clipped supplied codes, n≥2, EnJMS holds, ε1,n>0, and ε1,n≤{4R1exp(2CgR1)}−1, where R1 is the outer contour radius of Definition P-12 determined by the parameter record in force, then the certified contour statistic θspec,n of Definition P-13, built from the fixed primitive records, satisfies P∈PACE,nJMSsupQP,1−γ(∣θspec,n−θ0(P)∣)≤γnCspec, where the supremum ranges over laws in the published ACE class whose clipped supplied treatment and outcome codes are gˉn and qˉn.
- Upper-guarantee separation along sequences. For any sequence (pn) of parameter records whose sample size at index n is n, with the same fixed experiment constants as p0, the same radius sequences ε1,⋅,ε2,⋅, and pn.γ=γ, suppose that eventually n≥2, EnJMS holds, ε1,n>0, ε2,n≥0, ε1,n≤{4R1exp(2CgR1)}−1, and the published ACE class contains a law using the clipped supplied codes. If n−1/2ε1,nrε2,n+Cθε1,nr+1⟶∞, then BACE,n,γCspec/(γn)⟶0.
Gaussian Benchmark
- The JMS Gaussian comparison imposes Gaussian treatment noise and bounded observed outcomes.
- Under those simultaneous restrictions, the target coefficient is forced to zero.
- The fixed-code Gaussian minimax mean-squared risk and generalized-quantile risk are therefore zero whenever the clipped-code intersection is nonempty.
informal · Theorem T-8 In the bounded-outcome Gaussian JMS comparison class, every admissible law has θ0(P)=0, and the fixed-code Gaussian risks are zero when the class intersection is nonempty.
Explicit Mixtures
- A symmetric two-component Gaussian mixture has a known imaginary transform zero.
- The contour construction collapses to a sine-ratio estimator using Ysin(πZn/2) and Znsin(πZn/2).
- The L1(PX) treatment-code bound keeps the sine denominator bounded away from zero.
- The resulting clipped ratio attains mean-squared error at most C/n.
informal · Theorem T-6 For the symmetric Gaussian mixture treatment innovation with the stated independence, range, tail, sampling, and treatment-code conditions, the clipped sine-ratio estimator has mean-squared error at most C/n.
Local Benchmarks
- The Gaussian--Rademacher path makes weak non-Gaussianity explicit through the amplitude a.
- Its fourth cumulant magnitude is Δa=2a4, and the first positive transform zero determines the sine frequency.
- The mean-squared error bound grows with the reciprocal denominator scale as the path approaches Gaussian noise.
- The local ACE oracle envelope records the published ACE upper bound with shrinking cumulant threshold δn.
Let (δn)n∈N be a real sequence with δn>0 for every n, nonincreasing in n, and δn→0. Fix Cθ,Cg,Cq,ψξ∈R. The following two assertions hold.
- (Gaussian--Rademacher path.) There is a constant C>0, depending only on Cθ,Cg,Cq,ψξ, such that the following holds. Let p be any parameter record whose constants Cθ,Cg,Cq,ψξ are the ones just fixed and whose orders are k=4 and r=3. For any a∈(0,1] and any law P in a model m, assume that the treatment noise satisfies η∼1−a2G∘+aS, where G∘∼N(0,1), S is symmetric Rademacher, and G∘⊥S. Assume also Assumption A-1, Assumption A-2, Assumption A-3, Assumption A-4, Assumption A-5, Assumption A-6, Assumption A-9. At the sample size n of p, assume the supplied clipped treatment code gˉn satisfies x↦∣gˉn(x)−g0(x)∣∈L1(PX),∫∣gˉn(x)−g0(x)∣dPX(x)≤ε1,n. Then the conclusion of Lemma L-5 — the explicit mean-squared-error bound along the Gaussian--Rademacher path — holds with constant C for p,m,a.
- (Local ACE oracle envelope.) For every published ACE handle, every γ∈R, and every Cγ∈R, if the handle satisfies the conclusion of Theorem~5.4 of \citet{JinMackeySyrgkanis2025} at (γ,Cγ), then Cγ>0 and the following local ACE statement holds. For every parameter record p whose probability level is γ, writing n for its sample size, if ε1,n>0 and ε2,n>0, then for every deterministic treatment code g, deterministic outcome code q, and every model, whenever gˉn=Π[−Cg,Cg]gn,qˉn=Π[−Cq,Cq]qn, and, writing r and k for the orders of p and Δ=δn, the local ACE class conditions 1≤n,∣κk(η)∣≥Δ, hold together with Assumption A-2, Assumption A-3, Assumption A-4, Assumption A-5, Assumption A-6, Assumption A-9, the treatment-noise exponential envelope o↦exp{η(o)2/ψη2}∈L1(P),∫exp{η(o)2/ψη2}dP(o)≤2, and the nuisance-radius bounds ∥gˉn−g0∥Lr(PX)≤ε1,n,∥qˉn−q0∥Lr(PX)≤ε2,n, and whenever the local JMS eligibility quantities a1,n=2log(6(Cg+ψη)ε1,n−1),b1,n=log(γn/9), a2=4(Cg+ψη),b2,n(Δ)=max{ε1,n,ε2,n,(γn)−1/2(ψξ+Cθψη)}200min{1,Cθ}Δ satisfy a1,n=0,0<a1,nb1,n,0<a2b2,n(Δ),a2log{a2b2,n(Δ)}=0, and r≤min{a1,nb1,n−a1,n−1log(a1,nb1,n),a2log{a2b2,n(Δ)}b2,n(Δ)}, the published order-r ACE estimator obeys QP,1−γ(θACE,r,n−θ0(P))≤BACE,n,γ(δn;Cγ), where BACE,n,γ(Δ;Cγ)=Cγr!16rΔ−1(ε1,nrε2,n+Cθε1,nr+1+64(Cg+ψη)r(r2(Cg+ψη)+ψξ+Cθψη)(γn)−1/2).
Also in the Paper
informal · Lemma L-8 The JMS ACE class is contained in the non-Gaussian contour class, the Ls(PX)-restricted contour subclass is contained in the JMS ACE class, and the two coincide when s=r.
informal · Lemma L-4 A one-dimensional fixed-code submodel gives an n−1 minimax lower bound on the non-Gaussian and ACE classes.
informal · Lemma L-5 Along the Gaussian--Rademacher path, the sine denominator and cumulant scale are explicit, and the clipped sine estimator satisfies the displayed mean-squared-error bound.
Future Work
- The fixed-separation theorem holds with δ, the cumulant-separation level, fixed across the experiment.
- A local-to-Gaussian frontier lets the cumulant threshold δn shrink with n.
- The open statistical object is the sharp minimax mean-squared-error rate as a function of (n,ε1,n,ε2,n,δn).
- A full selector would compare ordinary debiased machine learning, the finite-order ACE estimator, and global-contour procedures under a uniform inference criterion.
A natural next question concerns triangular treatment-noise laws satisfying ∣κk(ηn)∣≥δnwithδn↓0 under the same supplied code sequences. The target is a data-driven selector among ordinary DML, finite-order ACE, and global-contour procedures that attains the sharp minimax mean-squared-error rate as a function of (n,ε1,n,ε2,n,δn), supports uniformly valid inference across these regimes, and admits a matching local minimax lower bound. Future work could determine the sharp-rate functional, the corresponding uniform-inference criterion, and the local lower-bound construction for this frontier.
Conclusion
- We construct contour instruments from zeros of the treatment-innovation transform.
- The population contour ratio identifies the partially linear treatment coefficient under zero-free boundary, zero-free nuisance, and positive-count conditions.
- The finite contour statistic attains fixed-code minimax mean-squared error of order n−1 under fixed cumulant separation and the stated stability and boundedness conditions.
- The ACE alignment places the contour guarantee and the published finite-order guarantee on a common clipped-code comparison class.
- The mixture benchmarks show the same mechanism in explicit sine-ratio form.