Contour Instruments for Partially Linear Models with Cumulant-separated Treatment Noise
Abstract
This paper studies estimation of the partially linear coefficient , the treatment coefficient, when the treatment innovation is independent of covariates, sub-Gaussian, and separated from Gaussianity by a fixed nonzero cumulant. The identifying resource is a zero of the treatment-innovation moment-generating function: the associated polynomial-exponential weight annihilates real shifts induced by treatment-code error, and a contour average of observable residual transforms identifies under the stated boundary zero-freeness, nuisance zero-freeness, and positive-count conditions.
The statistical theorem is a fixed-code result. For fixed primitive constants, i.i.d. product sampling, the partially linear conditional-mean restrictions, bounded coefficient and regression ranges, sub-Gaussian treatment and outcome-noise envelopes, a nonempty fixed-code class, and the displayed current treatment-code radius gate, where is the covariate marginal, a finite translated-dyadic contour bank and a total Borel statistic attain matched and mean-squared-error bounds on the non-Gaussian spectral class. For , where is the probability level, the same statistic controls the generalized lower -quantile of absolute error by order .
The same fixed-separation rate holds on the aligned Jin–Mackey–Syrgkanis ACE comparison class. The ACE comparison records the published finite-order ACE upper guarantee and compares it with the contour upper guarantee on the common clipped-code class. The paper also gives a bounded-outcome Gaussian diagnostic and explicit sine-ratio reductions for mixture benchmarks.
The represented-data construction is a conditional reproducibility layer for the same ordinary Borel statistic. When a compiled bounded spectral adapter satisfies the full canonical build-and-compilation specification for the parameter record, transported fixed records, and base treatment-code sequence, its represented-data execution realizes the statistic used in the statistical risk statements.
Introduction
Partially linear models are a standard way to separate a low-dimensional structural coefficient from flexible covariate regressions. In the model studied here, the observed data are triples , the treatment regression is , the outcome regression is , and the target is the partially linear coefficient in the conditional mean Classical semiparametric theory and modern double/debiased machine learning estimate this coefficient by orthogonalizing treatment and outcome variation around nuisance regressions (Robinson, 1988; Newey, 1990; Newey, 1994; Newey et al., 1994; Bickel et al., 1993; van der Vaart, 1998; van der Vaart et al., 1996; Belloni et al., 2014; Chernozhukov et al., 2018). Higher-order orthogonality and ACE methods refine this idea by exploiting additional structure in nuisance errors and treatment residuals (Robins et al., 2008; Liu et al., 2017; Newey et al., 2018; Liu et al., 2016; Mackey et al., 2018; Jin et al., 2025). This paper develops a complementary fixed-separation construction based on non-Gaussian treatment innovations.
The statistical decision problem fixes the supplied clipped treatment and outcome codes at the current sample size. The risk bounds compare laws compatible with that same code pair; if the codes are obtained from an external training stage, the theorem applies after conditioning on the training output. The finite-arithmetic material supplies a conditional reproducibility clause for represented-data executions that satisfy the displayed build contract.
Four statements organize the paper. First, the contour identification theorem recovers from an observable population contour ratio under the displayed zero-free and positive-count conditions. Second, under the theorem’s fixed-code nonemptiness, sampling, and radius conditions, the fixed-code minimax theorem gives matched and mean-squared-error bounds on the broad non-Gaussian spectral class and the aligned Jin–Mackey–Syrgkanis ACE class. Under the same fixed-code hypotheses, the same contour statistic also has a generalized lower -quantile upper bound of order . Third, the ACE alignment proposition compares the published finite-order ACE upper guarantee with the contour upper guarantee on the common clipped-code class. Fourth, the bounded-outcome Gaussian diagnostic characterizes the comparison class by its implied target and fixed-code risks. Here denotes the covariate marginal, the ACE order, the cumulant-separation level, the probability level, and the Gaussian innovation variance.
The identifying resource is a zero of the treatment-innovation moment-generating function. Let denote the treatment innovation and let . When has a zero of multiplicity , the polynomial-exponential weight annihilates every shifted innovation . This shift invariance is useful for code residuals: if , where is the clipped supplied treatment code, then the same zero generates conditional moments for . Theorem 1 gives the resulting known-zero ratio, and Lemma 1 and Theorem 2 turn the ratio into an observable contour identity.
The contour identity is the paper’s population mechanism. The residual transform factors as , where is the nuisance transform. The outcome-weighted transform decomposes into a derivative term carrying and an analytic contamination term. On any circle where is zero-free on the boundary, is zero-free on the closed disk, and the enclosed zero count is positive, the normalized contour integral of equals . This places the argument in the transform-based identification tradition (Feuerverger et al., 1981; Singleton, 2001; Chacko et al., 2003; Carrasco et al., 2007; Carrasco, 2017; Kagan et al., 1973; Mattner, 1992; D’Haultfoeuille, 2011; Darolles et al., 2011; Hu et al., 2022).
Estimator in words.
The statistic splits the sample deterministically. The pilot fold scans a fixed finite bank of circles, keeps circles whose empirical residual transform has a positive boundary modulus and a positive decoded winding count, and chooses an admissible circle by the displayed lower-modulus rule with an index tie-break. The evaluation fold computes the normalized contour ratio of the outcome-weighted transform to the residual transform on that selected circle. Clipping to the fixed coefficient range makes the statistic total on every ordinary sample.
The main statistical result is fixed-separation and fixed-code. Under the non-Gaussian spectral class in Definition 3, the treatment innovation is independent of covariates, sub-Gaussian, and has a fixed nonzero st cumulant bounded below by . Lemma 3 converts this cumulant separation and the treatment-tail scale into an explicit disk containing a transform zero. Lemma 8 constructs a finite translated-dyadic bank of circles with a positive dyadic boundary-modulus certificate. The treatment-code condition in Assumption 16 keeps separated from zero on the search disk, as shown in Lemma 10. Empirical transform concentration in Lemma 9 then supplies the stochastic control needed for the selected contour ratio.
The resulting estimator is the total Borel statistic in Algorithm 1. Theorem 3 states that, for fixed primitive constants and sufficiently small current treatment-code radius, this statistic attains while the fixed-code non-Gaussian minimax risk is bounded below by . Thus the fixed-code mean-squared-error rate is . For , the same theorem controls the generalized lower -quantile of absolute error by order , and it gives the parallel fixed-code result on the Jin–Mackey–Syrgkanis ACE comparison class. Theorem 4 records the corresponding sequence-level statement for experiments sharing the same primitive constants.
The comparison with ACE is stated as a comparison of upper guarantees on a common cumulant-separated domain. Lemma 5 aligns the published ACE class with the broad spectral class and its -restricted subclass. Proposition 1 then records the published order- ACE generalized-quantile guarantee from Jin et al. (2025) and the spectral guarantee on the same clipped-code ACE class. Along the displayed nuisance-dominance sequences, the ratio of the contour upper guarantee to the imported finite-order ACE upper guarantee tends to zero. This is an upper-bound comparison on the common clipped-code class. The minimax lower bound used for the ACE class is class-level over all estimators on that fixed-code class, as in Theorem 3.
The Gaussian comparison characterizes the bounded-outcome Gaussian intersection in Definition 4. That class combines a nondegenerate Gaussian treatment innovation with bounded observed outcomes and the exact partially linear conditional mean. Proposition 2 shows that these restrictions force , and consequently the corresponding fixed-code Gaussian mean-squared and generalized-quantile minimax risks are zero whenever the class intersection is nonempty. This result is a diagnostic for that simultaneous set of source assumptions.
Two explicit one-dimensional benchmarks illustrate the contour mechanism. Proposition 3 treats the symmetric mixture , whose transform has a known imaginary zero and whose sine-ratio estimator attains risk under the stated range, tail, independence, and treatment-code conditions. Lemma 7 and Theorem 5 give a Gaussian–Rademacher path and a local ACE oracle envelope, making the dependence on a shrinking fourth-cumulant magnitude explicit. These benchmarks connect the analysis to non-Gaussian identification ideas in independent component analysis and structural econometrics (Comon, 1994; Hyvärinen et al., 1997; Hyvärinen et al., 2001; Lanne et al., 2017; Lee et al., 2024; Hoesch et al., 2024; Reizinger et al., 2025).
The computational statement is a conditional output-correspondence theorem. Definition 17 and Algorithm 1 use the same primitive records for the ordinary Borel statistic and for represented-data execution. For any compiled bounded spectral adapter satisfying the full canonical build-and-compilation specification for the parameter record, transported fixed records, and base treatment-code sequence, the represented-execution clause in Theorem 3 states that the represented-data execution realizes the same statistic. The appendix records proof and deferred-detail information for the paper’s formal statements.
The paper proceeds as follows. The setup section defines the partially linear experiment, supplied-code restrictions, law classes, quantile criterion, and minimax risks. The zero-instrument section proves the known-zero and contour-identification identities. The main fixed-separation section gives zero localization, the contour statistic, and the fixed-code minimax theorem. The comparison section aligns the spectral, ACE, and Gaussian benchmark classes. The mixture section gives explicit sine reductions and local benchmarks. The limitations and future-work section states the local-to-Gaussian adaptation agenda. The appendices collect analytic, empirical-process, lower-bound, certified-construction, verification, and deferred-proof details.
Related work
The partially linear model originates in semiparametric regression, where a parametric coefficient is estimated alongside a nonparametric function of covariates (Engle et al., 1986; Speckman, 1988); Robinson (1988) gave the root- consistent residualized estimator that remains the reference construction, Härdle et al. (1993) developed the associated specification testing, and Härdle et al. (2000) give the monograph treatment. Partially linear models are a central workhorse for semiparametric estimation of low-dimensional structural coefficients in the presence of high-dimensional or otherwise flexible nuisance functions. The classical formulation and efficiency theory connect the coefficient , the partially linear coefficient, to residualized treatment variation after removing the treatment regression and outcome regression (Hansen, 1982; Newey, 1990; Newey, 1994; Newey et al., 1994; Bickel et al., 1993; van der Vaart, 1998; van der Vaart et al., 1996). Modern double/debiased machine learning builds on the same orthogonalization principle by allowing flexible first-stage estimates while preserving first-order stability of the target moment (Belloni et al., 2014; Chernozhukov et al., 2018). Within that tradition, the present paper focuses on a partially linear coefficient estimated from supplied treatment and outcome codes, and studies the additional identifying information carried by non-Gaussian treatment innovations.
A closely related line of work studies higher-order orthogonality and influence-function expansions for nuisance-robust semiparametric estimation. Higher-order constructions refine the first-order orthogonal-score idea by exploiting additional smoothness or distributional structure in nuisance errors (Robins et al., 1994; van der Laan et al., 2006; Robins et al., 2008; Liu et al., 2017; Newey et al., 2018; Liu et al., 2016; Chernozhukov et al., 2022). For the partially linear model, Mackey et al. (2018) identify the role of non-Gaussian treatment residuals in higher-order orthogonal learning. Jin et al. (2025) give the closest finite-order comparison point: their published order- ACE guarantee, with ACE order , controls the error quantile, where is the probability level, by a constant multiple of where and are the treatment- and outcome-code radii, is the coefficient-range bound, and is the sample size. The explicit prefactor , with denoting the cumulant-separation level, makes the comparison sensitive to the finite order and to the amount of non-Gaussian separation.
The direct gate in the contour theorem plays a different operational role from the JMS pair of treatment and outcome radii. The contour gate controls the treatment-code perturbation of the residual transform and preserves the zero geometry; the ACE eligibility quantities in Definitions 8 and 10 also use the outcome-code radius and the order- finite-difference scale. On bounded covariate-regression ranges, an treatment radius at level gives the corresponding treatment control, so the common ACE class lies inside the contour class through Lemma 5. The contour result therefore asks the treatment code to meet the displayed small-radius gate, while the published ACE guarantee asks the treatment and outcome radii to meet the JMS eligibility inequalities.
| Feature | Contour construction | JMS ACE comparison |
|---|---|---|
| Input accuracy | Current treatment-code radius and fixed primitive constants | Current treatment and outcome radii with JMS eligibility |
| Analytic resource | A localized zero of the treatment-innovation transform and a selected contour ratio | A finite order- ACE expansion under cumulant separation |
| Risk criterion compared | Fixed-code MSE minimax bounds and a generalized-quantile upper bound for the contour statistic | Published generalized-quantile upper guarantee and class-level fixed-code MSE lower bound |
| Fixed-separation dependence | Constants depend on the fixed cumulant-separation level and primitive envelopes | The displayed imported bound carries the finite-order prefactor and dependence |
This translation gives concrete nuisance-rate regimes for the comparison in Proposition 1. If and , the ACE nuisance part dominates sampling when , subject to the JMS eligibility and contour small-radius gates; in that regime the ratio of the contour upper guarantee to the imported finite-order ACE upper guarantee tends to zero on the common clipped-code class. If , the ACE upper guarantee and the contour upper guarantee are both governed by the sampling term up to constants and eligibility factors. These comparisons concern the available upper bounds on the aligned class.
The contour approach developed here represents the same source of identifying variation through zeros of the treatment-innovation moment-generating function and converts those zeros into instruments and contour ratios. Transform methods have a long history in econometrics and statistics, including characteristic-function estimation, empirical transform methods, and moment-generating-function arguments (Feuerverger et al., 1981; Singleton, 2001; Chacko et al., 2003; Carrasco et al., 2007; Carrasco, 2017). The classical foundation for reading distributional structure off a transform is Marcinkiewicz’s theorem, that a characteristic function of the form with a polynomial forces and hence a Gaussian law (Marcinkiewicz, 1939). In this paper’s centered sub-Gaussian envelope, the fixed cumulant-separation condition yields the explicit zero-localization conclusion in Lemma 3: the treatment-innovation MGF has a zero inside the displayed disk. The zero-based step also connects to classical results on distributions, transforms, and completeness (Kagan et al., 1973; Mattner, 1992; D’Haultfoeuille, 2011; Darolles et al., 2011; Hu et al., 2022). In the present setting, those tools are applied to residualized treatment variation formed with a supplied treatment code, so the analytic question becomes stability of transform zeros under treatment-code perturbations.
Non-Gaussianity is also a recurring identifying resource in independent component analysis, structural shock recovery, and related econometric models (Comon, 1994; Hyvärinen et al., 1997; Hyvärinen et al., 2001; Lanne et al., 2017; Lee et al., 2024; Hoesch et al., 2024; Reizinger et al., 2025). Relatedly, Andrews (2017) exhibits families that are -complete or boundedly complete, the property that makes a vanishing transform functional informative under the relevant class. The construction here uses non-Gaussianity through localized transform-zero geometry: it localizes a transform zero and integrates the observable ratio on a selected circle around the enclosed zero structure. The paper applies that resource in a partially linear estimation problem: cumulant separation supplies a transform zero, and the contour construction converts the resulting analytic geometry into a coefficient estimator. This places the contribution between higher-order semiparametric learning and non-Gaussian identification, with the comparison to ACE made on the common cumulant-separated class where both sets of assumptions are imposed.
The fixed-separation statistical statement is expressed in minimax terms. Under fixed cumulant separation and stable supplied treatment codes, the contour estimator attains the parametric mean-squared-error order over the non-Gaussian spectral class, while the comparison with Jin et al. (2025) records the published ACE generalized-quantile upper guarantee on the aligned ACE class. The sequence comparison in Proposition 1 is an upper-guarantee comparison on the common clipped-code class: when the displayed ACE nuisance terms dominate , the ratio of the fixed-separation contour upper guarantee to the imported finite-order ACE upper guarantee tends to zero.
Finally, the finite-construction component draws on computable analysis and interval arithmetic. The relevant background treats real-number computation, constructive numerical representations, and certified interval enclosures (Weihrauch, 2000; Ko, 1991; Moore, 1966). Those tools provide the arithmetic substrate for the contour bank and represented-data statistic, while the statistical exposition keeps the identification and risk arguments in standard econometric form.
Setup and assumptions
We use generic constants and only in later bounds, where their dependence is stated locally. All norms are taken under the law displayed in the surrounding statement, and denotes integration with respect to the covariate marginal. The sample index is fixed within each finite- experiment, while class constants and code-radius sequences are carried by the parameter record. Asymptotic comparisons in later sections vary the sample size through these records, with the order pair, tail scales, cumulant separation, and range bounds held according to the relevant theorem hypotheses.
The basic experiment is a partially linear model with deterministic supplied codes for the treatment and outcome regressions. The first definition records the numerical constants, the split convention, and the exact sub-Gaussian norm convention used throughout.
Fixed-code decision problem
The supplied treatment and outcome codes are inputs to the current finite- decision problem. At sample size , the estimator uses the clipped pair , and the risks below compare laws whose current clipped code pair agrees with that same input. The data used in the risk bound are then the i.i.d. product sample from the law under evaluation, with the code pair treated as fixed for that experiment.
This convention is the one used by the minimax criteria in Definition 9 and by the fixed-code classes in Theorem 3. When the supplied codes come from an external training stage, the theorem applies through the deterministic code pair obtained after conditioning on that training output, or through any deterministic code sequence fixed before the product-law sample is drawn. The theorem’s code-radius assumptions are assumptions about that current clipped pair under the covariate marginal .
A parameter record fixes a sample size , an order pair with a comparison exponent with , a probability level , positive constants and two code-radius sequences that are nonnegative and nonincreasing on positive indices; their values at are written and .
Two conventions are used throughout: denotes the Gaussian law on with mean and variance , and denotes the empirical average of over the sample. The sample is split into the two deterministic folds written for .
For a real random variable and , the Luxemburg sub-Gaussian convention is that means , and its conditional form means almost surely.
⊢ LeanThe observation-level model follows the usual partially linear decomposition, with and denoting the treatment and outcome regressions and with supplied code sequences clipped to the fixed ranges in Definition 1.
An observation is the triple , where takes values in a measurable covariate space and . A model over a parameter record consists of a probability law for under which and are integrable, a real coefficient , measurable regression functions on the covariate space with -almost surely, and deterministic measurable supplied code sequences and for and . The treatment and outcome innovations are and the clipped supplied codes at index are The covariate marginal of is written . A model is written , and when the law has to be displayed the coefficient is written .
⊢ LeanThe sampling and structural restrictions are collected next. They separate the finite-sample product experiment, the partially linear conditional mean, the bounded ranges, and the tail and cumulant conditions that define the statistical classes.
Under Definitions 2 and 1, the observations satisfy
⊢ LeanThe i.i.d. sampling condition is the standard finite-sample product-law setup for orthogonal and semiparametric procedures (Chernozhukov et al., 2018). It fixes the probability space on which all estimators and generalized quantiles below are evaluated.
Under Definitions 2 and 1, the treatment innovation is independent of the covariates:
⊢ LeanIndependent treatment noise is the standard residual independence condition used in the cumulant-based partially linear comparison (Jin et al., 2025). In this paper it also makes the treatment-innovation transform a law-level object separate from the covariate distribution.
Under Definitions 2 and 1, the outcome innovation has conditionally mean zero: almost surely.
⊢ LeanOutcome mean independence is the standard partially linear conditional mean restriction (Robinson, 1988). It pins down the coefficient through the conditional mean of given and .
Under Definitions 2 and 1, the partially linear coefficient satisfies
⊢ LeanThe coefficient range is the standard bounded target range condition in the JMS comparison framework (Jin et al., 2025). It supplies a common compact parameter range for minimax risks and clipped estimators.
Under Definitions 2 and 1, the treatment regression satisfies
⊢ LeanThe bounded treatment regression condition is standard in the same comparison class (Jin et al., 2025). It also aligns the supplied treatment code with the clipping radius used in Definition 2.
Under Definitions 2 and 1, the outcome regression satisfies
⊢ LeanThe bounded outcome regression condition is the standard common sub-Gaussian PLM range condition (Jin et al., 2025). Together with the treatment and coefficient bounds, it fixes the envelope in which the supplied outcome code is compared.
In the setting of Definitions 2 and 1, the outcome satisfies almost surely.
⊢ LeanThe bounded observed-outcome restriction is the standard bounded- condition in the simultaneous Gaussian comparison experiment (Jin et al., 2025). It is used only for that Gaussian intersection class.
In the setting of Definitions 2 and 1, with the Luxemburg convention the treatment innovation satisfies
⊢ LeanThe treatment-tail condition is the standard uniform sub-Gaussian treatment-noise requirement under the exact Luxemburg convention (Jin et al., 2025). The convention matches the concentration and transform-envelope bounds used later; background sub-Gaussian facts follow the usual conventions in Vershynin (2018) and Wainwright (2019).
In the setting of Definitions 2 and 1, the outcome innovation satisfies almost surely.
⊢ LeanThe outcome-tail condition is the standard uniform conditional sub-Gaussian outcome-noise restriction (Jin et al., 2025). It provides the conditional exponential-square control needed for empirical outcome-weighted transforms.
In the setting of Definitions 2 and 1, the th cumulant of the treatment innovation satisfies
⊢ LeanCumulant separation is the standard Jin–Mackey–Syrgkanis nonzero cumulant condition with a uniform separation level (Jin et al., 2025). In the contour analysis, the same lower bound supplies the non-Gaussian signal that later yields transform zeros.
The supplied-code restrictions specify how the deterministic first-stage codes enter the statistical classes. The conditions give a common language for comparison with ACE, the versions match the published order- condition, the Gaussian condition defines the comparison experiment, and the condition is the stability requirement used by the contour class.
In the setting of Definitions 2 and 1, for every ,
⊢ LeanThe treatment-code radius is a standard structure-agnostic supplied-code uncertainty set (Jin et al., 2025). It measures the current-index treatment-code error under the covariate marginal.
In the setting of Definitions 2 and 1, for every ,
⊢ LeanThe matching outcome-code radius is the standard supplied-code uncertainty set for the outcome regression (Jin et al., 2025). It places the outcome code and treatment code on the same comparison scale.
In the notation of Definitions 2 and 1, for every ,
⊢ LeanThe treatment-code radius is the standard Theorem 5.4 treatment-code uncertainty set in the JMS ACE guarantee (Jin et al., 2025). It is stated separately because the published finite-order bound is indexed by the ACE order .
In the notation of Definitions 2 and 1, for every ,
⊢ LeanThe outcome-code radius is the corresponding standard Theorem 5.4 outcome-code uncertainty set (Jin et al., 2025). It is paired with Assumption 13 in the ACE comparison class.
In the notation of Definitions 2 and 1, the treatment innovation satisfies
⊢ LeanThe Gaussian treatment-noise condition is the standard known-variance Gaussian experiment used in Theorem 3.2 of Jin et al. (2025). It defines the Gaussian benchmark class against which the non-Gaussian contour class is compared.
In the notation of Definitions 2 and 1, for every ,
⊢ LeanThe treatment radius is specific to this analysis. It is the stability condition that controls the code-residual perturbation of the treatment transform in the contour argument. Throughout, a code residual means formed with the deterministic supplied clipped code . All risk statements condition on deterministic supplied clipped treatment and outcome codes and evaluate fixed-code performance at the current sample index.
We now package these restrictions into the law classes used by the main theorem and by the ACE and Gaussian comparisons. The broad spectral class keeps the cumulant-separated treatment noise and the treatment-code control that are needed for the contour construction.
Using the notation of Definitions 2 and 1, and with the structural conditions recorded in Assumptions 2, 3, 4, 5, 6, 8, 9, and 10, define Thus membership in is exactly determined by the displayed law, range, tail, cumulant-separation, and treatment-code radius conditions.
⊢ LeanThe bounded-outcome Gaussian class uses the same partially linear residual functionals but imposes Gaussian treatment noise and the supplied-code restrictions at exponent . It is the source-stated Gaussian comparison class for the diagnostic and minimax criteria below.
Here are the residual functionals determined by in Definition 3.
⊢ LeanThe ACE comparison class records the published cumulant-separated setup at order . Its member conditions mirror the preceding assumptions while using the code radii that enter the JMS finite-order guarantee.
Fix an admissible parameter record with positive sample-size coordinate and with and with radius sequences satisfying At any positive class index , the JMS ACE comparison class consists of all model records whose observed-data component is a probability law on , carrying a real , measurable functions , and deterministic measurable supplied code sequences whose clipped versions are denoted and . The coordinates and are -integrable, and the conditional-mean identities hold -almost surely. Writing and , the member conditions in Assumptions 2, 3, 4, 5, 6, 8, 9, and 10 hold in the following form: with the conditional-mean identity for holding -almost surely; with the displayed bounds on and holding for -almost every ; with the conditional tail bound holding -almost surely; and
⊢ LeanRisk comparisons use a generalized lower quantile for absolute estimation error. Because , the controlled level lies below one half, so is a lower quantile of the absolute error and every quantile bound below is to be read on that scale. The definition is stated for an arbitrary real sample statistic so it can be applied both to estimator errors and to the published ACE quantile bound.
In the sampling setup of Assumption 1 and the notation of Definitions 2 and 1, let be an observed-data law, let , and let be a real-valued sample statistic. Under the product law , write For , the generalized lower -quantile of is This is the left-continuous generalized inverse of evaluated at probability level .
⊢ LeanThe contour factorization in the next section is expressed through the covariate-level error made by the supplied treatment code. The following notation isolates that error and the corresponding outcome-contamination transform.
Fix a law , an index , and the clipped supplied treatment code . The treatment-code error and the outcome contamination are the covariate functions and the contamination transform is their weighted exponential transform
⊢ LeanFor comparison with Jin et al. (2025), we also record the eligibility quantities that enter the published order- ACE condition. They combine code radii, sample size, tail scales, coefficient range, and cumulant separation into the finite-order admissibility check.
Under the notation and conditions in Assumption 4, Assumption 5, Assumption 8, Assumption 9, Assumption 10, Assumption 13, Assumption 14, Definition 2, and Definition 1, the four quantities entering the Jin–Mackey–Syrgkanis order- eligibility condition at index are
⊢ LeanThe minimax criteria are class-indexed and fix the supplied codes at the current sample size. This keeps the decision problem focused on laws compatible with the same clipped code pair.
For fixed supplied code sequences, let and denote their current-index clippings to and . All infima below range over measurable estimators based on , and each supremum is restricted to laws in the displayed class whose clipped treatment and outcome codes at index are and . Given a parameter record and supplied code sequences , a class of laws carries the notation for the subset cut out by this fixed-code restriction, and we write , , and for the three risks defined next whenever their dependence on the parameter record and the supplied codes has to be displayed. Define the non-Gaussian minimax mean-squared risk by Define the Gaussian JMS minimax mean-squared risk by For , define the Gaussian JMS minimax generalized lower-quantile risk by
⊢ LeanThe order- ACE guarantee is invoked in later comparisons through its eligibility predicate. The predicate below is the finite-sample condition formed from the four quantities in Definition 8.
With defined as in Definition 8, the JMS order- eligibility condition is on the indices for which the displayed expressions are defined.
⊢ LeanFinally, the common ACE comparison class is obtained by adding the treatment- and outcome-code restrictions to the broad spectral class. This subclass is the domain on which the contour and ACE upper guarantees are aligned in the comparison section.
The ACE comparison subclass of the non-Gaussian class in Definition 3 is It consists of laws satisfying the broad spectral-class conditions together with the displayed supplied-code restrictions.
⊢ LeanZero instruments and contour identification
The analytic mechanism starts from a zero of the treatment-innovation transform and turns it into an instrument for the residualized treatment. In the partially linear model, the supplied treatment code enters through the residual ; the point of the construction is that a transform zero of the innovation continues to generate usable population moments after the residual has been shifted by the treatment-code error. The use of transform zeros follows a classical identification tradition for moment-generating and characteristic functions (Kagan et al., 1973; Mattner, 1992; D’Haultfoeuille, 2011; Darolles et al., 2011; Hu et al., 2022), while the contour formulation below is aligned with transform-based estimation in econometrics (Feuerverger et al., 1981; Singleton, 2001; Carrasco et al., 2007; Carrasco, 2017).
The first object is the treatment-innovation transform and its zero-based instrument. If is a zero of multiplicity , differentiating the exponential tilt at that zero produces the polynomial-exponential weight used throughout this section.
Let , let , and let satisfy , , and analytic order for at . Define
⊢ LeanThe instrument is indexed by the complex zero and the positive multiplicity . Its role is local in the transform variable but global in the residual shift: the same weight will be evaluated at shifted innovations and at the code residual.
Fix a model-parameter record , a model with observed-data law , and an integer . Let Suppose that:
(Non-Gaussian class.) The model belongs to as specified in Definition 3.
(Zero and multiplicity.) For some and integer , and the analytic order of at is .
(Instrument.) is the zero-based instrument from Definition 12, equivalently
Then, for every , Moreover, and Whenever the denominator is nonzero, the coefficient is identified by
⊢ LeanTheorem 1 gives the moment-ratio form of the argument when a valid zero is known. The identity applies uniformly over real shifts , so conditioning on converts the residualized treatment into a valid zero moment. The denominator identity identifies the relevant scale as the product of the nonzero derivative at the transform zero and the nuisance transform ; when that scale is nonzero, the partially linear coefficient is the displayed population ratio.
For data-driven contour selection, the paper uses observable transforms of the code residual. The treatment-residual transform carries the zero geometry, the outcome-weighted transform carries the numerator information, and a contour average of extracts the coefficient over all enclosed zeros at once.
For , define Let be a positively oriented circle centered at the origin with positive radius, and suppose is analytic on a neighborhood of the closed disk bounded by . Let denote the number of zeros of strictly inside , counted with multiplicity. When is zero-free on and , define
⊢ LeanHere is the positively oriented circle used for the contour integral, is the enclosed zero count, and is the normalized contour functional. The normalization by the zero count makes the functional an average over the enclosed zero structure.
The next identity links this observable contour world to the latent treatment innovation. It factors the residual transform through the innovation transform and isolates the covariate-only contamination term that enters the outcome-weighted transform.
Let be a model with observed-data law in the non-Gaussian class of Definition 3. For define, for , and where is the outcome-contamination function in Definition 7. Then, for every , and
⊢ LeanThe proof is deferred to Section E.
Lemma 1 is the bridge from residual contamination to contour identification. The factorization states that the observable residual transform inherits zeros from the treatment-innovation transform whenever the nuisance transform stays away from zero. The second display decomposes into the derivative term carrying and a nuisance ratio; integrated around a valid contour, the latter contribution is analytic inside the contour and the former contributes through the enclosed zero count.
The population identification statement packages these ingredients into the contour ratio used later by the statistical construction. The contour must be valid for the observable residual transform and for the nuisance transform on the same disk.
Fix , a law as in Definition 3, and a certified contour-bank input as in Definition 17. For any bank index , let be the positively oriented circle with radius , and let Assume:
(Residual zero-freeness.) for every .
(Nuisance zero-freeness.) for every on or inside .
(Positive contour count.) The zero count , counted with multiplicity inside , satisfies .
Then the contour functional of Definition 13 identifies the partially linear coefficient and is given by the normalized contour integral:
⊢ LeanTheorem 2 establishes the population contour identity on any bank circle satisfying the stated zero-free and positive-count conditions. Residual zero-freeness makes well defined on the integration path; nuisance zero-freeness keeps the zeros enclosed by the residual transform aligned with the innovation-zero geometry supplied by Lemma 1; and the positive count gives a nontrivial normalization. Under those conditions, the contour average of the observable ratio is exactly the partially linear coefficient.
The section therefore reduces identification with code residuals to a geometric task: find a circle on which the observable residual transform is separated from zero, whose interior contains at least one residual-transform zero, and on whose closed disk the nuisance transform is separated from zero. The next section supplies fixed-separation conditions and a finite contour bank under which such circles are available with enough empirical stability to yield the stated fixed-code risk bounds.
Main fixed-separation result
The preceding section gave the population contour identity. The fixed-separation result turns that identity into a uniform statistical construction by combining three ingredients: an explicit zero-localization radius from cumulant separation, an treatment-code gate that preserves the relevant zero geometry, and empirical transform control on a finite contour bank. The empirical-process component follows the same uniform-supremum logic used in semiparametric and double/debiased learning analyses (Bickel et al., 1993; van der Vaart, 1998; van der Vaart et al., 1996; Chernozhukov et al., 2018), while the contour bank uses finite analytic localization of zeros of exponential transforms (Michelen et al., 2019; Eremenko et al., 2021; Dinh et al., 2021).
The first auxiliary bound fixes the Luxemburg normalization used throughout the paper. It converts the sub-Gaussian convention in Definition 1 into a concrete moment-generating-function envelope, so the constants entering the zero-localization radius remain explicit.
Following the Luxemburg sub-Gaussian convention in Definition 1, let be a Luxemburg scale and let be a real-valued measurable random variable on a probability space with law . Assume:
(Integrability.) and .
(Centeredness.) .
(Luxemburg envelope.)
Then, for every ,
⊢ LeanThe proof is deferred to Section E.
Lemma 2 supplies the analytic tail envelope used to control treatment-innovation transforms on complex disks. With the exact convention fixed in Assumption 8, the bound carries the displayed factor , which then feeds directly into the localization constants.
Cumulant separation gives the treatment-innovation transform a zero inside an explicit disk. This step is the bridge from the population identification formulas in Theorems 1 and 2 to a finite search region.
Fix a primitive parameter tuple and a law for . Assume:
Define Then there exists such that and .
⊢ LeanThe proof is deferred to Section E.
The radius is determined by the fixed cumulant order, the treatment-noise Luxemburg scale, and the separation level. The contour construction expands this disk to the outer radius , allowing the finite bank to place candidate circles around the localized zero while retaining a buffer for perturbation control.
The numerator transform also needs a fixed envelope on the search region. The next bound records a population scale for the observable outcome-weighted transform on the disk used by the contour bank.
Let be a parameter record as in Definition 1, and let be a model with observed-data law as in Definition 2. Write , let be the outer search radius from Definition 17, and define the population numerator envelope Suppose that:
(Non-Gaussian class.) The model , equivalently its law , lies in the non-Gaussian class of Definition 3.
(Search window.) The complex point satisfies .
Then the observable outcome-weighted transform from Definition 13 satisfies
⊢ LeanThe proof is deferred to Section E.
Lemma 4 keeps the outcome side of the contour ratio on the same fixed scale as the denominator search region. The envelope depends on the primitive bounds and , so the later risk constants can be indexed by the fixed class constants.
The estimator is a finite, split-sample contour program. The pilot fold selects a circle using the empirical winding number and a denominator lower-modulus certificate; the evaluation fold computes the normalized contour ratio on that selected circle and clips the real midpoint to the parameter range. The certified-record clauses state the corresponding finite-rational implementation contract for the same statistic.
The estimator is the packaged adaptive contour estimator on the public domain where the current supplied treatment-regression code underlying is measurable. It takes as inputs , the supplied treatment-regression code sequence through its current clipped values , the parameter record , and fixed certified records and representing , and , with Here contains , with and , with , , positive constants , and nonnegative nonincreasing sequences on positive indices.
Define the deterministic folds For each real observation coordinate and integer , define the floor-dyadic interval
Form the residualized treatment values For each fold , derivative order , and complex , define and Thus an empty fold contributes the empty sum and is normalized by one.
For represented observation data, define where , , and are certified-real records naming , , and . The full represented-data input is For , put
Invoke the contour bank determined by . For the th circle, define At error tolerance one, refine , , and using their own executable moduli, and let , , and be the largest absolute endpoint of the corresponding rational intervals. For a canonical observation name this is the interval with the convention above; for it is the certified subtraction interval obtained from and , whose modulus uses the two half-error input moduli. Using for and , form rational upper bounds for the empirical derivative sums on each .
For each , compute a rational bracket of width at most enclosing Let be the lower endpoint of this bracket. When , define Use the rational Lipschitz bound
Compute a rectangle enclosure, with coordinate widths strictly below , for Define when the real coordinate contains a unique nonnegative integer and the imaginary coordinate contains zero. A circle is admissible when Let If the admissible set is empty, set
On the evaluation fold for the selected circle , compute a lower endpoint If set
Otherwise define the evaluation-fold contour integrand Use the rational Lipschitz bound For each rational tolerance and rational Lipschitz bound , the schedule spends one third of the tolerance on discretization and uses so that .
Enclose the unnormalized contour integral Divide the resulting rectangle by to enclose with coordinate radius at most . Take to be the real midpoint of that enclosure, so that approximates the normalized contour integral to within : Only rational quantities are ever formed, so the estimator returns this certified rational approximation rather than the integral itself; the tolerance is negligible beside the statistical error.
Output
The statistic is a total Borel rule because every branch has a displayed output, including the fallback value used when the empirical admissibility or evaluation-fold lower-modulus checks fail. The represented-data clauses give an executable correspondence for implementations satisfying the bounded contour-arithmetic specification of Definition 16; the statistical risk statements below concern the same ordinary-sample statistic.
The main theorem gives the fixed-code minimax statement under fixed cumulant separation. The fixed-code restriction keeps the supplied clipped treatment and outcome codes common across the laws in the risk criterion, matching the decision-theoretic convention in Definition 9.
Fix and positive constants . Then there are constants and , depending only on these fixed primitive constants, such that the following statements hold.
For every measurable covariate space, let be a parameter record whose primitive entries equal the fixed constants above, and fix one experiment-wide record consisting of the certified contour-bank input and the represented range input for . For every parameter record with the same primitive entries as , and for every base model satisfying the usual measurability, conditional-mean, and integrability requirements, write for the sample size of . Let be the total Borel statistic of Algorithm 1, constructed from the transported fixed records and the base treatment-code sequence. Define the fixed-code non-Gaussian class Define the fixed-code JMS ACE class
(Measurable statistic.) The statistic is measurable.
(Represented execution.) For every compiled bounded spectral adapter satisfying the full canonical build-and-compilation specification for , the transported fixed records, and the base treatment-code sequence, the represented-data execution realizes the same as in Algorithm 1.
(Class relations.) Every law in belongs to . If , then coincides with the -restricted subclass of Definition 11.
(Non-Gaussian fixed-code risk.) If the sample law is the i.i.d. product law of Assumption 1, is nonempty, , and where is the search radius from the contour-bank construction in Algorithm 1, then and with generalized lower quantiles as in Definition 6.
(ACE fixed-code risk.) If is nonempty, , and then where the infimum is over measurable real-valued estimators of . Moreover,
Theorem 3 establishes the fixed-separation statistical guarantee. On the non-Gaussian fixed-code class, the minimax mean-squared risk is sandwiched between constants times , and the contour statistic attains the upper bound. The same statistic also attains the MSE order and the displayed generalized-quantile bound on the fixed-code JMS ACE class whenever that class is nonempty and the treatment-code radius satisfies the stated small-radius gate.
The intuition is the same as in the population contour identity, with sampling error added. Cumulant separation localizes a zero of ; the radius controls the perturbation from to , preserving a usable contour inside the finite bank; and the split empirical transforms approximate the population transforms on the selected circle at the usual stochastic scale. Squaring the coefficient error gives the MSE order, while the lower bound records the matching parametric difficulty inside the same fixed-separation experiment.
The sequence-level statement packages the same construction for common primitive constants. It also records the bounded-outcome Gaussian diagnostic under the JMS comparison class, separating the fixed-separation non-Gaussian rate statement from the Gaussian benchmark risk calculation.
Fix and positive constants . For every measurable covariate space, the following hold.
(Non-Gaussian sequence.) Let have these fixed values of , and fix the primitive records and of Definition 17 and Algorithm 1. Let be any parameter sequence with the same fixed values of as , and with common supplied radii and . Let be any measurable deterministic code pair, and write for the clipped pair. Suppose that and, for all sufficiently large , Then there are constants such that, with the translated-dyadic contour bank generated from and , for all sufficiently large the transported primitive bank for generates the same bank , the statistic of Algorithm 1 is measurable, every compiled bounded spectral adapter satisfying the full canonical build and compilation specification is represented by the corresponding execution on the same transported primitive records, and Moreover, the same statistic witnesses the upper bound:
(Gaussian class.) For every parameter record , writing for its sample size, every in Definition 4 satisfies For every measurable deterministic code pair , whenever the clipped-code intersection of is nonempty,
Theorem 4 gives two common-experiment conclusions. Along any sequence with the same fixed primitive constants and common supplied radii, the translated-dyadic contour bank stabilizes and the non-Gaussian minimax MSE remains of order , with the contour statistic attaining the upper bound. On the bounded-outcome Gaussian JMS class, the partially linear target is identically zero, so both the fixed-code MSE risk and the fixed-code generalized-quantile risk equal zero whenever the clipped-code intersection is nonempty.
The non-Gaussian sequence clause reflects the finite nature of the contour bank: fixed primitive constants determine the same search geometry for all sufficiently large indices satisfying the displayed radius gate. The Gaussian clause records a diagnostic property of the simultaneous bounded-outcome and Gaussian treatment-noise restrictions in Definition 4; under those restrictions, the target itself is pinned to zero, and the corresponding minimax risks inherit that degeneracy. Together, Theorems 3 and 4 provide the fixed-separation rate statement used for the ACE comparison in the next section.
Comparison with ACE and Gaussian benchmarks
The fixed-separation theorem in Theorem 3 gives a contour guarantee on the broad spectral class and on the fixed-code ACE class. This section records the comparison with the published order- ACE guarantee of Jin et al. (2025) on the same clipped-code experiment. The comparison has two parts: first, the class relations align the cumulant-separated spectral and ACE restrictions; second, the two generalized-quantile upper bounds can be read on the common class under the same JMS eligibility condition.
The following proposition states the alignment and the resulting upper-guarantee comparison. The constant is fixed by the primitive constants in the contour construction, while and the estimator come from the published ACE handle. The displayed quantity is the order- generalized-quantile upper guarantee from Jin et al. (2025). The sequence clause compares the displayed upper bounds on the common clipped-code class.
Fix an ACE order and positive constants . There is a constant , depending only on these fixed constants, such that the following statements hold.
Let , let be the constant in the published order- ACE generalized-quantile guarantee, and write for the estimator supplied by that published ACE handle. Let be a parameter record whose , and coordinates equal the displayed constants, and fix the certified primitive bank and range records associated with . For any parameter record sharing the same fixed experiment constants as and the same probability level , and for any supplied code sequences whose clipped versions are and , the following hold, under the usual measurability and integrability conditions.
Class relations. At the sample size of , If the comparison exponent of satisfies , then the common clipped-code convention gives Here the classes and the comparison subclass are those of Definitions 5, 3, and 11.
ACE upper guarantee. If a law uses the clipped supplied treatment and outcome codes , if the JMS eligibility condition of Definitions 10 and 8 holds, and if and , then where is the generalized quantile in Definition 6 and
Spectral upper guarantee. If contains at least one law using the clipped supplied codes, , holds, , and where is the outer contour radius of Definition 17 determined by the parameter record in force, then the certified contour statistic of Algorithm 1, built from the fixed primitive records, satisfies where the supremum ranges over laws in the published ACE class whose clipped supplied treatment and outcome codes are and .
Upper-guarantee separation along sequences. For any sequence of parameter records whose sample size at index is , with the same fixed experiment constants as , the same radius sequences , and , suppose that eventually , holds, , , and the published ACE class contains a law using the clipped supplied codes. If then
Proposition 1 places the spectral and ACE guarantees on a shared statistical domain. The class relations use the fact that the ACE code control implies the treatment-code control needed by the spectral class, while the -restricted spectral subclass matches the published ACE code restrictions when . Under JMS eligibility, the ACE bound retains its finite-order nuisance terms, and the contour bound is governed by the fixed-separation sampling term . Along the displayed nuisance-dominance sequences, the ratio of the contour upper guarantee to the imported finite-order ACE upper guarantee tends to zero. The conclusion is an upper-bound comparison on the common clipped-code class; the ACE-related lower-bound statement used in the paper is the class-minimax lower bound over all measurable estimators on the fixed-code class.
For later reference, the class relation used inside Proposition 1 is also recorded separately. This isolates the comparison-class bookkeeping from the estimator-specific upper bounds.
Fix a measurable covariate space and a parameter record with , , , satisfying , , positive constants , and nonnegative nonincreasing code-radius sequences on positive indices. Then, for every class index ,
(JMS inclusion.) For every model , if as defined in Definition 5, then as defined in Definition 3.
(Comparison inclusion.) For every model , if as defined in Definition 11, then .
(Equal-exponent equivalence.) If , then for every model ,
The proof is deferred to Section E.
Lemma 5 is the set-inclusion component of the comparison. It identifies which restrictions are shared by the broad cumulant-separated class and the published ACE class, and it pins down the exact equality case under the common clipped-code convention. This makes the comparison in Proposition 1 an upper-guarantee comparison on the same law class and the same clipped-code experiment.
The second benchmark is the bounded-outcome Gaussian comparison class of Definition 4. Proposition 2 gives a bounded-outcome Gaussian diagnostic for the intersection of the exact partially linear conditional mean, bounded observed outcome, and nondegenerate Gaussian treatment innovation. In every nonempty fixed-code intersection of that class, the decision-theoretic mean-squared and generalized-quantile risks are zero because the member restrictions determine the target value itself.
Fix a measurable covariate space and a parameter record with sample size , orders satisfying and , exponent , probability level , positive constants , and nonnegative nonincreasing code-radius sequences on positive indices. Then every model satisfying Definition 4 has Moreover, for any deterministic supplied code sequences , if there exists at least one whose clipped codes at sample size satisfy for every , then the fixed-code Gaussian minimax risks are both zero:
⊢ LeanProposition 2 gives a diagnostic for the source-stated Gaussian comparison class. A bounded observed outcome has bounded conditional support, while the partially linear term is driven by nondegenerate Gaussian treatment noise under Definition 4; the proposition records the resulting target value and the induced zero minimax risks on each nonempty fixed-code intersection. Thus the Gaussian benchmark characterizes that simultaneous set of restrictions, whereas the fixed-separation comparison with ACE in Proposition 1 is the active finite-sample upper-guarantee comparison on the cumulant-separated class.
Together, Proposition 1, Lemma 5, and Proposition 2 separate the two comparison roles. The ACE alignment places the published finite-order guarantee of Jin et al. (2025) and the contour guarantee on a common cumulant-separated domain, with the sequence clause identifying the regime in which the ratio of displayed upper guarantees tends to zero. The Gaussian proposition explains the bounded-outcome Gaussian class by its implied target and the corresponding decision-theoretic risks, matching the diagnostic role already summarized in Theorem 4.
Explicit mixture reductions and local benchmarks
The preceding comparison keeps the cumulant separation level fixed. This section records two constructive weak-non-Gaussian benchmarks that make the transform-zero mechanism explicit in one dimension. The first benchmark treats a symmetric two-component Gaussian mixture, where the relevant zero is known and the contour search collapses to a sine moment. The second follows a Gaussian–Rademacher path indexed by a mixture amplitude, so the denominator scale and the fourth cumulant can be read from the same parameter. These examples connect the contour construction to non-Gaussian identification ideas in independent-component and structural-shock settings (Comon, 1994; Hyvärinen et al., 1997; Hyvärinen et al., 2001; Lanne et al., 2017; Lee et al., 2024; Hoesch et al., 2024; Reizinger et al., 2025).
The first construction is the clipped sine-ratio estimator . It uses the known imaginary zero of the symmetric mixture transform and estimates the corresponding ratio directly from the empirical moments of and .
Write for the projection of a real number onto . For observed triples , let and let , where is the supplied treatment code clipped to . The clipped sine-ratio estimator is
⊢ LeanThe threshold in Definition 14 is tied to the population denominator scale that appears in the next proposition. In the symmetric mixture experiment, the treatment innovation has a known transform zero at , and control of the supplied treatment code keeps the sine denominator bounded away from zero.
For any real constants , there is a constant , depending only on these four constants, such that the following holds. Let be a parameter record whose constants are the ones just fixed, write for its sample size, and let be a law in the corresponding model. Assume the usual measurability and integrability conditions, and suppose that
(Sampling.) are drawn according to the i.i.d. product law in Assumption 1.
(Symmetric mixture.) The treatment innovation has law
(Independence and ranges.) The treatment innovation is independent of as in Assumption 2, the outcome innovation satisfies Assumption 3, and the range and tail conditions in Assumptions 4, 5, 6, and 9 hold.
(Treatment code.) The supplied treatment code obeys
Then the fourth cumulant of is the treatment-innovation moment-generating function satisfies and, writing and , with Moreover, for the sine estimator in Definition 14,
⊢ LeanProposition 3 gives a fully explicit root- benchmark for a symmetric non-Gaussian treatment innovation. The mixture has fourth cumulant , its transform vanishes at , and the sine denominator equals a positive contamination factor times . The resulting empirical ratio attains mean-squared error under the stated range, tail, independence, and treatment-code conditions.
The next elementary inequality is the algebraic device behind the sine-risk calculations. It separates numerator fluctuation from denominator fluctuation after clipping, using only a lower bound on the population denominator mean.
Use the empirical-average notation from Definition 1. Let and be real-valued functions of one observation, let , and let be any sample. Assume:
(Sample size.) .
(Clipping range.) The target value lies inside the clipping bound , in the sense that .
(Denominator scale.) The scale is strictly positive.
(Population denominator level.) The denominator mean satisfies .
Define the clipped ratio statistic at threshold by Then, pathwise on this sample,
⊢ LeanThe proof is deferred to Section E.
Lemma 6 is deterministic. In applications, is the sine-weighted residualized treatment, is the centered outcome-noise contribution to the sine numerator, and is the population denominator scale. Once the sine moment gives , the ratio error is controlled by two empirical averages.
The Gaussian–Rademacher path provides a local family approaching Gaussian treatment noise through an explicit fourth-cumulant parameter. Let denote the centered Gaussian–Rademacher treatment innovation in the next statement. Its first positive transform zero, denominator scale, and cumulant magnitude are all explicit functions of .
Fix real constants . There is a constant , depending only on these four constants, such that the following statement holds. Let be an admissible parameter record whose constants are the ones just fixed and whose cumulant order is , let , and let be a law in the model . Assume:
(Gaussian–Rademacher path.) The treatment innovation has the law of where , is symmetric Rademacher, and and are independent.
(Sampling and model restrictions.) The law satisfies Assumptions 1, 2, 3, 4, 5, 6, and 9.
(Current treatment-code radius.) For the clipped supplied treatment code , under the usual integrability condition for this quantity.
Define Then the treatment-innovation moment-generating function satisfies the fourth cumulant satisfies , and is the first positive zero of the characteristic transform: and for every . Moreover, for every ,
Let and . If , then and With denoting empirical averaging over , define where clips to . Its mean squared error obeys The same smallness condition also gives and the equivalent cumulant-parameterized bound
⊢ LeanThe proof is deferred to Section E.
Lemma 7 shows how the explicit sine reduction behaves as the treatment-noise law moves toward Gaussianity. The first zero is , while the denominator scale decays exponentially as decreases. Expressing the same bound through makes the dependence on fourth-cumulant magnitude explicit. The condition is exactly the treatment-code accuracy gate used by this pathwise calculation.
The section closes with a local benchmark statement. It combines the Gaussian–Rademacher path just described with a localized version of the published ACE oracle envelope from Jin et al. (2025). The sequence supplies the shrinking cumulant threshold, and the ACE clause evaluates the published bound after replacing the fixed separation level by the local value .
Let be a real sequence with for every , nonincreasing in , and . Fix . The following two assertions hold.
(Gaussian–Rademacher path.) There is a constant , depending only on , such that the following holds. Let be any parameter record whose constants are the ones just fixed and whose orders are and . For any and any law in a model , assume that the treatment noise satisfies where , is symmetric Rademacher, and . Assume also Assumptions 1, 2, 3, 4, 5, 6, and 9. At the sample size of , assume the supplied clipped treatment code satisfies Then the conclusion of Lemma 7 — the explicit mean-squared-error bound along the Gaussian–Rademacher path — holds with constant for .
(Local ACE oracle envelope.) For every published ACE handle, every , and every , if the handle satisfies the conclusion of Theorem 5.4 of Jin et al. (2025) at , then and the following local ACE statement holds. For every parameter record whose probability level is , writing for its sample size, if and , then for every deterministic treatment code , deterministic outcome code , and every model, whenever and, writing and for the orders of and , the local ACE class conditions hold together with Assumptions 2, 3, 4, 5, 6, and 9, the treatment-noise exponential envelope and the nuisance-radius bounds and whenever the local JMS eligibility quantities satisfy and the published order- ACE estimator obeys where
Theorem 5 gives two procedure-specific local upper benchmarks. The Gaussian–Rademacher clause reuses Lemma 7 for the path under the stated treatment-code condition. The ACE clause specializes the published order- generalized-quantile guarantee of Jin et al. (2025) to a local cumulant threshold , with the displayed eligibility quantities and the resulting envelope . The last algebraic clause records how an oracle upper bound that is valid under two sets of hypotheses can be summarized by the smaller of the two displayed upper inputs.
These local statements complement the fixed-separation results by making the dependence on the non-Gaussian signal explicit in tractable submodels. For the symmetric mixture, the sine denominator is bounded by a fixed positive scale and the risk is . Along the Gaussian–Rademacher path, the same sine construction has a denominator scale and a cumulant magnitude , yielding the explicit exponential dependence displayed in Lemma 7. The localized ACE envelope gives the corresponding published finite-order benchmark under its stated eligibility and nuisance-radius conditions.
Limitations and future work
The fixed-separation results in Theorems 3 and 4 use a cumulant-separation constant that is fixed across the experiment. The local benchmarks in Theorem 5 point to a complementary triangular-array program in which the non-Gaussian signal is allowed to shrink with the sample size. This section records that agenda separately from the fixed- contour theorem and from the procedure-specific local upper bounds of the preceding section.
A useful starting point for that program is a finite, sample-facing object that stores the contour information available at a given library depth. The following definition names the empirical handle used to organize candidate contours, empirical winding information, contour moments, and angular variability.
Fix a parameter tuple with positive sample size , fixed ACE order , cumulant order , , , strictly positive , and nonnegative nonincreasing sequences and on positive indices. Work on a measurable covariate space with observations , and fix a probability law with real functional , measurable conditional means and , integrable and , and deterministic measurable supplied regression-code sequences for treatment and outcome.
For each library depth , let be the deterministic finite library of rational circles represented by the positive rational radii , . For a realized sample indexed by , form the empirical transforms and on the prespecified deterministic fold , using indices below for and indices at least for . For a circle , , on which the empirical boundary modulus is certified positive, the candidate record consists of quantities definitionally computed from the empirical contour integrands Specifically, the record contains a rational interval enclosure of the boundary modulus , a complex rational interval enclosure for , a complex rational interval enclosure for , and a rational interval enclosure of the angular variance Each enclosure is accompanied by its soundness certificate for the displayed empirical quantity, and the actual empirical transforms on the stated deterministic fold are the transform inputs of the schema. These records are candidate data for one realized sample, one deterministic inference fold, and one rational circle.
⊢ LeanDefinition 15 provides a concrete data structure for studying adaptive contour selection in local regimes. It keeps the deterministic library, fold-specific empirical transforms, certified boundary separation, contour moments, and angular variance in one record. For shrinking cumulant separation, such records give a way to describe what a selector observes before choosing among contour-based procedures and more classical alternatives.
The open statistical question is naturally triangular. Let denote shrinking positive cumulant-separation thresholds. The fixed-separation theorem controls risk when the primitive separation constant is fixed, while the local benchmarks in Theorem 5 describe particular upper envelopes when the non-Gaussian signal is weak. The remaining problem is to identify the sharp joint rate and the corresponding inference theory across the ordinary DML, finite-order ACE, and contour regimes (Chernozhukov et al., 2018; Newey et al., 2018; Mackey et al., 2018; Jin et al., 2025; Lee et al., 2024; Hoesch et al., 2024).
A natural next question concerns triangular treatment-noise laws satisfying under the same supplied code sequences. The target is a data-driven selector among ordinary DML, finite-order ACE, and global-contour procedures that attains the sharp minimax mean-squared-error rate as a function of supports uniformly valid inference across these regimes, and admits a matching local minimax lower bound. Future work could determine the sharp-rate functional, the corresponding uniform-inference criterion, and the local lower-bound construction for this frontier.
⊢ LeanRemark 1 states the future-work target as an open research agenda rather than as an identification or risk theorem. Its rate arguments are the sample size, the two supplied-code radii, and the local cumulant-separation threshold. The agenda also includes a uniform-inference criterion and a local minimax lower-bound construction, so any completed theory would need to align estimation, coverage, selection, and lower-bound witnesses on the same triangular experiment.
This separation clarifies the scope of the paper’s established results. The fixed-separation contour theorem gives the MSE characterization under fixed cumulant separation and the stated treatment-code stability condition. The local mixture and ACE benchmarks give constructive upper comparisons under their own displayed hypotheses. The shrinking-separation program in Remark 1 is the natural next step for a unified adaptive theory across Gaussian-adjacent and fixed non-Gaussian regimes.
Appendices
Proofs for identification and analytic localization
This appendix collects the analytic statements that support the identification argument in Theorem 1, Lemma 1, and Theorem 2 and the fixed-bank localization used in Theorem 3. The common theme is that zeros of moment-generating functions can be controlled by classical complex-analytic arguments and then transferred to residualized-treatment transforms. The zero and completeness perspective is standard in identification arguments based on transforms (Kagan et al., 1973; Mattner, 1992), and the finite-bank construction uses the same analytic geometry that underlies modern zero-localization bounds for moment-generating functions (Michelen et al., 2019; Eremenko et al., 2021; Dinh et al., 2021).
For Theorem 1, the analytic order of a zero of determines which polynomial-exponential weight annihilates shifted treatment innovations. Independence of the treatment innovation from , as imposed in Assumption 2, then turns the shifted identity into a conditional moment for the code residual . The denominator calculation is the corresponding differentiated transform identity, and the partially linear conditional-mean restriction in Assumption 3 supplies the numerator relation.
The factorization in Lemma 1 isolates the effect of treatment-code contamination. Since , independence separates the residual transform into the treatment-innovation transform and the nuisance transform. The outcome-weighted transform has the derivative term carrying plus the covariate-only contamination transform from Definition 7. This decomposition is the algebraic input for Theorem 2: on a contour where the residual transform is separated from zero and the nuisance transform has no zero on the closed disk, the logarithmic-derivative contribution counts the enclosed zeros and the analytic contamination contribution integrates to zero.
The remaining analytic step is uniform localization. Lemma 3 gives a disk, determined by the fixed cumulant order, sub-Gaussian scale, and separation constant, that contains at least one zero of the treatment-innovation transform. The finite-bank statement below turns that disk-level information into a deterministic list of circles, with explicit dyadic bookkeeping and a positive boundary-modulus certificate. Its role is to provide the law-blind contour geometry used later by the empirical statistic.
Let be a primitive parameter record with , , , , positive constants , and nonnegative nonincreasing code-radius sequences . Let be a certified-bank input record for in the sense of Definition 17, so that For any sample size , any model with observed-data law , and any membership certificate as defined in Definition 3, form the bank data of Definition 17 for and . Then , and the components of satisfy The bank radii are strictly increasing and lie in the search annulus: Writing and for the treatment-innovation moment-generating function, there exists an index such that
⊢ LeanLet be the bank produced from . The construction in Definition 17 gives and Thus and the displayed bookkeeping identities hold.
The radii have the explicit form Hence is strictly increasing. Since and , each bank radius satisfies
It remains to prove the treatment-noise certificate. The error-one certified refinements and the ceiling construction give Let The conditional-mean identity for centers . The treatment-tail condition in , together with Lemma 2, gives for every The same real exponential integrability gives the entire complex moment-generating function, normalization at the origin follows from the probability law, and the complex modulus estimate yields
Let be the number of zeros of in the open disk of radius , counted with multiplicity. Jensen’s zero-counting bound applied on gives
Since , Therefore
List the zeros in , with multiplicity, as The finite Blaschke factorization on the disk of radius gives a holomorphic function , zero-free in , such that where
The grid spacing is , and the mesh identities give , , and . Thus one listed zero modulus can be within distance of at most one grid radius: two distinct grid radii are separated by at least , while two distances would put them less than apart. Also The pigeonhole argument for this separated dyadic grid gives an index such that
By Lemma 3, has a zero with . Since , this zero lies strictly inside . If had a zero on , then , and the factorization together with the zero-freeness of would force for some . This would give contradicting the separation of the chosen grid point. Hence the contour has positive zero count:
It remains to lower-bound on this circle. Put On , the Blaschke product has , so the boundary envelope for gives . Also and , hence . Applying Harnack’s inequality to the positive harmonic function gives, for every , Since , Thus, for every ,
Since the elementary dyadic-exponential comparison for natural , applied to , gives
Now take . For each zero , Also , and therefore For each Blaschke factor,
Multiplying over the at most factors and using gives
Combining the last two lower bounds, Thus for this bank index , and the asserted uniform treatment-noise certificate follows.
∎Lemma 8 is the population geometric certificate used by the empirical selector in Algorithm 1. The proof below records the dyadic bookkeeping and the zero-localization argument behind that certificate.
Empirical process, stability, and lower-bound lemmas
This appendix collects the auxiliary statistical statements used by the fixed-separation contour theorem. The first two results control the analytic transforms that enter the contour statistic: one gives uniform concentration of the empirical residual and outcome transforms on a fixed disk, and the other gives a deterministic stability bound for the nuisance factor. Together they supply the stochastic and population inputs behind the contour-risk bounds in Theorem 3. The concentration statement uses the same sub-Gaussian scale conventions and product-sample structure as Definition 1 and Assumption 1, in line with standard empirical-process and high-dimensional probability arguments (van der Vaart, 1998; van der Vaart et al., 1996; Vershynin, 2018; Wainwright, 2019).
Fix real numbers and . There is a constant such that the following holds. Let be a parameter record with and sample size . Let be a model in , as in Definition 3, and let be sampled under the i.i.d. product law , as in Assumption 1. For each , write for the corresponding deterministic inference fold and suppose . Define, for , where , and let Then
⊢ LeanSet and define The constants and are strictly positive, so ; it depends only on the displayed primitive constants and .
Fix a fold for which is nonempty, and let , , and . Since the contour-bank radius in Definition 17 satisfies with , we have . Write , so . The range bound for in Definition 3 and the clipping of in Definition 2 give almost surely. Hence, for almost every observation, where the second inequality is the Young inequality obtained by expanding . Therefore Integrating and using the Luxemburg treatment-tail bound from Definition 3 yields
The outcome-weighted envelope is Indeed, Definition 2 gives . For almost every observation, the algebraic bound combines with the bounds in Definition 3, namely almost surely and , to give almost surely. The Luxemburg moment bounds give and , and hence Together with and the preceding residual envelope applied at radius , this gives the stated bound.
We use the following coefficient-series bound. Let be a nonempty deterministic finite subset of sample indices, and write . Let be measurable real functions, let , let , and suppose Define Then, under the i.i.d. product law for , The envelope gives because for ; the hypotheses and make the displayed majorant nonnegative. Under the product sample law, centering over the block gives The series of majorants is summable, and the centered analytic-series supremum bound yields
The same coefficient construction represents the empirical transform errors exactly. For the same block and for , with the exchange of summation, finite averaging, and integration justified by the doubled exponential envelope.
Apply the preceding bound first with the block , , , , and . Since , this choice has , and
Apply it again with the block , , , , and . Since , this choice has , and
Adding the two displays yields This proves the assertion.
∎Lemma 9 is the stochastic input for the selected-contour risk proof. The proof below supplies the disk-uniform bound used after the pilot fold selects a circle.
The next statement is the population stability input. It uses the direct treatment-code radius in Definition 3 to keep the nuisance transform close to one on the contour search region.
For every satisfying Definition 3, the associated treatment-code discrepancy obeys Consequently, for the corresponding factor and radius , If , then for every , and and have the same zeros in , counted with multiplicity.
⊢ LeanLet Since almost surely and by clipping,
The treatment-code condition in gives
For , the exponential remainder bound gives, with , Together with , this yields Taking the supremum over gives the displayed uniform bound.
Assume Then the preceding display gives throughout the disk, and therefore
It remains to record why the multiplicity comparison is legitimate at each point of the disk. The almost-sure bound implies that for every real , so the complex MGF is analytic in a neighborhood of every . The treatment-tail condition in , under the Luxemburg convention of Definition 1, gives the same local exponential-moment condition for : for every real , completing the square bounds by a constant multiple of , which is integrable. Hence the treatment-noise MGF is analytic in a neighborhood of every point of the disk as well.
By Lemma 1, Fix . The preceding lower bound gives , so the analytic order of at is zero. Applying additivity of analytic order to the displayed factorization gives Thus and have exactly the same zeros in , counted with multiplicity.
∎Lemma 10 is the deterministic transfer from treatment-transform zeros to observable residual-transform zeros. The proof below uses the clipping bound and the radius to control the nuisance factor on the search disk.
A separate tail calculation is used when Gaussian outcome innovations are embedded in lower-bound and benchmark paths. The following elementary scale bound matches the Luxemburg convention fixed in Definition 1.
Use the Gaussian-law notation of Definition 1. For every positive scale , let . Then is integrable under the law of , and
⊢ LeanLet Since , this variance is positive. Under , the density identity is, for every , because .
Since , the function is integrable on . The density identity therefore transfers this integrability to under , so is integrable under the law of .
For the expectation, the same density identity rewrites the Gaussian-law integral as The Gaussian integral formula applied with coefficient gives
Finally, since after squaring both sides this is Therefore
∎Lemma 11 is the tail-scale normalization used in the lower-bound construction below.
The final auxiliary result supplies the decision-theoretic lower-bound path. It constructs one-dimensional submodels in which the observed treatment law, regression functions, and supplied clipped codes remain fixed while the coefficient varies locally at scale. Such paths are the standard input for parametric lower bounds in semiparametric experiments (Bickel et al., 1993; van der Vaart, 1998).
Fix and real constants with . There exist constants such that For every parameter record whose constants are the ones just fixed, and every , if the non-Gaussian class in Definition 3 is nonempty, then there are a base law and a family such that:
(Class membership and target.) For every , and .
(Fixed observed treatment law.) For every , the law of under equals its law under .
(Fixed regressions and supplied codes.) For every , the functions and the clipped supplied codes under equal those under .
(Gaussian outcome innovation.) For every , the outcome innovation under has law and is independent of .
(Outcome equation.) For every , almost surely under .
(Product relative entropy.) For all real with , , and one has, with the Kullback–Leibler divergence,
Moreover, for every starting law in the class of Definition 5, there is a family such that:
(JMS ACE membership and target.) For every , and .
(Fixed observed treatment law.) For every , the law of under equals its law under .
(Fixed nuisance objects.) For every , the functions , the clipped supplied codes , and the treatment innovation under equal those under , and the treatment and outcome code-radius conditions in Definition 5 hold at .
(Gaussian outcome innovation.) For every , the outcome innovation under has law and is independent of .
(Outcome equation.) For every , almost surely under .
(Product relative entropy.) For all real with , , and one has
Finally, for every pair of clipped-code functions , if there is a law with and , then the fixed-code JMS ACE minimax risk satisfies
⊢ LeanSet Then , , and . For any admissible choice of parameters whose displayed constants agree with the fixed constants, Definition 1 gives , and hence .
The two-point testing bound at relative-entropy budget supplies a constant , depending only on that budget, such that whenever two -sample laws have relative entropy at most , every measurable sample statistic with finite quadratic risks satisfies The minimax reduction applies this bound after symmetrically clipping an arbitrary estimator to a range containing both target values. Define
Fix now an admissible choice of parameters with the displayed primitive constants and a class index . Assume first that is nonempty, and choose , represented by a base model . For each , define by preserving the joint law of under , drawing an independent and setting Equivalently, the path is the product of the retained -law and the centered Gaussian innovation, transported through this affine outcome map.
Along this path, and the clipped supplied codes satisfy The covariate-treatment marginal is also retained:
The treatment innovation is unchanged along the path: The outcome innovation is the fresh Gaussian coordinate: so
For , the coefficient range holds because . All treatment-side assumptions, the regression ranges, the cumulant separation, and the treatment code-radius condition are inherited from . The conditional mean condition follows from the displayed outcome equation, the independence of from , and its centering. Finally, Lemma 11 applied at scale gives Since is independent of , the conditional Luxemburg part of Assumption 9 is Hence
The same construction gives the asserted outcome equation:
It remains, for this branch, to record the product relative-entropy bound. For the affine Gaussian path, one observation satisfies The moment bound used here is the Luxemburg consequence Tensorization over independent observations gives If and , then Thus the required entropy display holds.
Now let . Define by the same affine Gaussian outcome construction from this starting law: The displayed preservation identities remain valid, and the target coordinate is the path parameter: and
Consequently the two JMS ACE code-radius inequalities are transported unchanged: Together with the transported treatment-side and range conditions, the coefficient bound , and the Gaussian conditional Luxemburg calculation above, this gives The Gaussian residual law, residual independence, outcome equation, and product entropy bound are the same as in the non-Gaussian branch: whenever .
It remains to prove the fixed-code ACE lower bound. Suppose a starting law has clipped codes . Set and let , , using as the starting law. Since , Both laws belong to the fixed-code JMS ACE class with clipped codes , and the entropy bound gives
Applying the two-point lower bound with target values and yields, for every measurable estimator whose two quadratic risks are finite, The minimax reduction clips arbitrary measurable estimators to , so the preceding finite-risk display applies inside the fixed-code risk. Taking the supremum over the fixed-code JMS ACE class and then the infimum over estimators gives
∎Lemma 12 is the lower-bound ingredient for the minimax statements. The proof below builds the fixed-code local path and applies the relative-entropy calculation used in the testing-to-estimation step.
The ACE branch of the proof verifies the same fixed-code path inside , with the treatment innovation and clipped nuisance codes preserved along the path.
Certified construction and executable correspondence
This appendix records the finite arithmetic contract behind the contour statistic in Algorithm 1. The statistical arguments in the main text use ordinary empirical transforms and contour ratios; the material here specifies a rational interval substrate, a deterministic contour-bank record, and the deterministic perturbation statement that connects a selected represented contour calculation to the population ratio. The construction follows the computable-analysis and interval-arithmetic convention that real inputs are accessed through nested rational enclosures with effective moduli (Weihrauch, 2000; Ko, 1991; Moore, 1966).
The first definition fixes the elementary arithmetic layer. It gives represented real and positive-real records, complex rectangles, guarded division, elementary functions on bounded intervals, circle-node construction, and certified mesh integration. These operations are the only numerical primitives used by the represented contour program.
The bounded-build specification consists of the following finite rational arithmetic operations and their stated enclosure properties.
A represented real number is given by rational closed intervals and an executable modulus such that A positive represented real number additionally supplies a rational lower bound with . The executable data are the interval sequence, the modulus, and the positive lower bound.
For rational intervals and , interval addition, subtraction, and multiplication are A complex rectangle is . Complex addition, subtraction, conjugation, and multiplication use these endpoint operations together with For , define The squared modulus of is enclosed by adding the corresponding lower and upper square bounds. Rational bisection between nonnegative rational endpoints gives lower and upper square-root enclosures whose width after bisections is at most the initial width times . Division is the guarded operation performed when the modulus routine returns a rational lower bound , with outward division by the positive rational interval enclosing .
The represented value of is constructed from Machin’s identity For , Intersecting the first requested enclosures gives rational nested enclosures. If is in lowest terms, the choice makes the propagated Machin remainder smaller than , giving an effective modulus.
For a rational interval contained in , with rational , the Taylor polynomials for , , and are evaluated by the preceding interval operations, with Lagrange remainders bounded by For a requested rational error , take After index , successive absolute Taylor terms decrease by at least a factor ; since and , the displayed gives error at most . Finite intersections make the output rectangles nested. Consequently, has a nested complex-rectangle extension with an effective modulus on every bounded input rectangle. Combining the constructed representation of with these trigonometric operations gives, from a positive radius representation and integers and , represented circle-node values
Finite rectangle sums and products are primitive recursive. If a complex-valued function on has a supplied rational Lipschitz bound , then a uniform mesh with has oscillation at most on each cell. Evaluating every node to a rational error budget divided by the displayed finite number of operations yields sound finite infimum, supremum, and trapezoidal-integral enclosures. The required finite computation consists of node evaluations, the displayed Taylor cutoffs, square-root bisection cutoffs, and finitely many endpoint operations. Absolute value of a real interval is implemented by the square-free endpoint rule implicit in and , with the same compositional modulus property.
A concrete implementation satisfies when it implements these algorithms, exports their soundness and effective-modulus contracts, and compiles the resulting module. Later executable conclusions are conditional on a compiled implementation satisfying this bounded-build specification.
⊢ LeanDefinition 16 supplies the finite representation interface used throughout the executable clauses. Nested intervals and moduli give a standard Type-2-style name for real values, while the endpoint operations and guarded division give sound complex-rectangle enclosures for the contour integrands. The mesh rule turns a supplied Lipschitz bound into finite enclosures for boundary moduli and contour integrals, matching the interval-analysis approach of Moore (1966) within the computable-real framework of Weihrauch (2000); Ko (1991).
The next definition specializes that arithmetic substrate to the contour bank. The primitive record stores certified names for the fixed separation, treatment-noise scale, zero-localization radius, and outer search radius. From those names and the parameter record, the bank construction produces a deterministic translated dyadic list of circles and a positive dyadic lower-modulus certificate.
Let be a parameter record with sample size , , orders satisfying and , exponent satisfying , probability level , constants , and sequences satisfying for and whenever , for . Let denote a supplied certified-real record with value , and let denote such a record together with a rational lower-bound witness satisfying . As part of the experiment convention, fix the supplied record where is exact with and The value contracts are Writing for the numerical factor, this reads . The search radius of the record is this outer radius: . The record is the primitive certified input shared by the bank and the estimator for this parameter record. The bank data assembled below from and is written .
Invoke the supplied moduli of and at error one. If the returned intervals have rational upper endpoints and , define Set The bank therefore contains circles, indexed by . For each , use certified addition on the fixed record and the rational point record for to form The associated real-valued radius field is
Define the integers Define the rational dyadic certificate and its real value by Every output datum is obtained by the displayed finite modulus calls, rational endpoint operations, certified additions with rational point records, and primitive recursion over the displayed integer bounds. The construction is a finite certified procedure in the supplied record .
⊢ LeanDefinition 17 converts the analytic radius from Lemma 3 into a finite bank used by Algorithm 1. The radii fill a translated dyadic grid outside the zero-localization disk and inside the outer search region, and the certificate gives the deterministic denominator scale used by the selector. Because the construction depends on fixed primitive records and finite modulus calls, the same bank can be transported across the common-experiment sequences described in Theorem 4.
The represented-data transducer in Algorithm 1 applies Definition 16 to the sample records . It first forms certified residual records, then evaluates the pilot-fold winding integrand on each bank circle, and finally evaluates the selected contour ratio on the evaluation fold. The fallback branches are part of the total rule: a failed admissibility check or a failed evaluation-fold modulus check returns the displayed value , while a successful run returns the clipped midpoint of a certified real interval. Thus the ordinary statistic and the represented execution have the same branch structure and the same clipped value whenever the compiled arithmetic implementation satisfies the bounded-build contract.
The final lemma is the deterministic accuracy statement used by the statistical risk proof. It compares the evaluation-fold empirical transforms with arbitrary population maps and on the selected circle, assumes the selected contour identifies a real target , and turns uniform transform errors into a bound for the clipped contour statistic.
Fix a parameter record as in Definition 1, the certified primitive and range records and of Definition 17 and Algorithm 1, a supplied treatment-regression code sequence, and observations with as in Definition 2. Let be the contour bank generated by and , let be its dyadic lower-modulus certificate, let have radius , and write . Let , let , and let be the target value.
Assume:
(Target range.) The target satisfies
(Selector event.) The canonical represented input built from the clipped current code and the sample satisfies and there are a bank index and an integer such that the certified selector of Algorithm 1 selects , the pilot winding output on decodes to , , the evaluation-fold modulus enclosure on has lower endpoint at least , and every satisfies where and are the split-fold empirical transforms of Definition 15.
(Contour integrability.) For every bank index selected by the canonical selector, both contour ratios are integrable on the circle .
(Exact identification.) For every bank index and every integer , if the canonical selector selects and the pilot winding output on decodes to , then
Then the total Borel contour statistic of Algorithm 1, computed on the same canonicalized sample, satisfies
⊢ LeanOn the selector event, choose the selected bank index and the decoded integer . Write for the selected circle and set The exact-identification hypothesis for the selected decoded contour gives .
For each , the selector event gives Together with , the triangle inequality yields the empirical denominator margin Since and , both denominators are nonzero on . Hence and the two terms have moduli bounded by and , respectively. Thus, pointwise on ,
The contour-integrability hypothesis for the selected index permits integration of this pointwise bound around the circle of radius . Therefore Multiplication by and taking real parts give
It remains to relate this contour value to the returned statistic. For the same selected and decoded , the construction of in Algorithm 1 uses the rational midpoint of the real coordinate of the certified evaluation rectangle. That rectangle contains the normalized empirical contour value whose real part is , and its real-coordinate width is at most . Hence The statistic is the clipping of this rational value to . Since , projection onto this interval is contractive relative to , so
By the bank construction in Definition 17, the selected radius satisfies , and the decoded count satisfies . Using and , Combining the preceding displays proves the asserted bound.
∎Lemma 13 is the deterministic bridge between certified execution and the risk calculation in Theorem 3. The selector event supplies a lower bound for the population denominator and an evaluation-fold lower-modulus check; the exact-identification clause supplies the population target represented by the same selected contour. The conclusion then bounds the statistic by the uniform denominator and numerator errors on the selected circle, with the final term coming from the certified interval radius used when the real contour moment is converted into the clipped rational midpoint.
Combining Definition 16, Definition 17, and Lemma 13 gives the executable correspondence needed by the main contour theorem. The arithmetic substrate provides sound finite enclosures, the bank handle supplies the fixed finite search geometry, and the perturbation lemma states the deterministic accuracy of the selected represented statistic on the good event used in the statistical argument.
Verification scope and crosswalk
This appendix records how the reader-facing statistical claims in the paper align with the formal statements displayed in the main text and technical appendices. The crosswalk is organized by mathematical role: model primitives and law classes, analytic identification, fixed-separation estimation, comparison results, local benchmarks, and certified execution.
The model primitives are fixed by Definitions 1 and 2. Sampling and structural restrictions are the assumptions in Assumptions 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, and 16. The law classes and risk criteria used in the statistical comparisons are the objects in Definitions 3, 4, 5, 6, 7, 8, 9, 10, and 11.
The analytic identification chain is the sequence Definition 12, Definition 13, Theorem 1, Lemma 1, and Theorem 2. These statements fix the treatment-innovation transform, the observable residual transforms, and the contour ratio that identifies the partially linear coefficient under the stated zero-free and positive-count conditions.
The fixed-separation estimation result is anchored by Lemma 3, Theorem 3, and Theorem 4. These statements provide the zero-localization radius, the finite-bank contour statistic, the minimax mean-squared-error scale on the fixed-code non-Gaussian class, the corresponding generalized-quantile bound, and the Gaussian-class diagnostic.
The comparison with ACE is recorded in Proposition 1, Lemma 5, and Proposition 2. The published order- ACE guarantee enters through the cited Jin–Mackey–Syrgkanis result (Jin et al., 2025); the displayed proposition states the exact imported bound and the class relations used to compare it with the contour guarantee.
The explicit mixture and local benchmark statements are Definition 14, Proposition 3, Lemma 7, and Theorem 5. The future-work formulation in Definition 15 and Remark 1 records the sample-level contour information and the triangular local-to-Gaussian adaptation target.
The certified execution layer is specified by Definition 16, Definition 17, and Algorithm 1. These statements connect the ordinary Borel statistic used in the statistical risk bounds with the finite represented-data construction based on the same primitive records.
Verification scope.
The theorem statements and proofs corresponding to the displayed formal environments are machine-checked in Lean 4 under their stated assumptions and citation interfaces. The ACE-alignment theorem in Proposition 1 verifies the class reduction, the application of the imported bound, and the comparison algebra conditional on the Jin–Mackey–Syrgkanis conclusion cited in its theorem-local footnote.
Proofs of the main results
Put as in Definition 7, and write as in Definition 12. From the definitions of and in Definitions 2 and 13, The definitions of in Definition 2 and above also give, by algebra, Membership in , together with the clipping convention in Definition 2, supplies the integrability and boundedness used below. Pulled back to the observation space, Definition 3 gives and the clipping gives . Hence The treatment and outcome sub-Gaussian clauses in Definition 3 imply real exponential moments of every order for and ; the displayed bound on then gives the corresponding real exponential moments for . These facts justify the differentiated transforms, conditional expectations, and products appearing below.
Since is analytic at and has analytic order there, Moreover, for each such , by differentiating the moment-generating function under the expectation. Hence, for any , This is the shifted annihilation identity.
We next show the conditional identity. Let be any measurable covariate event. Since , the binomial expansion gives The treatment-noise independence in Definition 3 factors each summand as and the first factor is . Thus the integral over every covariate event is zero, which characterizes
Define, for , Since , the same independence gives the factorization with as in Definition 13. Also, Differentiating times at yields All terms with vanish by the analytic-order calculation above, and the remaining term is Therefore
Finally, using the algebraic decomposition , The second term is zero because is covariate-measurable and the conditional identity above gives The third term is zero because is measurable with respect to , while Definition 3 gives : Thus Whenever , division gives
Fix . The definitions give Moreover, Definitions 2 and 3 give the pointwise clipping bound and the treatment-regression range bound almost surely, hence Together with the Luxemburg treatment-tail condition for , this bound gives every real exponential moment of , and hence the exponential integrability used below. Since is a function of and is independent of under Definition 3,
For the second identity, first note the pointwise residual decomposition because Definition 7 gives , while and . The preceding exponential-moment envelope for also gives differentiation under the expectation:
Thus The contamination weight is bounded almost surely by where the last line uses the outcome-regression and coefficient range bounds in Definition 3 together with the displayed bound on . This boundedness and the same exponential envelope for give integrability of the contamination term.
The outcome-noise term is integrable as well. The -tail condition in Definition 3 gives , and the learned-residual envelope gives By Cauchy–Schwarz, . Since , the real and imaginary parts of are measurable with respect to . Applying from Definition 3 separately to these integrable real products yields
Finally, independence of and factors the contamination term:
Substituting the last three displays and the derivative identity gives Since was arbitrary, both identities hold for every .
∎Write for the radius of the selected positively oriented bank circle . The bank construction gives . Put and By Lemma 10, The clipping defining , the range conditions in , and give Consequently and are complex differentiable at every point. Indeed, on the unit neighborhood of any , and so dominated differentiation under the integral applies to the weighted transforms. Separately, the sub-Gaussian exponential-moment condition for the treatment innovation , together with and the almost-sure bound on , gives for every real . The standard moment-generating-function analyticity criterion therefore gives analyticity of on a neighborhood of the closed disk .
On the closed disk , the nuisance zero-freeness assumption gives . Hence the quotient is well-defined on the closed disk, differentiable in the open disk, and continuous on the closed disk. Cauchy’s theorem on the circle gives
For , Lemma 1 gives where . Since and on , also on . Therefore, pointwise for , Integrating around and using the Cauchy identity for the quotient yields
The function is analytic near the closed disk and zero-free on , so the argument principle gives Thus The positive-count assumption gives , so . Dividing by this nonzero factor gives
By Definition 13, the right-hand side is . Therefore
∎Fix and put The square-exponential envelope first gives integrability of : for every outcome point, because . Hence is dominated by .
Consider first the case . For every real , Indeed, when , the Taylor-remainder bound gives . When and , one has , so . When and , one has , whence . With , Also, where the two inequalities use and . Combining these estimates with monotonicity of the exponential gives
Therefore Integrating and using and ,
It remains to treat . Completing the square gives After integration, Since and implies ,
The two cases prove the asserted bound for every .
∎Let From and , one has . The treatment-tail assumption gives together with the corresponding integrability, and the conditional-mean identity for gives Applying Lemma 2 to gives, for every real , Consequently, for every , The same exponential-integrability envelope gives analyticity of the complex MGF on all of , and normalization at the origin gives .
Suppose, toward a contradiction, that for every . Then is harmonic on the open disk . Choose an analytic function on this disk whose real part is , and set Then and, for ,
The Borel–Carathéodory estimate on the disk of radius , applied to the normalized function , gives on the sphere Cauchy’s estimate on the disk of radius therefore yields
Near the origin, , so takes values in the slit plane on a small disk and the principal logarithm is analytic there. The analytic function has identically zero real part on that small connected disk and vanishes at ; hence it is identically zero there. Thus Since is the real part of this derivative and ,
By the definition we have Substitution in the preceding bound gives
This contradicts the cumulant separation condition . Therefore there exists with and .
∎Put , , and The radius is nonnegative, since and by its displayed construction in Definition 17.
First, Indeed, ; the range bounds give and , while the Luxemburg bounds give and, after integrating the conditional bound, . Applying the elementary fourth-power bound to the three summands gives the display.
Next, This is the residual exponential envelope applied with radius . In detail, , where satisfies almost surely. Hence , and Young’s inequality at this radius gives The Luxemburg bound then yields the displayed exponential bound.
Therefore The pointwise inequality used here is and the preceding two displays bound the two expectations.
Let Then and . For , because . Hence Since , the claimed bound follows.
∎Define the primitive search radius where is the numerical factor in Definition 17. Let be the constant supplied by Lemma 9 for . For , put and define Thus . Set and Since , , and , one has . The two fixed-code converse arguments below give positive constants and , depending only on the same primitive tuple: comes from the fixed-code non-Gaussian converse, and comes from the separate fixed-code JMS ACE converse. We take Then and .
Fix the covariate measurable space, , the fixed records, , and the base model as in the statement. The transported-record hypothesis gives Let and denote the transported bank and range records. If is the lower-modulus certificate of the actual contour bank, then the width-one certified refinement containing and gives Consequently the actual selector constant with satisfies The monotonicity is exactly the monotonicity of the bank exponent, the map , and positive reciprocals.
The statistic is measurable. Indeed, is obtained by running the finite rational contour program of Algorithm 1 and then clipping its raw real output: On each finite branch trace, the queried observation endpoints are floor-dyadic functions of the sample coordinates, and the remaining operations are rational arithmetic and finite comparisons. The terminal traces are countable and exhaustive, so is Borel measurable; clipping by constants preserves measurability.
The represented execution clause follows from the canonical compilation contract. For every sample, the represented trace equals the ordinary trace, and the represented output record certifies The last identity is the elementary three-case equality used with the positive range record for .
The class relations are those of Lemma 5: for every model , and, if , The proof uses monotonicity of -norms on a probability space to pass from the published treatment-code condition to the treatment-code condition, and when the comparison subclass imposes the same code-radius requirement.
We record the pointwise risk bound used in both fixed-code classes. Let , assume , and assume where . For and every sample define Set On , the canonical selector packet supplies exactly the hypotheses of Lemma 13. The packet is obtained as follows. First, Lemma 8 gives a bank contour for the treatment-innovation transform with positive zero count and boundary modulus at least . By Lemma 10, the smallness condition gives on the search disk and preserves the zero count from the treatment-innovation transform to ; hence has positive -zero count, is zero-free on and inside the contour, and on its boundary.
The pilot-fold inequality , together with the certified modulus width and the certified winding enclosure in Algorithm 1, makes this population contour admissible: its pilot lower endpoint is at least , and its decoded pilot integer is the positive population zero count. The selected contour maximizes the pilot lower endpoint among admissible contours, so for the selected contour and decoded integer , while the evaluation-fold inequality gives the evaluation modulus margin required by the selector and The same certified node specifications give and Lemma 4 gives on . Continuity and the positive denominator margins give the two contour-integrability hypotheses for the empirical and population ratios. Finally, Theorem 2 applies to the selected contour, because is zero-free on , is zero-free on and inside , and the decoded integer is the positive population zero count; therefore the selected population contour ratio has real part .
Applying Lemma 13 with gives, on , Hence, on , The folds have cardinality at least , and Lemma 9 gives Markov’s inequality therefore yields Since the estimator and the target both lie in , on . Integrating the two displays gives
The quantile bound is the Markov consequence of the preceding mean-square bound. For the risk bound gives Equivalently, and the defining property of the generalized lower quantile in Definition 6 gives
It remains to identify the two lower constants. Put For a nonempty fixed-code non-Gaussian class, choose in that class. For , define The affine Gaussian perturbations and keep the law of , , and the supplied clipped codes from , and use the outcome equation with independent of . Since , and since Lemma 11 gives the required outcome-innovation Luxemburg envelope at scale , both and remain in the same fixed-code non-Gaussian class. Their target separation is The pair is a one-dimensional fixed-code path of the kind supplied by Lemma 12: the observed treatment law, the regression functions, and the clipped supplied codes are held fixed while the target moves by , and the divergence along such a path is quadratic in . Carrying out that Gaussian likelihood ratio calculation here, with , gives Le Cam’s two-point MSE inequality therefore supplies a constant , depending only on , such that every measurable estimator satisfies Thus the fixed-code non-Gaussian minimax risk is at least This is the fixed-code non-Gaussian converse.
For a nonempty fixed-code JMS ACE class, the separate fixed-code ACE converse applies the same two-point argument starting from a law in that JMS ACE class. The affine perturbations preserve JMS ACE membership, the target values and , the law of , the nuisance functions , the treatment innovation , and both clipped supplied codes; this is the JMS ACE branch of Lemma 12. The same product relative-entropy calculation and Le Cam inequality give a constant , depending only on the primitive tuple, such that
For the fixed-code non-Gaussian class, assume the hypotheses in the theorem. The lower bound from the preceding step and give The statistic is measurable by Step 3, so it is an admissible rule in the minimax infimum, and therefore For every in this fixed-code class, the current clipped treatment code agrees with the base clipped treatment code. Since the canonical inputs in Algorithm 1 use the current clipped treatment code, the statistic assembled from the base treatment-code sequence agrees pointwise with the statistic assembled from the class member’s treatment-code sequence. Applying Step 6 to the latter statistic and then using this pointwise equality gives The same pointwise equality reduces the absolute-error statistic in the generalized-quantile display to the statistic covered by Step 7, giving the displayed generalized-quantile bound over the same supremum.
For the fixed-code JMS ACE class, assume its nonemptiness, , and the same smallness condition on . The ACE lower bound from Step 8 and give The estimator is measurable, hence admissible, so the infimum is bounded above by its worst-case risk. By Step 5, every JMS ACE law lies in . The fixed-code restriction gives the same current clipped treatment code as the base model, so the pointwise congruence from Algorithm 1 identifies the statistic built from the base treatment-code sequence with the statistic built from the class member’s treatment-code sequence. Step 6 therefore yields The same pointwise equality and Step 7 give the asserted generalized-quantile bound on this class as well.
Fix the displayed constants and choose the constants and supplied by Theorem 3. Set Then , and its dependence is only through .
For the class relations, apply Lemma 5 at the displayed parameter record and the displayed sample size . It gives, law by law, and Since these pointwise implications are exactly the membership predicates for the displayed law classes, they yield The same result also gives, when , hence the asserted equality of the two classes under the common clipped-code convention.
Fix a model with observed-data law using the clipped supplied codes , and suppose , , and . If , the positive-radius part of the assumed published order- ACE generalized-quantile guarantee, with the objects and gates stated in Proposition 1, applied with the supplied handle and the same clipped codes, gives
It remains to treat the boundary case . For any , let be the parameter record obtained from by replacing the whole second radius sequence by and leaving every other coordinate unchanged. This sequence is again nonnegative and nonincreasing. Read the same observed-data law, regressions, and supplied code sequences as a model over ; its observed-data law is still . If , then the class membership is preserved: every condition in is unchanged except the outcome-code radius, and at index Moreover, with the denominator entering the second JMS eligibility quantity is unchanged: because and . All other eligibility quantities are unchanged by construction, so holds for at the same index.
Let The assumed ACE guarantee includes , and the displayed hypotheses give . For any real , set Then , and the positive-radius result applied to gives By the definition of , so the same quantile is at most . Since this holds for every , the desired boundary inequality follows:
For the spectral guarantee, assume the displayed nonemptiness condition and choose a model with observed-data law using the clipped supplied codes . Define the fixed-code ACE class This is the ACE-class set associated with the base model in Theorem 3. The statistic computed from the supplied treatment-code sequence agrees pointwise with the statistic computed from the base model’s supplied treatment-code sequence, because both constructions use the current treatment code only through the residuals The ACE-class part of Theorem 3, applied with the fixed primitive and range records, , and therefore yields Since is exactly the class of laws in the statement with the clipped supplied codes , this is the asserted spectral upper guarantee.
Finally consider a sequence satisfying the hypotheses in the last item of the proposition. Define and The assumed ACE guarantee includes , so . Along the sequence, the fixed-constant identities in the hypotheses rewrite the ACE bound as The second summand in brackets is nonnegative, hence eventually The assumed divergence is For , Once and , we have The right-hand side tends to , so the squeeze theorem gives
Apply Theorem 3 with the fixed constants It gives constants , depending only on these displayed constants, with the fixed-code non-Gaussian lower and upper bounds stated there. Also apply Proposition 2 on the chosen measurable covariate space. It gives, for every parameter record , writing for the sample size of , and every , and it gives the two zero Gaussian minimax conclusions for every nonempty clipped-code intersection at that same sample size.
Fix , the fixed primitive records, the sequence , the supplied measurable deterministic codes , and the displayed eventual non-Gaussian event. Write for the sample size of , and put For every sufficiently large , the displayed event in the theorem supplies a base model such that together with Let and be the supplied code sequences carried by . The two clipped-code equalities above imply equality of the two fixed-code law sets, The statistic also agrees pointwise when the supplied treatment code is replaced by : because Algorithm 1 forms the represented observation record with the current clipped values . Applying Theorem 3 to with base model , and using the i.i.d. product sampling law of that base model, yields the minimax bounds and, after the preceding law-set and statistic equalities, the estimator upper bound The measurability of follows from the same pointwise equality and the measurability conclusion of Theorem 3. The transported primitive bank record reuses the same certified names for , and . By Definition 17, the translated-dyadic contour-bank data are a deterministic function of those names and their error-one refinements. Thus, writing for the bank generated from and , the bank generated from and is also .
For the compiled statement, fix a compiled bounded spectral adapter satisfying the full canonical build-and-compilation specification for . The represented-data soundness implication for compiled bounded spectral adapters states that this full canonical build-and-compilation specification entails the represented-execution contract for the same parameter record, primitive bank record, range record, and treatment-code sequence. Its proof rewrites the pilot-boundary, winding-quadrature, and evaluation-quadrature entry points to the reference finite programs, identifies the compiled represented program with the ordinary finite-rational program on canonical represented inputs, and then applies the clipping certificate for . Therefore the compiled represented-data execution realizes with the same ordinary finite-rational result, trace, and certified clipped output required by the represented-execution clause of the theorem.
For the Gaussian part, fix a parameter record , write for its sample size, and fix . The first conclusion of Proposition 2 gives Now fix measurable deterministic codes whose clipped-code Gaussian intersection at this sample size is nonempty. The second conclusion of Proposition 2, applied to this same code pair, gives These are precisely the two Gaussian assertions.
Set and Choose Since and , this constant is positive and depends only on . Fix , , and the model objects satisfying the hypotheses. Define, on the observation space, so that .
The stipulated law gives the treatment-innovation transform explicitly. Indeed, for every , Therefore Near zero, The third derivative at zero of is zero, while Hence the fourth logarithmic derivative of at zero is , which is the asserted fourth cumulant: All differentiations are justified by the Gaussian mixture’s finite exponential moments.
The same transform computation gives the weighted characteristic moment used for the denominator. Since evaluating at yields Using , the independence of and , and the identity , Taking imaginary parts gives which is the displayed identity after substituting .
The treatment-code condition gives Since , Using , Combining this with the denominator identity gives
Define the residual sine score and the denominator sine score With the outcome equation gives The zero implies By independence of and , The conditional mean restriction for the outcome innovation gives Thus
The mixture second moment is bounded by The range condition for , together with clipping of , gives Since , Also, The Luxemburg sub-Gaussian condition in Assumption 9 gives Therefore
By Definition 14, the statistic is the clipped ratio with denominator score , remainder score , frequency , and threshold , because Let The preceding lower bound gives . Applying Lemma 6 pathwise and then averaging under the i.i.d. product law yields Using the two score bounds, Together with the identities and denominator lower bound proved above, this proves all claims.
Write Since ,
Consider first the branch . Then , and the unclipped ratio obeys The projection onto fixes , because , and is -Lipschitz. Hence Squaring and adding the nonnegative term involving gives
It remains to handle the branch , where . Since , Also , hence . Therefore The two branches give the claimed pathwise bound.
∎1. Fix . By Lemma 7, there is a constant , depending only on these four displayed constants, such that whenever an admissible parameter record has primitive entries equal to the displayed constants and cumulant order , every satisfying the Gaussian–Rademacher law assumption on , the sampling and model restrictions, and the treatment-code radius satisfies the full Gaussian–Rademacher path conclusion with constant . Applying this rule to the theorem’s records with and gives the first assertion.
2. Now fix a published ACE handle, , and , and assume that the handle satisfies the theorem’s published ACE interface at . Unpacking that hypothesis gives first It also gives the following specialization rule: for every parameter record whose probability-level coordinate is , every real , and every pair of strictly positive current code radii every supplied code pair and every model satisfying the clipped-code identities, the local ACE class conditions at , and the local JMS eligibility inequalities at obey For the theorem, take The standing assumption supplies the required positivity of . Together with the displayed strict positivity conditions and , the theorem’s clipped-code identities, local ACE class conditions, and local JMS eligibility inequalities are exactly the premises of the specialization rule at this . Substitution therefore yields where
∎Fix a model in the displayed Gaussian class, and write for its observed-data law. Define The conditional mean restrictions in Definitions 4 and 2 give and and are measurable with respect to . Hence conditional-expectation linearity yields The bounded-outcome condition gives -a.s. Conditioning the two inequalities on gives and therefore Conditioning the same bounded-outcome inequalities on , and using gives Combining the two displayed bounds,
Suppose, for contradiction, that , and define The previous step implies so The Gaussian membership in Definition 4 gives , and the parameter record in Definition 1 gives . Thus the variance is nonzero and For a Gaussian law with positive variance, Lebesgue measure is absolutely continuous with respect to that Gaussian law. Hence the last display implies But This contradiction proves
Now fix deterministic code sequences and assume the fixed-code Gaussian class in the statement is inhabited. Consider the estimator For every model in that fixed-code class, the first part gives , and hence Taking the supremum over the fixed-code class and then the infimum over all estimators in Definition 9 gives Since this risk is nonnegative by definition,
It remains to evaluate the generalized-quantile criterion for the same zero estimator. Since , For every model in the fixed-code Gaussian class, Thus the law of this loss is the point mass at zero. Its distribution function is for negative thresholds and at every nonnegative threshold, so the generalized lower -quantile of Definition 6 is Consequently the supremum of the fixed-code quantile loss for the zero estimator is , and the minimax infimum is at most . Nonnegativity in Definition 9 gives
Fix the parameter record, an index , and a model . For every code index , the clipped-code discrepancies are -measurable, because the supplied codes and the regressions are measurable and clipping preserves measurability. The covariate marginal is a probability law, since it is the image of the observed-data probability law under the covariate map.
Assume first that . The positive class-index condition and the structural, range, tail, and cumulant requirements of are the corresponding member conditions in Definition 5. The treatment-code requirement in Definition 3 follows from the JMS ACE treatment radius. Put The JMS ACE radius bound gives a finite bound for . Since is a probability law and , exponent monotonicity gives and Together with and the nonnegativity of at the positive index , this yields with the displayed integrand integrable. Thus .
Next assume . The inherited non-Gaussian member conditions supply the positive class-index condition and the common structural, range, tail, and cumulant requirements appearing in Definition 5. The subclass restrictions in Definition 11 give Because and is a probability law, exponent monotonicity gives These are the two JMS ACE code-radius requirements, so .
Finally suppose . The preceding paragraph gives Conversely, if , the first inclusion gives , and the two JMS ACE code bounds become exactly the two subclass bounds in Definition 11. Hence
∎Choose Then . Fix satisfying the assumptions, and write Since , and .
Independence of and gives the product transform The fourth logarithmic derivative at zero of this transform is the Rademacher contribution, hence
On the imaginary axis, The exponential factor is nonzero. Therefore the first positive zero occurs when , namely
For any ,
Let so that The derivative of the transform at the first zero is equivalently
Using , independence of and , and , Taking imaginary parts gives
Assume now that . The condition gives Since , Together with the preceding identity and , this yields
Define the two scores The zero-shift identity and the conditional mean condition imply Indeed, writing the covariate term is killed by the sine-shift identity after conditioning on , while Thus
The Gaussian–Rademacher path has the second-moment envelope With and , we have and therefore
Similarly, the range conditions give and the conditional Luxemburg bound for gives Consequently The two displayed square-moment bounds are exactly the components entering .
Let The estimator in the statement is the clipped ratio based on with threshold . Since , Lemma 6 gives the pathwise inequality Taking expectation under the i.i.d. product law and using gives
The clipping range also gives the deterministic bound Combining this with the preceding display and using yields
The smallness condition is equivalent to the stated radius bound. Since and we obtain
Finally, The first risk factor is bounded by the cumulant-parametrized factor: Substituting this inequality into the previous risk display gives
∎References
- Robinson, P. M. (1988). Root-\(N\)-consistent semiparametric regression. Econometrica. doi
- Chernozhukov, Victor and Chetverikov, Denis and Demirer, Mert and Duflo, Esther and Hansen, Christian and Newey, Whitney and Robins, James (2018). Double/debiased machine learning for treatment and structural parameters. The Econometrics Journal. doi
- Mackey, Lester and Syrgkanis, Vasilis and Zadik, Ilias (2018). Orthogonal Machine Learning: Power and Limitations. Proceedings of the 35th International Conference on Machine Learning. arXiv
- Jin, Jikai and Mackey, Lester and Syrgkanis, Vasilis (2025). It's Hard to Be Normal: The Impact of Noise on Structure-Agnostic Estimation. Advances in Neural Information Processing Systems. arXiv
- Dinh, Tien-Cuong and Ghosh, Subhroshekhar and Tran, Hoang-Son and Tran, Manh-Hung (2021). Gaussian fluctuations for spin systems and point processes: near-optimal rates via quantitative Marcinkiewicz's theorem. arXiv preprint. arXiv
- Mattner, Lutz (1992). Completeness of location families, translated moments, and uniqueness of charges. Probability Theory and Related Fields. doi
- Hu, Yingyao and Shiu, Ji-Liang (2022). A simple test of completeness in a class of nonparametric specification. Econometric Reviews. doi
- Newey, Whitney K. and Robins, James M. (2018). Cross-fitting and fast remainder rates for semiparametric estimation. arXiv preprint. arXiv
- Chacko, George and Viceira, Luis M. (2003). Spectral GMM estimation of continuous-time processes. Journal of Econometrics. doi
- Lee, Adam and Mesters, Geert (2024). Locally robust inference for non-Gaussian linear simultaneous equations models. Journal of Econometrics. doi
- Reizinger, Patrik and Mackey, Lester and Brendel, Wieland and Krishnan, Rahul G. (2025). Estimating Treatment Effects with Independent Component Analysis. arXiv preprint. arXiv
- Engle, Robert F. and Granger, C. W. J. and Rice, John and Weiss, Andrew (1986). Semiparametric estimates of the relation between weather and electricity sales. Journal of the American Statistical Association. doi
- Speckman, Paul (1988). Kernel smoothing in partial linear models. Journal of the Royal Statistical Society: Series B (Methodological).
- H{\"a}rdle, Wolfgang and Mammen, Enno (1993). Comparing nonparametric versus parametric regression fits. The Annals of Statistics. doi
- Härdle, Wolfgang and Liang, Hua and Gao, Jiti (2000). Partially Linear Models. Physica-Verlag. doi
- Hansen, Lars Peter (1982). Large sample properties of generalized method of moments estimators. Econometrica. doi
- Newey, Whitney K. and McFadden, Daniel (1994). Chapter 36 Large sample estimation and hypothesis testing. Handbook of Econometrics. doi
- Newey, Whitney K. (1990). Semiparametric efficiency bounds. Journal of Applied Econometrics. doi
- Newey, Whitney K. (1994). The asymptotic variance of semiparametric estimators. Econometrica. doi
- Bickel, Peter J. and Klaassen, Chris A. J. and Ritov, Ya'acov and Wellner, Jon A. (1993). Efficient and Adaptive Estimation for Semiparametric Models. Johns Hopkins University Press.
- {van der Vaart}, Aad W. (1998). Asymptotic Statistics. Cambridge University Press.
- van der Vaart, Aad W. and Wellner, Jon A. (1996). Weak Convergence and Empirical Processes. Springer. doi
- Robins, James M. and Rotnitzky, Andrea and Zhao, Lue Ping (1994). Estimation of regression coefficients when some regressors are not always observed. Journal of the American Statistical Association. doi
- {van der Laan}, Mark J. and Rubin, Daniel (2006). Targeted maximum likelihood learning. The International Journal of Biostatistics. doi
- Chernozhukov, Victor and Escanciano, Juan Carlos and Ichimura, Hidehiko and Newey, Whitney K. and Robins, James M. (2022). Locally robust semiparametric estimation. Econometrica. doi
- Belloni, Alexandre and Chernozhukov, Victor and Hansen, Christian (2014). Inference on treatment effects after selection among high-dimensional controls. Review of Economic Studies. doi
- Robins, James and Li, Lingling and Tchetgen, Eric and {van der Vaart}, Aad W. (2008). Higher order influence functions and minimax estimation of nonlinear functionals. Probability and Statistics: Essays in Honor of David A. Freedman. doi
- Lin Liu and Rajarshi Mukherjee and James Robins and Eric Tchetgen Tchetgen (2016). Adaptive Estimation of Nonparametric Functionals. Journal of Machine Learning Research. arXiv
- Liu, Lin and Mukherjee, Rajarshi and Newey, Whitney K. and Robins, James M. (2017). Semiparametric efficient empirical higher order influence function estimators. arXiv preprint. arXiv
- Feuerverger, Andrey and McDunnough, Philip (1981). On the efficiency of empirical characteristic function procedures. Journal of the Royal Statistical Society: Series B (Methodological).
- Singleton, Kenneth J. (2001). Estimation of affine asset pricing models using the empirical characteristic function. Journal of Econometrics. doi
- Carrasco, Marine and Florens, Jean-Pierre and Renault, Eric (2007). Chapter 77 Linear Inverse Problems in Structural Econometrics Estimation Based on Spectral Decomposition and Regularization. Handbook of Econometrics. doi
- Carrasco, Marine (2017). Efficient estimation using the characteristic function. Econometric Theory. doi
- Kagan, A. M. and Linnik, Yu. V. and Rao, C. R. (1973). Characterization Problems in Mathematical Statistics. Wiley.
- Michelen, Marcus and Sahasrabudhe, Julian (2019). Central limit theorems from the roots of probability generating functions. Advances in Mathematics. doi
- Eremenko, Alexandre and Fryntov, Alexander (2021). Stability in the Marcinkiewicz theorem. Zhurnal Matematicheskoi Fiziki, Analiza, Geometrii. doi
- D'Haultfoeuille, Xavier (2011). On the completeness condition in nonparametric instrumental problems. Econometric Theory. doi
- Darolles, Serge and Fan, Yanqin and Florens, Jean-Pierre and Renault, Eric (2011). Nonparametric Instrumental Regression. Econometrica. doi
- Comon, Pierre (1994). Independent component analysis, a new concept?. Signal Processing. doi
- Hyvärinen, Aapo and Oja, Erkki (1997). A Fast Fixed-Point Algorithm for Independent Component Analysis. Neural Computation. doi
- Hyvärinen, Aapo and Karhunen, Juha and Oja, Erkki (2001). Independent Component Analysis. Wiley. doi
- Lanne, Markku and Meitz, Mika and Saikkonen, Pentti (2017). Identification and estimation of non-Gaussian structural vector autoregressions. Journal of Econometrics. doi
- Hoesch, Lukas and Lee, Adam and Mesters, Geert (2024). Locally robust inference for non-Gaussian SVAR models. Quantitative Economics. doi
- Weihrauch, Klaus (2000). Computable Analysis. Springer. doi
- Ko, Ker-I (1991). Complexity Theory of Real Functions. Birkh{\"a}user. doi
- Moore, Ramon E. (1966). Interval Analysis. Prentice-Hall.
- Vershynin, Roman (2018). High-Dimensional Probability. Cambridge University Press. doi
- Wainwright, Martin J. (2019). High-Dimensional Statistics. Cambridge University Press. doi
- Marcinkiewicz, J{\'o}zef (1939). Sur une propriété de la loi de Gauß. Mathematische Zeitschrift. doi
- Andrews, Donald W.K. (2017). Examples of <mml:math xmlns:mml="http://www.w3.org/1998/Math/MathML" id="mml1" display="inline" overflow="scroll" altimg="si1.gif"><mml:msup><mml:mrow><mml:mi>L</mml:mi></mml:mrow><mml:mrow><mml:mn>2</mml:mn></mml:mrow></mml:msup></mml:math>-complete and boundedly-complete distributions. Journal of Econometrics. doi
Comments on earlier versions
Anchored to: