Uniform Expected Risk for Distance-Based Boundary Regression Designs
Abstract
This paper establishes two new unconditional logarithmic minimax lower bounds for distance-compressed boundary regression. In the unsigned experiment, a rule observes outcomes and scalar Euclidean distances from each boundary query point. Over the compact nonparametric law class in Definition 9, with and , both the CTY common-map risk in Definition 12 and the larger point-indexed outer risk in Definition 14 have minimax expected boundary sup-loss bounded below on the scale The support-boundary hypercube constructs separated boundary perturbations with matching scalar-distance information at the queried point, yielding the same lower-bound scale even when the rule may use point-indexed Borel sections. Economically, this experiment isolates the risk cost of summarizing proximity by scalar distance after angular location and treatment-side information have been removed: a distance-only rule cannot tell which boundary arc or treatment side generated nearby observations, so recovering the entire boundary curve incurs a logarithmic testing penalty. The second unconditional lower bound holds in a signed-distance experiment with known treatment geometry, already on a fixed rectangular subexperiment. Conditional on the three CTY-style analytic inputs collected in Definition 33, the stabilized local-polynomial estimator and that lower bound characterize the signed-distance expected outer-risk rate on the same scale for every fixed polynomial order , moment exponent , and envelope . Thus the unsigned contribution is lower-bound sharpness, while the signed two-sided frontier is explicitly conditional on .
Introduction
Regression discontinuity designs use discontinuities in treatment assignment to identify local causal effects near assignment thresholds. The classical scalar-cutoff design begins with the threshold assignment framework of Thistlethwaite et al. (1960) and the continuity-based identification analysis of Hahn et al. (2001); modern econometric treatments develop local comparisons, bandwidth selection, bias correction, and inference around that threshold (Imbens et al., 2008; Lee et al., 2010; Imbens et al., 2012; Calonico et al., 2014; Calonico et al., 2020; Armstrong et al., 2020; Kolesár et al., 2018; Imbens et al., 2019). Boundary and geographic RD designs extend the same logic to bivariate or multivariate scores, where treatment changes across a curve or interface rather than at a scalar cutoff (Keele et al., 2015; Keele et al., 2015; Keele et al., 2016; Papay et al., 2011; Reardon et al., 2012; Wong et al., 2013; Choi et al., 2018; Díaz et al., 2023; Sawada et al., 2024; Merlano, 2025; Kendall et al., 2026).
Distance-to-boundary reductions are central in this setting. They turn a two-dimensional local problem into a one-dimensional running-coordinate problem and are used both as empirical summaries of proximity to the boundary and as coordinates for local-polynomial fitting. Cattaneo et al. (2026) develop distance-based identification and estimation tools for boundary discontinuity designs, including local-polynomial methods for boundary treatment-effect curves. This paper studies the minimax expected-risk consequences of two precise distance experiments: an unsigned experiment in which angular and side information is compressed away, and a known-geometry signed experiment in which the rule retains the treatment side induced by the boundary.
The first experiment is an unsigned distance-compressed boundary regression problem. The law class in Definition 9 consists of compact bivariate random-design regression laws with bounded continuous design density, Hölder regression smoothness of order , and uniformly nondegenerate conditional variance. At a boundary query point , the data available to the rule are as defined in Definition 10. The target is the boundary regression function , and the loss is expected supremum error over .
For this unsigned experiment, Theorem 1 establishes that, for every integer and every , the common-map risk in Definition 12 is bounded below on the scale . Theorem 2 gives the same lower scale for the larger point-indexed class in Definition 13, evaluated with the outer-expectation risk in Definition 14. Thus the logarithmic lower bound is tied to the unsigned distance information set itself: even law-independent Borel sections chosen separately at each boundary point face the same order of uniform expected loss over .
The proof of the unsigned converse combines a direct-product testing inequality with a support-boundary packing. Lemma 1 converts many small coordinatewise overlaps into a lower bound on the probability of at least one decoding error. Lemma 2 builds many separated boundary points on a square support and assigns an independent binary regression perturbation to each local cell. The perturbations are smooth enough to remain in , their boundary values differ by order , and the unsigned distance laws at each coordinate have Kullback–Leibler divergence of logarithmic size. The support-boundary packing yields separated boundary signals of order and logarithmic compressed-distance KL control.
The second experiment is a known-geometry signed-distance boundary RD problem. The law class in Definition 22 is the displayed Euclidean-distance, uniform-kernel specialization of the CTY Assumptions 1–2 boundary design, with rectangular support, bounded continuous density, smooth potential-outcome regressions, continuous variances, a moment envelope, a rectifiable interior interface, two-sided local mass, and a population Gram lower bound. The rule observes signed-distance data from Definition 21 and may depend on the known geometry from Definition 20. The target is the interface treatment-effect curve in Definition 23.
For the signed-distance experiment, Theorem 3 proves a lower bound on the signed-distance minimax risk in Definition 26. For every fixed order , there is an envelope threshold such that, for every and , the risk is bounded below on the scale. The same lower bound is already present in the fixed rectangular subexperiment in Definition 27. The hard family in Lemma 6 places independent local perturbations along the bottom edge of a known treated rectangle and keeps all vertices inside the same known-geometry design.
The signed-distance upper bound uses an explicit local-polynomial rule. Definition 28 defines a degree- estimator that winsorizes outcomes at , inverts the empirical Gram matrix only on a uniform conditioning event, and clips the resulting treatment-effect estimate to . Under the distance-identification, first-order bias, and expected maximal inputs collected in Definition 33, Proposition 1 shows that this estimator with bandwidth has expected outer risk bounded by a constant multiple of uniformly over .
Under , the signed-distance upper bound combines with the lower bound to give the conditional frontier in Theorem 4. For every fixed , , and , the signed-distance minimax risk satisfies The same theorem records that the winsorized, Gram-stabilized signed-distance local-polynomial estimator attains the upper side of this bound under the same three analytic inputs. A short reproducibility note records the checked scope for the anchored mathematical statements in Section D.
The comparison between the two experiments is organized by law class, information set, and risk. In the unsigned problem, scalar Euclidean distances from each query point create a boundary-indexed compression in which many angularly separated alternatives remain hard to distinguish, producing unconditional lower bounds. In the signed problem, known boundary geometry and treatment side supply the one-dimensional coordinate used by local-polynomial RD methods; its fixed-geometry lower bound is unconditional, while its two-sided rate characterization uses the displayed analytic inputs. The common normalization therefore records a shared logarithmic lower scale, with a conditional upper side for the signed experiment.
The remainder of the paper proceeds as follows. Section 3 defines the compact nonparametric unsigned distance experiment, its decision classes, and its minimax risks. Section 4 defines the known-geometry signed-distance RD experiment, the A1/A2 law class, the treatment-effect target, and the stabilized local-polynomial estimator. Section 5 states the unsigned lower bounds, the signed-distance lower bound, the conditional upper bound, and the matched signed-distance frontier. Section 6 discusses the interpretation for boundary RD designs. Sections A, B, C, D, and E contain the direct-product and packing arguments, the signed-distance empirical-process ingredients, the fixed-geometry signed hypercube, the reproducibility note, and deferred proofs.
Related Literature
Regression discontinuity designs originate in the threshold assignment framework of Thistlethwaite et al. (1960) and the identification analysis of Hahn et al. (2001), with modern econometric treatments emphasizing local comparisons, bandwidth choice, and inferential refinements (Imbens et al., 2008; Lee et al., 2010; Imbens et al., 2012). Geographic and boundary designs extend this logic to assignment regions in space, where the score is vector-valued and the estimand is indexed by an interface rather than a scalar cutoff (Keele et al., 2015; Keele et al., 2015; Keele et al., 2016; Papay et al., 2011; Reardon et al., 2012; Wong et al., 2013; Choi et al., 2018; Díaz et al., 2023; Sawada et al., 2024; Merlano, 2025). Within that literature, Cattaneo et al. (2026) develop distance-based identification and estimation tools for boundary regression designs. Their analysis establishes identification from distance-to-boundary reductions, sequence-scoped bias expansions, and probability-order stochastic control for local-polynomial procedures under their maintained design conditions.
Theorem 6 of Cattaneo et al. (2026) studies the unsigned minimax risk where one law-independent measurable map is used at every boundary point. Their theorem proves a positive lower bound after multiplication by . The printed conjecture immediately following that theorem replaces by , matching the scale of their uniform upper bound. The unsigned common-map result here proves that conjectured strengthening without changing the law class, transformed data, target, decision class, or expected supremum loss. The point-indexed result changes only the decision rule: a law-independent Borel section may now be selected separately for every query point, while the transformed data and regression target remain the same; outer expectation supplies the measurable envelope for the boundary supremum. It proves the same logarithmic lower scale for this larger class.
The signed result concerns a different econometric target and information set. Motivated by CTY’s signed-distance identification and local-polynomial analysis, it studies the boundary treatment-effect curve when the assignment geometry and treatment side are available to the rule. The paper proves a new minimax lower bound already on one fixed known rectangular geometry. Its expected-risk upper bound is stated under the displayed identification, bias, and expected maximal hypotheses. Table 1 records these three deliverables in parallel.
| Deliverable | Nearest CTY result | Exact increment in this paper |
|---|---|---|
| Unsigned CTY decision problem | Theorem 6 of Cattaneo et al. (2026): , data , one common map , boundary regression target, and expected boundary sup-loss; lower scale . | Keeps the law class, information, target, decision class, and risk fixed and proves the conjectured lower scale . |
| Point-indexed unsigned rules | CTY’s lower-bound theorem uses one law-independent common map across query points. | Allows a separate law-independent Borel section for each query point, uses outer expectation for the boundary supremum, and still proves the lower scale. |
| Signed known-geometry treatment effects | Theorems 1–2 and Supplemental SA-8.3–SA-8.4 provide the nearest identification, bias, and stochastic-process ingredients for signed-distance local-polynomial analysis. | Changes the target to the boundary treatment-effect curve and the information to signed distance with known geometry; proves an unconditional lower bound on a fixed rectangular subexperiment and, under , a stabilized expected outer-risk upper bound at the same scale. |
This paper studies the uniform expected-risk consequences of those distance reductions. In the unsigned experiment, the relevant object is the compact bivariate nonparametric regression class in Definition 9, where a rule observes outcomes and unsigned Euclidean distances as in Definition 10. The decision problem uses the common-map class in Definition 11 and its completed-expectation risk in Definition 12. The resulting converse is a same-class statement: over the law class and information set just described, the minimax expected boundary sup-loss has the logarithmic scale introduced in Definition 8. The point-indexed extension uses the sectionwise-Borel outer-risk setup in Definitions 13 and 14, thereby characterizing the same compression barrier for rules that may choose a law-independent Borel section separately at each boundary point.
This focus differs from the usual role of distance in geographic RD applications. Empirical distance reductions often serve as low-dimensional summaries of proximity to a boundary or as practical coordinates for local fitting (Keele et al., 2015; Keele et al., 2015; Keele et al., 2016; Kendall et al., 2026). Here the distance reduction is part of the statistical experiment: the rule receives the outcome and the scalar unsigned distance from the query point, and the risk is the expected supremum over boundary points. The lower bound therefore speaks to information loss under the compressed observation scheme, rather than to specification choice in a geographic RD regression. Related diagnostic and design tools, including manipulation tests and geographic balance analyses (McCrary, 2008; Cattaneo et al., 2022; Cattaneo et al., 2019; Cattaneo et al., 2023; Cattaneo et al., 2025), address complementary features of RD credibility; the present comparison is organized around the target, the available data transformation, and the uniform loss criterion.
The signed-distance part connects to local-polynomial boundary estimation. Local-polynomial methods provide the standard approximation architecture for nonparametric regression and RD estimation (Fan et al., 1996; Ruppert et al., 1994; Calonico et al., 2014; Calonico et al., 2020; Imbens et al., 2019). Robust bias correction and optimal inference results refine confidence interval construction and coverage for scalar or boundary-local targets (Calonico et al., 2014; Calonico et al., 2020; Armstrong et al., 2020; Kolesár et al., 2018). The known-geometry signed-distance experiment fixes the CTY Assumptions 1--2 potential-outcome boundary design through in Definition 22, observes signed-distance data in Definition 21, and evaluates treatment-effect rules in Definitions 25 and 26. Assuming the CTY distance-identification, uniform first-order bias, and expected maximal inputs collected as in Definition 33, the paper combines the signed-distance lower bound with the winsorized, Gram-stabilized local-polynomial upper bound to obtain the two-sided expected-risk rate.
The minimax structure draws on classical nonparametric lower-bound and empirical-process ideas. Packing and modulus arguments underlie the rate calculations in nonparametric estimation (Stone, 1982; Tsybakov, 2009), while boundary and set-indexed empirical-process tools support the stochastic controls used in local smoothing problems (Mammen et al., 1995; Cuevas et al., 2004; Vapnik et al., 1971; Pollard, 1984; van der Vaart et al., 1996). The lower-bound construction in this paper adapts those ideas to boundary-indexed distance compression: many separated boundary locations carry independent binary perturbations, while scalar distance observations constrain the coordinatewise information available to any admissible rule. The empirical-process component enters the signed-distance upper result through the stated analytic inputs and the bounded-envelope maximal inequality used for the winsorized score. Together, these ingredients place the paper between geographic RD methodology and minimax nonparametric theory: the estimand is a boundary treatment-effect or boundary regression curve, the information set is a distance-compressed experiment, and the criterion is uniform expected loss over the displayed law classes.
Setup: Distance-Compressed Boundary Regression
This section fixes the statistical experiment used for the unsigned lower bounds. The object of interest is a boundary-indexed regression function under a compact bivariate random-design law, and the information available at a boundary point is the outcome together with the scalar Euclidean distance from that point. The normalization used throughout the converse arguments is introduced first.
For a random element taking values in a measurable space under a probability law , write for the distribution of ; when is clear, write . For an event with , write for the conditional distribution satisfying for every .
For a real symmetric matrix , we write , equivalently means for every vector .
For and , we write for the closed Euclidean ball .
We say when is on , satisfies for every and every , and its st Fréchet derivative is -Lipschitz on : for all .
We say when is a probability law of with square-integrable outcome, whose carried objects , , , and are tied to the law as follows: is the topological support of the -marginal, the -marginal has Lebesgue density , is an -marginal-almost-everywhere version of , and is an -marginal-almost-everywhere version of .
For a law of on , we say is the bivariate design support when , the topological support of the -marginal under .
For the two probability laws and on the common measurable sample space of , define their total variation distance by where the supremum ranges over all measurable sets in that sample space.
These preliminary conventions give the basic probabilistic and geometric vocabulary. In particular, the law-space convention ties the support, density, regression, and conditional variance to the same underlying law, while the Hölder and ball notation provide the local smoothness and Euclidean localization language used in the packing arguments. Total variation enters later through the direct-product comparison of coordinatewise compressed experiments.
The frontier normalization is
⊢ LeanThe scale in Definition 8 is the common normalization for the unsigned boundary lower bounds. Its logarithmic factor reflects the uniform boundary supremum loss rather than a pointwise loss criterion.
The nonparametric class now fixes the bivariate random-design model. It combines compact support, a bounded continuous design density, Hölder regression smoothness, and uniformly nondegenerate conditional variance.
For and , the class consists of all random-design regression laws satisfying the following conditions.
(Parameter regime.) and .
(Sampling.) is a probability law for one observation , is square-integrable under , and for every , are sampled independently and identically from .
(Design.) The -marginal has Lebesgue density , where is the exact topological support of the -marginal. Write for this density. It is continuous on , is compact and satisfies , and for every . The boundary admits a Lipschitz parametrization by .
(Regression.) The function is an -marginal-almost-everywhere version of and belongs to .
(Variance.) The function is an -marginal-almost-everywhere version of , is continuous on , and satisfies for every .
Equivalently,
⊢ LeanThe class in Definition 9 is a compact-support version of a random-design nonparametric regression model. The density bounds keep local sample sizes comparable across admissible laws, the Lipschitz boundary condition gives a one-dimensional interface over which a supremum can be evaluated, and the Hölder condition calibrates how quickly regression values may change near separated boundary locations. These conditions are standard in nonparametric minimax analysis, with compactness and smoothness playing the same organizational role as in classical regression lower bounds (Stone, 1982; Tsybakov, 2009).
At a boundary query point, the unsigned experiment compresses each covariate to its distance from that point. The rule retains the outcome and the scalar distance, so all angular information around the query point is absorbed into this one-dimensional summary.
For , the distance data are
⊢ LeanThe observation in Definition 10 is indexed by the boundary point because the same sample is re-expressed relative to each query location. This formulation matches the distance-to-boundary reductions studied by Cattaneo et al. (2026), while the present risk criterion evaluates the resulting boundary-indexed regression rule uniformly over .
The first decision class uses one law-independent map for all query points. This common-map restriction captures procedures that process the distance data by a single measurable rule once the sample size is fixed.
For a represented rule, write for its selected law-independent measurable map. Define the sample-size-indexed class by
⊢ LeanThe common-map minimax risk is
⊢ LeanTogether, Definitions 11 and 12 define the completed-expectation minimax problem for common-map rules. The expectation is the ordinary expectation under the product law generated by , and the loss is the supremum error along the support boundary. The infimum therefore ranges over estimators that share the same measurable distance-data transformation across all admissible laws and all boundary points.
For comparison, the point-indexed decision class lets the measurable section depend on the query point while remaining law independent. This keeps the unsigned distance information set fixed and enlarges the admissible class by allowing a separate Borel rule at each .
The point-indexed decision class is This class consists exactly of such law-independent fixed- Borel sections.
⊢ LeanFor an extended nonnegative function on the sample space, define the outer expectation by The point-indexed distance risk is
⊢ LeanThe outer expectation in Definition 14 follows the standard empirical-process convention for suprema that may require measurable majorants (van der Vaart et al., 1996). With this convention, Definitions 13 and 14 define a sectionwise Borel problem that preserves the same distance-compressed observations as Definition 10. The main lower bounds compare these two unsigned risks after the common notation and signed-distance setup have been fixed.
Setup: Known-Geometry Signed-Distance Designs
The second experiment keeps the boundary geometry available to the rule and records which side of the interface each nearby observation occupies. This is the setting in which local-polynomial methods use distance as a one-dimensional coordinate while preserving the sign induced by treatment assignment, as in boundary regression designs developed by Cattaneo et al. (2026). We first fix the assignment, interface, basis, and coordinate conventions that underlie the signed-distance reduction.
We say is the arm- assignment region of , for , when and are Borel subsets of satisfying and , with deterministic treatment assignment .
We say is the known treatment interface for the fixed hypercube geometry when it is the boundary of the treated rectangle,
For an integer and , define .
We say is the score-coordinate projection when, for every potential-outcome observation , .
The partition in Definition 15 makes treatment deterministic given location, so the side of the boundary is a feature of the design rather than an additional random assignment variable. The fixed interface in Definition 16 is the rectangular geometry used in the lower-bound construction, while Definition 17 supplies the signed-distance polynomial basis used by the local fits. The projection in Definition 18 records that the score coordinate is the Euclidean covariate.
The known-geometry object packages the support, assignment regions, interface, Euclidean metric, and smoothing weight. The canonical smoothing weight is the one-dimensional uniform kernel , which selects observations whose signed distance lies within the bandwidth window.
The fixed uniform kernel is
⊢ LeanDefine to be the class of admissible known geometries , where are a Borel partition of , For a member law , write for its assignment regions and write . Define its known design object by
⊢ LeanFor , define the signed distance and the signed-distance sample
⊢ LeanThe known design object in Definition 20 gives the rule the relevant boundary and side labels. For an interface point , the signed distance in Definition 21 retains the Euclidean distance magnitude and assigns its sign according to the treatment region. Thus the signed-distance sample provides the one-dimensional running coordinate used by local-polynomial RD estimators (Fan et al., 1996; Ruppert et al., 1994; Calonico et al., 2014; Calonico et al., 2020; Imbens et al., 2019).
The law class below is the displayed Euclidean, uniform-kernel specialization of the CTY Assumptions 1–2 boundary design. Its clauses have four statistical roles. The support, assignment regions, interface, metric, and signed coordinate define the experiment and information available to a rule. The density, regression smoothness, conditional variance, and moment bounds provide uniform design and outcome envelopes. The population-Gram, arm-mass, distance-slice, and VC conditions keep local-polynomial fitting uniformly conditioned along the interface. The separate inputs , introduced in Definition 33, supply identification, first-order approximation, and expected uniform stochastic control for the conditional upper bound.
For integers , real numbers , and , let be the set of laws satisfying the following conditions.
Write for the sampled potential outcomes of observation , and write for the one-dimensional Hausdorff measure of the common interface.
(Sampling and consistency.) For every , the vectors , , are independently and identically distributed,
(Rectangular support and density.) The support has the form The variable has a Lebesgue density , continuous on , such that
(Smooth potential-outcome regressions.) For each , the regression has a extension to an open neighborhood of , and
(Conditional variances.) For each , the conditional variance is continuous and satisfies
(Moment envelope.) For each and every ,
(Assignment geometry.) The sets and form a Borel partition of , and their known common interface is a compact rectifiable curve satisfying
(Metric, kernel, and VC bound.) The known metric and kernel are The collection is a VC class with VC index at most four.
(Population Gram lower bound.) Put and For every , every , and both ,
(Small-bandwidth arm mass.) For every , every , and both ,
(Distance-slice density.) For every , both , and every , the one-dimensional Hausdorff integral of over is finite and strictly positive.
This class is the displayed -uniformized Euclidean-distance, uniform-kernel specialization of the two baseline smoothness and design conditions, augmented by the displayed envelope restrictions and quantitative small-bandwidth thresholds.
⊢ LeanThe assumptions in Definition 22 fall into three groups. First, the support, assignment partition, rectifiable interface, signed coordinate, and distance-slice conditions define the design and the information observed by the signed-distance procedure. Second, the density, moments, smoothness, variance, arm-mass, population-Gram, and VC conditions provide uniform outcome, conditioning, and complexity envelopes for local fitting. Third, in Definition 33 collects the identification, bias, and expected maximal inputs invoked specifically by the conditional upper bound. The unconditional signed lower bound uses the fixed-geometry subexperiment inside the first two groups; the stabilized estimator and conditional frontier use all three groups.
The potential outcomes and deterministic treatment in Definition 22 give the boundary target its causal RD interpretation. The arm-specific regressions support degree- signed-distance approximation, while the arm-specific variances and moment envelope keep the outcome scale uniform. Geometrically, the class includes compact rectifiable interfaces of controlled length inside rectangular supports and requires two-sided local mass and stable signed-distance Gram matrices at every interface point. These quantitative conditions also cover the corners of the fixed treated rectangle: Lemma 6 verifies the Gram and arm-mass thresholds there rather than relying on a smooth-boundary shortcut.
The treatment-effect target is the jump between the two smooth potential-outcome regression surfaces at the interface. The population local-polynomial coefficient records the signed-distance normal equations that the estimator approximates.
For a law and an interface point , define the treatment-effect curve by , where is the continuous conditional-mean version for arm .
We say is the unwinsorized population degree- signed-distance local-polynomial coefficient when satisfies the normal equations
The curve in Definition 23 is evaluated uniformly over . The coefficient in Definition 24 is a population projection onto the signed-distance polynomial basis; its intercept difference is the population analogue of the treatment-effect estimator, and its approximation error is the first-order bias component isolated later.
Rules in the signed-distance experiment may depend on the known geometry and the query point through law-independent Borel sections. The risk then uses the outer expectation already introduced for point-indexed suprema in Definition 14.
Let be the class of rules for which there is one collection selected independently of the unknown member law, such that every fixed section is a Borel map from to and, for every , every , and every , The imposed regularity is fixed- Borel measurability of each section.
⊢ LeanUsing the outer expectation defined in Definition 14, for the signed-distance minimax risk is
⊢ LeanWe define as the signed-distance outer minimax risk with the law supremum restricted to the fixed hard geometry:
The decision class in Definition 25 gives procedures access to the known geometry while keeping the section family fixed before the member law is chosen. The signed-distance risk in Definition 26 measures the expected worst interface error for treatment-effect estimation. The fixed-geometry risk in Definition 27 is the same criterion restricted to the rectangular hard family, which is the family used for the signed-distance converse in Theorem 3.
It remains to name the estimator used for the upper bound. The construction follows the standard local-polynomial template, with three stabilizing features: outcomes are winsorized at a bandwidth-dependent level, the empirical Gram matrix is inverted only when it has a uniform quadratic-form lower bound, and the final treatment-effect estimate is clipped to the natural envelope scale.
For , define the winsorization map by For each , query point , and bandwidth , define With , define The stabilized coefficient is With , the clipped signed-distance local-polynomial treatment-effect estimator is
⊢ LeanThe winsorization map in Definition 28 converts the moment envelope into a bounded-score component at bandwidth . The empirical Gram matrix estimates the population Gram , and the fallback rule keeps the coefficient map stable on samples where the empirical design is ill conditioned. The resulting estimator is the concrete rule used in the signed-distance upper result.
The following three deviation objects separate approximation, design, and score variation. They are stated here because the main theorem collects analytic inputs in terms of these quantities before applying the estimator in Definition 28.
For , define the uniform first-order bias ratio by
For a law and bandwidth , we write for the extended nonnegative sample-space random variable
Define as the sample-dependent uniform centered local-polynomial score deviation with the norm taken in .
We say is the winsorized local-polynomial residual score when, for , sample size , arm , interface point , bandwidth , winsorization level , and coordinate , it is the -random variablewith the unwinsorized population local-polynomial coefficient at .
The bias ratio in Definition 29 measures the uniform first-order approximation error of the population signed-distance fit. The Gram deviation in Definition 30 tracks the uniform empirical conditioning of the local design, and the raw-score deviation in Definition 31 records the centered stochastic part before winsorization. The winsorized score in Definition 32 is the bounded-envelope score component used by the maximal inequality in the upper-bound appendix. These objects place the later risk result in the usual bias, design-stability, and empirical-process decomposition for local-polynomial RD estimators (Fan et al., 1996; Ruppert et al., 1994; Calonico et al., 2014; Calonico et al., 2020; van der Vaart et al., 1996).
Main Results
The results now put the two distance-compressed experiments on a common risk scale. The unsigned problem uses the nonparametric regression class and distance data introduced in Definitions 9 and 10; the signed problem uses the known-geometry law class, signed-distance data, and treatment-effect risk in Definitions 22, 21, and 26. Throughout the section, the normalizing sequence is the rate in Definition 8.
The first result is the same-class converse for the unsigned common-map experiment. It evaluates exactly the decision class in Definition 11 under the completed-expectation risk in Definition 12.
For every integer and every , there exists a constant such that, for the common-map minimax risk in Definition 12 and ,
⊢ LeanThe theorem establishes that a single law-independent distance rule incurs expected boundary sup-loss at least of logarithmic order over . The packing calculation behind the rate is easiest to read from the separated boundary construction used in the appendix. If the signal amplitude is , then the construction places active locations along the boundary and uses cells with area scale . Unsigned distance compression leaves each coordinate with information of order , so the many-coordinate testing barrier is calibrated by The smoothness order affects the packing constants and the geometric calibration of , while the exponent is governed by the fourth-order distance-compression information calculation. In the signed-distance local-polynomial results below, the polynomial order likewise affects constants and regularity thresholds while the displayed rate normalization remains .
The next result carries the same logarithmic converse to the point-indexed unsigned class. The point-indexed class in Definition 13 permits a separate Borel section at each query point, and the risk is evaluated with outer expectation as in Definition 14.
For every and satisfying and , there exists a constant such that, with the point-indexed outer-expectation minimax risk in Definition 14 and ,
⊢ LeanThus the lower-bound scale is a property of the unsigned distance information itself, even when the rule may choose its measurable section point by point. Outer expectation in this statement is the same measurable-majorant criterion introduced in Definition 14, which is the standard way to state uniform risks over potentially nonseparable boundary index sets (van der Vaart et al., 1996).
The signed-distance lower bound uses the known-geometry experiment in Definitions 25 and 26. It is stated for all local-polynomial orders after an envelope threshold that depends on the order.
For every nonnegative integer , there exists a real number such that:
(Envelope threshold.) .
(Moment exponent.) .
(Uniform envelope.) .
(Risk definition.) is the signed-distance minimax risk in Definition 26.
(Fixed-geometry risk.) is the infimum over the same known-geometry point-indexed decision rules of the same worst-case outer-expected interface sup-loss, with the law supremum further restricted to laws in satisfying the fixed hard-geometry condition.
For every and satisfying these conditions, there exists a constant such that, writing , and
⊢ LeanTheorem 3 establishes the same logarithmic lower scale for the known-geometry signed-distance minimax risk. The second inequality records that the lower bound is already present inside the fixed rectangular geometry from Definition 27. Consequently, the converse is calibrated by a concrete Euclidean boundary experiment rather than by variation over arbitrary interfaces.
The upper result for the signed-distance experiment combines the estimator in Definition 28 with three analytic inputs. These inputs are distance identification, uniform first-order bias control, and expected maximal inequalities; every signed-distance upper or matched-rate statement below is conditioned on this collection.
Define as the following three signed-distance analytic inputs used by the expected-risk upper bound.
(CTY distance identification.) For every A1/A2 law and metric , the support is rectangular, the density is continuous and positive on the support, both potential outcomes are integrable, both arm regressions have extensions, the boundary is a positive-length rectifiable curve, is a metric uniformly equivalent to Euclidean distance on the support, and every sufficiently small positive armwise distance slice at every boundary point has finite positive density mass. Under these conditions there is one source-coherent selected conditional law of the observed outcome given signed distance, with armwise mean versions whose one-sided limits at zero satisfy
(Uniform first-order bias.) There is , depending only on , such that for every and every positive antitone deterministic sequence with ,
(Expected maximal bounds.) There is such that, along every positive deterministic satisfying all sufficiently large and every satisfy and
The three components of are statistical assumptions with distinct roles. Identification connects the one-sided signed-distance conditional-mean limits to the boundary treatment-effect curve (Cattaneo et al., 2026). Approximation bounds the population local-polynomial intercept error uniformly at first order in the bandwidth (Fan et al., 1996; Ruppert et al., 1994). Stochastic control bounds the expected uniform Gram and score deviations over arms, boundary points, and laws (Vapnik et al., 1971; Pollard, 1984; van der Vaart et al., 1996). The next proposition is conditional on all three displayed inputs and derives the expected-risk rate for the stabilized estimator from them.
Let be an integer, let , and let . Assume:
(Distance identification.) For every A1/A2 law , every distance map , and every satisfying the CTY identification assumptions at order , there are signed-distance conditional-mean versions and one-sided limits and such that
(Uniform first-order bias.) There is a constant , depending on and , such that for every and every positive antitone deterministic bandwidth sequence with and ,
(Expected maximal bounds.) There is a constant such that, for every positive deterministic bandwidth sequence with and the following bounds hold for all sufficiently large , uniformly over : and
With as in Definition 8, take the bandwidth and the winsorization level in the winsorized, Gram-stabilized signed-distance local-polynomial estimator clipped to . Then there are constants and such that, for every ,
⊢ LeanThe proposition gives a concrete signed-distance procedure with expected outer risk bounded by a constant multiple of . The bandwidth choice balances the first-order signed-distance bias, of order , with the uniform stochastic scale . At , both terms have order , and the additional heavy-tail term in the assumed maximal bound is controlled under the displayed moment condition. The winsorized-score lemma used in the appendix supplies the bounded-envelope score component needed for this decomposition. The full proposition additionally uses the assumed uniform Gram control and the assumed raw-score, law-uniform, heavy-tail maximal bound with exponent , so the displayed expected-risk conclusion is tied to the full collection in Definition 33.
Conditional signed-distance frontier.
Under the analytic inputs in Definition 33, the signed-distance lower and upper statements combine to yield the matched outer-expected rate over .
For every integer , there exists a constant such that, for every and every , the following conditions imply the frontier conclusion below:
(Distance identification.) For every A1/A2 law , every map satisfying the CTY identification assumptions for and , and every , there are selected source-coherent signed-distance conditional-mean versions whose right and left limits at zero exist and identify
(Uniform first-order bias.) There is a constant depending only on and such that, uniformly over every and every positive antitone deterministic bandwidth sequence with , the finite limsup of the normalized uniform first-order bias ratio is bounded by that constant.
(Expected maximal bounds.) For this , there is a constant such that, along every positive bandwidth sequence with for all sufficiently large and every , the expected outer Gram-deviation supremum satisfies and the expected outer centered raw-score supremum satisfies
Then there exist constants and with such that, for the signed-distance minimax risk of Definition 26 and , Moreover, the explicit winsorized, Gram-stabilized degree- signed-distance local-polynomial estimator with bandwidth and winsorization level , clipped to , has uniform outer risk over satisfying
⊢ LeanUnder the analytic inputs in Definition 33, Theorem 4 characterizes the signed-distance known-geometry minimax risk up to constants under the same envelope threshold as the lower bound. The lower inequality is supplied by Theorem 3, while the upper inequality is attained by the explicit winsorized, Gram-stabilized local-polynomial estimator from Definition 28. The result places the signed-distance experiment on the logarithmic expected outer-risk scale for every fixed polynomial order , moment exponent , and envelope above the stated order-dependent threshold, with the upper side governed by the three analytic inputs.
Discussion and Extensions
The preceding results separate two roles of distance in boundary regression designs. For the unsigned experiment, Theorems 1 and 2 establish logarithmic lower bounds for boundary sup-loss when the rule observes only outcomes and scalar Euclidean distances from the query point. The common-map statement applies to the decision class and completed-expectation risk in Definitions 11 and 12, while the point-indexed statement applies to the sectionwise Borel class and outer-expectation risk in Definitions 13 and 14. These conclusions identify an information barrier created by unsigned distance compression over the compact bivariate law class in Definition 9.
The signed-distance experiment has a different information set. There the rule observes signed distances generated by the known geometry in Definitions 20 and 21 and is evaluated over the A1/A2 law class and risk in Definitions 22 and 26. Assuming the analytic inputs collected in Definition 33, Proposition 1 and Theorem 4 combine the signed-distance lower bound with the stabilized local-polynomial upper argument to obtain a two-sided expected outer-risk rate. The role of Cattaneo et al. (2026) is therefore most direct in the signed design: their distance-identification framework supplies the population interpretation of the one-dimensional signed-distance coordinate, while the conditional result here packages the displayed inputs with the new lower bound and stabilized estimator.
The common normalization in Definition 8 indexes different delivered objects across the two experiments. The unsigned experiment has unconditional common-map and point-indexed lower bounds: many separated boundary perturbations remain difficult because scalar distance hides their angular locations and treatment sides. The signed experiment has an unconditional fixed-geometry lower bound and, under the three inputs in Definition 33, a conditional upper bound attained by a procedure that uses the treatment side, known boundary geometry, and local-polynomial structure. Accordingly, the shared normalization compares lower scales; a two-sided rate conclusion is available here for the signed experiment conditional on .
The signed-distance result is stated for the known Euclidean geometry in Definition 20 and the one-dimensional uniform kernel in Definition 19. The law class in Definition 22 imposes rectangular support, bounded continuous density, smooth potential-outcome regressions, continuous variances, a moment envelope, rectifiable interior interface, two-sided small-bandwidth mass, and a population Gram lower bound. These conditions describe the boundary designs for which the signed-distance local-polynomial estimator has uniformly controlled bias, stable local design matrices, and expected stochastic fluctuations. They also align with the geographic RD setting in which boundary location and treatment side are part of the design information (Keele et al., 2015; Keele et al., 2015; Keele et al., 2016; Kendall et al., 2026).
The Euclidean and envelope restrictions also clarify the intended empirical interpretation. Geographic RD applications often use distance to a boundary as a practical running coordinate while retaining design knowledge about assignment regions and local geography (Papay et al., 2011; Reardon et al., 2012; Wong et al., 2013; Choi et al., 2018; Díaz et al., 2023; Sawada et al., 2024; Merlano, 2025). The signed-distance analysis formalizes that information set through the known-geometry class, whereas the unsigned analysis formalizes a more compressed experiment in which angular and side information are absent from the observation available at a query point. The results therefore organize comparisons by the data transformation and target risk, rather than by a universal ranking of distance-based empirical strategies.
Open questions.
The unsigned results in Theorems 1 and 2 are lower bounds only and establish no matching unsigned upper bound here. The signed-distance matched rate in Theorem 4 is conditional on , so a fully self-contained derivation of the signed-distance upper ingredients within the same displayed law class remains an open direction. Further extensions include non-Euclidean metrics, kernels beyond , alternative interface regularity classes, and risk criteria adapted to inference rather than uniform expected loss.
Appendices
Direct-Product and Support-Boundary Lower Bound
This appendix records the finite-experiment ingredients used to derive the unsigned logarithmic lower bounds. The argument has two components. The first is an overlap product inequality for many binary coordinates after each coordinate has been compressed. The second is a support-boundary hypercube inside the nonparametric class of Definition 9, constructed so that the boundary regression values differ at many separated points while the corresponding unsigned distance experiments remain close coordinate by coordinate. These ingredients yield the common-map converse in Theorem 1; after the strict inclusion statement below, the same construction yields the point-indexed converse in Theorem 2.
The direct-product step is stated abstractly because the same form is also useful for the signed-distance lower experiment. It converts many small coordinatewise overlaps into a lower bound on the probability of at least one decoding error.
Fix an integer . Assume the following conditions.
(Spaces.) The spaces carrying are standard Borel.
(Binary coordinates.) The variables are independent and uniformly distributed on .
(Conditional product structure.) Conditionally on , the variables are independent, the law of depends only on for , and the law of is independent of .
(Compressed coordinate decoding.) For each , for a measurable map . The decoder is a measurable -valued function of and is unchanged when its raw coordinate vector is altered only in coordinate while is fixed. Define
Then every such decentralized decoder satisfies Moreover, for every real , if then
Write for the measurable space carrying , and let be the measurable map representing the th decoder, so that and so that the invariance hypothesis reads: for every , every , and all raw vectors with for every , The two laws compared at coordinate are the compressed conditional laws of Definition 1, and with the total variation distance of Definition 7. Since , we have for every .
Step 1: a coordinatewise coupling whose compressions agree with probability . The raw bit-conditional law at coordinate is , and is its image under . Let be the common part of and , that is, the measure on with density with respect to . Its total mass is because , both densities integrate to one, and half the distance between them is the total variation distance.
Each residual and then has total mass . When , the maximal coupling of and is the law on given by and when it is the image of under . In both cases its two marginals are and , and its diagonal carries mass at least .
Because all the spaces are standard Borel, the regular conditional distributions of given exist under both raw bit-conditional laws; feeding the two coordinates of the maximal coupling into them produces a law on pairs of raw coordinate blocks, whose coordinates we write , such that The first two equalities hold because the lift preserves each marginal, and the third because the image of under is exactly the maximal coupling above, whose diagonal has mass at least .
Step 2: the coupled experiment reproduces the original one at every vertex. Let be the product of the coordinatewise couplings, that is, the law of the array determined by and let , which by the conditional product structure does not depend on . Write for the law under which has law and is independent of the array. For a vertex define the selected raw vector by Since the two marginals of are the two raw bit-conditional laws at coordinate , the selection map pushes forward to the product over of , so that which is precisely the conditional product law assumed in the statement.
Consequently, writing for the simultaneous-correctness event at the vertex , the uniform hypercube prior gives
Step 3: the total-disagreement event. Let be the event that every coupled coordinate has unequal compressed values. It is a product event across coordinates, so Step 1 gives each factor being the complement of an event of probability at least .
Step 4: a pointwise bound on the number of correctly decoded vertices. Fix one realization of and of the coupled array . We claim On this is immediate: each of the indicators is at most one, and the right-hand side equals .
Off there is a coordinate with . Let be with coordinate flipped. Selection at a coordinate reads the same bit under and , while at coordinate the two selected blocks have equal compressions by the choice of ; hence The invariance hypothesis applied to these two raw vectors therefore gives so this single value cannot equal both and , and consequently
Summing this over all and using that is an involution of , hence a bijection under which the sum is unchanged, that is, , which is the claim off .
Step 5: the simultaneous success bound. Integrating the pointwise bound of Step 4 against and using Step 3, Dividing by and inserting the identity of Step 2 gives the first assertion,
Step 6: calibrating the product by the Kullback–Leibler budget. Fix a real and assume Since we have , and we split on the sign of .
If , then , so and the budget reads . The Bretagnolle–Huber overlap bound applied to the pair states so monotonicity of together with the budget gives
Every factor is nonnegative, so the elementary inequality applied coordinatewise and then the common floor just obtained yield the final equality because .
If , then , so and the budget forces , hence and , that is, for every . Since , the product below has at least one factor, so
Step 7: passing to the error probability. Write and . Then , because every factor is a total variation distance; , because ; and by Step 6. Hence so subtracting the success bound of Step 5 from one gives
Finally, along any sequence with fixed, and hence , so the displayed lower bound is .
∎The invariance requirement in Lemma 1 matches the distance-compressed experiment: when a rule estimates the value attached to coordinate , the information about the th perturbation that remains after compression is summarized by , while the remaining raw coordinates and the background block may be held fixed. The total-variation overlap coefficient measures how often the compressed coordinate laws can be coupled to agree. The Kullback–Leibler consequence gives the logarithmic calibration used in the boundary packing, in the same testing spirit as standard minimax lower-bound arguments (Stone, 1982; Tsybakov, 2009).
The next statement supplies the geometric and probabilistic certificate for the unsigned law class. It places many separated locations on the boundary of a square support and assigns one binary regression perturbation to each local half-disc.
For every integer and every envelope , there exist constants such that, for all sufficiently large , writing , there are an integer , a radius , a mass , boundary points , laws indexed by , and values such that:
(Packing size and scale.)
(Boundary packing.) With , each lies on , and for , The cells are pairwise disjoint.
(Model class and support.) For every , the law belongs to and has covariate support .
(Cell masses.) For every and every ,
(Local dependence on one bit.) If , then the restricted laws of on the cell agree under and . The restricted laws of on are the same for all .
(Boundary signal.) For every and ,
(Radial invariance.) Let be with coordinate flipped. For every and , the laws of under and agree. Moreover, the corresponding one-point distance laws agree after restriction to radii .
(Compressed-distance KL bound.) For every and ,
Fix and , and write .
Step 1: the fixed profiles. Let be the normalized radial bump, that is, the smooth radially symmetric function with and let be the smooth transition function, which vanishes on and equals one on . For each integer , fix a nonnegative derivative envelope satisfying These envelopes are finite fixed constants attached to the bump profile.
Step 2: the constants. Define the smoothness-dependent derivative scale and the one-observation divergence constant by and set Since every , the scale is finite and satisfies Hence all six constants are positive and .
The three displayed inequalities give the divergence budget used below: the first step because and , the second because .
Step 3: the eventual regime. With as in Definition 8, set
Because , for all sufficiently large we have , , and
Moreover, since , while gives ; hence, enlarging once more,
Finally, for every , because makes .
Fix any in this eventual regime; all remaining claims are verified at that .
Step 4: the packing points and cells. Take the equispaced grid on the middle half of the lower edge of , and let as in the statement, with the closed Euclidean ball of Definition 3. Each has second coordinate and first coordinate of modulus strictly below , so ; and since , each is the closed upper half-disc , where .
From the definition of , . The small-radius bound also gives , and therefore Thus the grid spacing obeys so that for and the closed cells , of radius around centers at distance at least , are disjoint.
The radius bounds hold with equality since . For the packing size, forces , and for , so
Step 5: the laws . Fix . Define the regression profile on by
and the design density on by where the radial tilt is and is a unit vector of (the bump is radial, so the choice is immaterial).
Let be the joint law on observation coordinates obtained as follows. Its -marginal is the measure with Lebesgue density , and at score the outcome kernel is the probability measure Equivalently, is the image under of the product-kernel measure generated by this -marginal and this outcome kernel. The declared regression and variance profiles of the resulting law are They are -marginal-almost-everywhere versions of the conditional mean and variance in the sense of Definition 9.
Since on , the bumps have disjoint supports , and , so the mixture weight is legitimate and the clipping of to needed off , and to used in Step 10, are both silent on .
Step 6: model class, support, and cell masses. The angular tilt obeys uniformly, and the bumps attached to distinct centers have disjoint supports, so whence for ; is continuous on , and is compact with and with Lipschitz-parametrized by . The conditional variance is continuous and lies in . Finally, for , using , and ; since the th derivative of the localized bump is bounded by , these bounds give on for together with -Lipschitz continuity of , that is in the sense of Definition 4. Consequently as in Definition 9, with .
For the cell masses, the only -dependent part of on is , and reflecting in the vertical line through changes the sign of while fixing ; hence that term integrates to zero over and where is common to all because the cells are translates of each other along the lower edge.
Step 7: locality in one bit. The functions and vanish outside , and the cells are pairwise disjoint. Therefore, on both and depend on only through , so and on we have and , so
Step 8: boundary signal. Put At every bump attached to another center vanishes, because for , while ; hence
Step 9: radial invariance. Let be with coordinate flipped. For every measurable , the reflection of Step 6 applied to the annular slice again cancels the -dependent density term, so
For the tail statement, fix . The last display of Step 3 gives , so and ; therefore the cutoff is fully active,
Let be measurable and let be the part of the cell on which the two vertices differ. Assume , so that ; the reverse case is symmetric. On the difference of the two outcome-weighted densities is because on the cell all other bumps and angular corrections vanish.
Splitting the brace into the radius-only part and the remainder , the first part contributes by the same horizontal reflection, and the second contributes
Since the average of over each half-circle equals , the two remaining terms cancel:
Together with the equality of radial slice masses this says that, for every measurable , the two experiments assign the same mass and the same conditional-mean mass to the slice ; since the outcome mixture is determined by its conditional mean, the restricted one-point distance laws agree:
Step 10: the compressed-distance divergence. Below the fully active radius the same cancellation is replaced by a localized setwise bound. For every measurable , where denotes the covariate marginal: outside the radii the cutoff is fully active and Step 9 gives exact cancellation, while on those radii the uncancelled increment is bounded by the bump amplitude per unit radial mass.
Moreover, since and a disc of radius has area ,
The complete radial laws of Step 9 provide a common marginal law for under the two adjacent vertices. Disintegrating the one-observation laws with respect to this common radial marginal gives measurable radial success parameters, clipped to and equal almost everywhere to the corresponding conditional Bernoulli weights, such that the preceding setwise bound implies Here the factor is the envelope used for the almost-everywhere Radon–Nikodym comparison, and the middle-half clipping is silent by Step 5. For two Bernoulli-plus-Gaussian mixtures with success parameters in , the divergence at a fixed radius is at most four times the squared difference of the success parameters. Averaging over the common radial marginal therefore gives using .
The distance data of Definition 10 are independent copies of the one-observation pair , so the divergence tensorizes; combining with the eventual budget of Step 3,
Steps 4–10 supply every displayed clause with the constants of Step 2, for all sufficiently large .
∎Lemma 2 is the concrete bridge between smooth boundary regression and the abstract product experiment. The Hölder restriction determines the relation between the signal height and the cell radius , while the boundary packing supplies separated query points. Local dependence on one bit makes each cell a coordinate of the hypercube, and the radial invariance clause expresses the angular information loss created by unsigned distance observations. The compressed-distance Kullback–Leibler bound then puts those coordinates in the range of Lemma 1. This is the finite-sample testing structure that underlies the rate calculation summarized in the main text.
For common-map rules, the decoder associated with a proposed estimator applies the same measurable map to each query point. If the estimator were uniformly accurate on all ’s, the signs of the boundary signals in Lemma 2 would be decoded correctly across all coordinates. Combining the signal separation with the overlap lower bound gives the positive lower scale in Theorem 1. The argument uses the law class and completed-expectation criterion exactly as defined in Definitions 9, 11, and 12.
The point-indexed class enlarges the set of admissible rules by allowing the Borel section to vary with the query point. The following inclusion statement records that relationship before applying the same hypercube construction to the outer-expectation risk.
For parameters satisfying
(Sample size.) .
(Smoothness index.) .
(Envelope.) .
the common-map decision class and the point-indexed decision class of Definitions 11 and 13 satisfy
⊢ LeanThroughout, a point of is written , and a sample point is written , so that the distance data of Definition 10 evaluated at and at a query point are
Inclusion. Let and let be a measurable map, selected independently of , representing it as in Definition 11. Define a point-indexed family by Each section is the single Borel map , hence Borel measurable for every fixed , and the family is selected independently of because is. Moreover, for every , every sample point, and every , so in the sense of Definition 13. Therefore
A point-indexed rule outside the common-map class. Let the rule return the first coordinate of the query point, ignoring the data: For each fixed the section is constant in , hence Borel, and the family does not depend on the member law. Thus .
Suppose, for contradiction, that this same rule also belongs to , with common measurable map . Apply Lemma 2 at the given and , at any sample size beyond the threshold it provides, and take the vertex ; its model-class and support clause supplies a law
Set Each of and has second coordinate and first coordinate of modulus , so it is a non-corner point of the lower edge of and therefore
Take the deterministic sample point whose every observation equals , that is and for all . Since , and share their second coordinate, the Euclidean distances reduce to horizontal ones:
Consequently the compressed samples at the two query points coincide:
Both and lie on , so the assumed common-map representation applies at each of them and, with the displayed equality of distance data, gives But and , a contradiction.
Hence , so the inclusion proved first is proper:
∎The strict inclusion in Lemma 3 places the two unsigned risks in the expected order of admissible decision classes. For the converse, the relevant fact is that each point-indexed section remains law independent and receives the same unsigned distance data from Definition 10. Thus the coordinatewise testing reduction induced by Lemma 2 continues to apply after replacing a single common map by fixed- Borel sections. The outer expectation in Definition 14 supplies the measurable-majorant convention for the boundary supremum, and the resulting lower bound is the point-indexed logarithmic converse in Theorem 2.
Signed-Distance Geometry, Maximal Inequalities, and Upper Bound
This appendix collects the empirical-process ingredients used by the signed-distance upper bound in Proposition 1. The objects are those of the known-geometry design in Definitions 20, 21, 22, and 28: local neighborhoods are Euclidean balls generated by the uniform weight , the outcome moment envelope is handled through winsorization, and the stochastic term is measured by an outer-expected supremum over arms and interface points. The two statements below organize these roles before the estimator bound is assembled in Proposition 1.
The first ingredient is geometric. Because the local windows are closed Euclidean balls, the relevant index class has finite VC complexity, giving the covering control used by the maximal inequality (Vapnik et al., 1971; Pollard, 1984; van der Vaart et al., 1996).
For every finite set of score points in , if , then is not shattered by the class of closed Euclidean balls Equivalently, no such has the property that every subset of can be written as for some center and radius .
⊢ LeanSuppose, for contradiction, that a finite set with is shattered by closed Euclidean balls, that is, for every subset there are a center and a radius with Since we may fix four pairwise distinct points .
Writing , put
An algebraic identity. Let be arbitrary coefficients and let be the corresponding combination, For a center and a radius , expanding the squared distance gives, for each , which is one and the same linear combination of the four entries of , with coefficients not depending on . Multiplying by and summing therefore replaces those entries by the entries of :
A sign contradiction. We next record the contradiction used in both cases below. Suppose the coefficients are not all zero, that their combination satisfies , and that Reading off the first entry of and using for every ,
so, some being nonzero, at least one coefficient is strictly negative; fix an index with .
Apply the shattering hypothesis to the subset obtaining a center and a radius with . Because are distinct, holds exactly when . Hence every term of the weighted sum is nonpositive, if then , so , the bracket is nonpositive, and the product is nonpositive; if then , so , the bracket is strictly positive, and again the product is nonpositive.
At the second alternative occurs with a strictly negative coefficient, so
and summing the four terms gives which contradicts the assumed nonnegativity at this .
Case 1: are linearly independent. Four linearly independent vectors in form a basis, so the fourth standard basis vector has an expansion with coefficients that are not all zero, since . Here , so and , and the algebraic identity gives, for every and every , The sign contradiction applies.
Case 2: are linearly dependent. Choose coefficients , not all zero, with Here , so in particular , and the algebraic identity gives, for every and every , The sign contradiction applies again.
Both cases are impossible, so no finite with is shattered by closed Euclidean balls.
∎The bound in Lemma 4 is used with the ball notation introduced in Definition 3. In the signed-distance design, the event selected by the uniform kernel is a side-restricted Euclidean neighborhood of the query point. Finite VC complexity therefore converts the continuum of possible interface centers into an empirical-process class with logarithmic entropy, matching the factors in the risk calculation.
The second auxiliary statement supplies the expected maximal bound for the centered winsorized local-polynomial score. It combines the VC geometry of Lemma 4 with the bounded envelope created by winsorizing the outcome at level , using the outer-expectation convention already built into Definition 26.
Fix a nonnegative integer and a real number . There is a constant such that the following holds for every real .
(Law class.) The law belongs to the A1/A2 class of Definition 22.
(Sample size.) The sample size satisfies .
(Bandwidth and clipping.) The bandwidth and winsorization level satisfy and .
Let denote the -fold i.i.d. law of the potential-outcome observations. For each and , let the winsorized centered local-polynomial score be the vector whose th coordinate is where the empirical score uses the uniform kernel, the polynomial basis , bandwidth , and the clipped outcome transform , centered around the corresponding population score. Then
⊢ LeanThroughout, denotes one potential-outcome observation, its observed outcome, and for we write for the coordinate maximum, which is the norm appearing in the displayed conclusion. All constants below are fixed before the moment exponent , the law , the sample size , the bandwidth , and the winsorization level .
Step 1: a coefficient radius depending only on and . Set which is positive. We claim that for every , every , every , every and every , Write for the right-hand side of the population normal equations in Definition 24. Its coordinates obey Indeed, the uniform kernel forces , hence and , so the integrand is dominated pointwise by . The density envelope of Definition 22 and the area of a planar disk give The moment envelope in Definition 22 gives, since , the conditional second-moment estimate using . Localizing this bound to gives for each arm, so the displayed coordinate bound follows from
The population Gram lower bound of Definition 22 and the normal equations of Definition 24, written with , then give the first step because is one of the summands. Hence , which is the claim.
Step 2: the bounded score class. For let be its componentwise clipping, and let For define the scalar function of one observation Only the kernel window uses the enlarged bandwidth ; the polynomial arguments keep , so the target score is the member with . The half-open enlargement supplies the support slack used in Step 6, and clipping the coefficient gives a constant-envelope class.
Step 3: the score is a member of the class. Fix , and , and put . By Step 1 the clipping is inactive at this index, , so comparing with the definition of the winsorized score in Definition 32,
The envelope bound of Step 5 makes integrable, so averaging the previous display over the independent and identically distributed observations gives
Step 4: reduction to one scalar empirical process. For each fixed sample the coordinate maximum is attained at some , and every pair with together with that produces an index of by Step 3. Hence, pointwise on the sample space,
Outer expectation is monotone and pulls out the positive finite constant , so
Step 5: envelope, radius and covering witnesses. Put There are constants and , depending only on and and therefore only on and , such that for every real , every and all with , the class satisfies the following three properties.
(Envelope.) For every and every observation , On the kernel window , so for every . Together with and , this gives where the middle inequality uses .
(Population radius.) For every , On the almost-sure support of , the squared separable score is bounded pointwise by The kernel window gives the radius because . The probability of this disk is at most , while the two localized arm-square integrals are each bounded by , using the conditional second-moment estimate from Step 1. Hence is bounded by Taking square roots and using for gives the displayed bound. The winsorization level is absorbed through , so this radius is uniform in .
(Polynomial covering.) For every non-empty finite list of observations and every there is a finite set with and with . The certificate is obtained as follows. First, the fixed-dimensional radial residual-polynomial class with bounded response , bounded arm weights, and coefficient box has polynomial empirical covers with envelope ; this is applied with radial ranges for the positive and negative arm pieces and for the zero-distance and outside-support pieces. Second, Lemma 4 bounds the shattering of the planar closed-ball class, so the classifier family indexed by the center-radius parameter has Vapnik–Chervonenkis dimension at most three; the standard passage from a finite Vapnik–Chervonenkis dimension to empirical covering numbers (Pollard, 1984; van der Vaart et al., 1996) then supplies a fixed polynomial cover for that classifier family. Product closure combines each radial piece with the ball classifier for the positive and negative pieces, and finite-sum closure combines the four pieces On the negative arm the factors are absorbed into the clipped coefficient vector, and the zero-distance and outside-support terms are degree-zero radial residual terms. A pullback then ties the ball center, radial center, arm, coefficient vector and coordinate to the same index . These finitely many product, sum, envelope-enlargement and pullback operations preserve polynomial empirical entropy, changing only the constants; after the final envelope enlargement from to , the witnesses are the displayed and .
The three witnesses , , are chosen before , , and , which is what makes the final constant uniform in the moment exponent.
The parameter-regime clause of gives , so and ; the certificate therefore applies at the bandwidth and clipping level of the statement.
Step 6: countable reduction. Equip with the product topology inherited from , the discrete arm and coordinate sets, the relative topology on , and the Euclidean topology on . Choose a fixed countable dense sequence in this space. Let The set has -probability one: almost surely, and because the interface has finite one-dimensional Hausdorff measure by Definition 22 and hence vanishing planar Lebesgue measure under the displayed density.
Fix . For set and consider the target point . By density of , choose so that satisfies with and . Then , and the one-sided construction gives Thus the approximating support bandwidth converges to the target bandwidth while staying above it by enough to cover the center perturbation. For every , this slack preserves the closed-kernel decision when , including the endpoint case ; when , ordinary continuity of distance and of makes the approximating kernel decision eventually agree with the target one. Since , the assignment partition in Definition 22 fixes the signed-arm indicator along the approximating boundary centers, and continuity handles the polynomial and clipped-coefficient factors. Therefore The class is dominated by the integrable envelope so dominated convergence also gives convergence of the corresponding population integrals. Consequently, for every , almost surely under ,
Step 7: the variance-adaptive maximal inequality. We use the following countable-reduction form of the variance-adaptive VC maximal inequality. Let be a probability law and a real-valued class. Suppose that , , , every is measurable, , , every countable subfamily has empirical covers of size at most at radius for , and a countable subfamily realizes the centered empirical supremum almost surely under every finite product law. Then, with one has, for every , Indeed, choose a reducing sequence . The envelope makes the countable supremum measurable, integrable, and bounded by , so the outer expectation of the continuum supremum is bounded by the ordinary expectation of the countable supremum. Countable-class symmetrization bounds that expectation by twice the Rademacher complexity. The Rademacher bound applies Dudley chaining to the empirical polynomial covers, uses the population radius , and controls the empirical radius by symmetrizing the squared bounded class; the resulting quadratic inequality yields the displayed leading -term and second-order -term, with the numerical constant rounded to .
Apply this inequality with , , , , and with the entropy and countability witnesses from Steps 5 and 6. Writing we obtain, for every ,
Step 8: calibrating the logarithm. Since , and , Also gives , so and the maximum in is attained at the second entry. Putting , because .
Step 9: conclusion. Abbreviate both nonnegative. Multiplying the two terms of Step 7 by and using Step 8, Therefore, chaining Steps 4, 6, 7 and the two displays above and setting we obtain The constant is positive and, since , , and were fixed from and alone, depends only on and .
∎The two terms in Lemma 5 have the usual variance and envelope forms for bounded empirical processes indexed by VC-type sets (Pollard, 1984; van der Vaart et al., 1996). The scaling reflects the two-dimensional local window around an interface point, while the supremum over and matches the signed-distance risk criterion in Definition 26. Together with the Gram lower bound in Definition 22, the statement controls the stochastic perturbation of the intercepts in the stabilized estimator.
These ingredients enter the upper-bound calculation as follows. The population part is separated into the first-order local-polynomial approximation term and the winsorization term, the latter controlled by the conditional moment envelope in Definition 22, which bounds the winsorization remainder by a multiple of . The sample part is decomposed into empirical Gram fluctuation and centered score fluctuation, with the score controlled by Lemma 5 and the Gram term handled by the expected maximal condition stated in Proposition 1. The stabilization rule in Definition 28 converts these uniform matrix and score bounds into a uniform intercept bound for each arm, and clipping keeps the treatment-effect estimate on the envelope scale.
At the bandwidth from Definition 8 and winsorization level , the first-order bias, winsorization bias, and stochastic terms in Proposition 1 are all bounded by a constant multiple of for sufficiently large . Thus the auxiliary results in this appendix support the expected outer-risk upper bound for the explicit signed-distance local-polynomial estimator, completing the upper side of the matched-rate comparison stated in Theorem 4.
Signed-Distance Hypercube and Matched Rate
This appendix records the hard signed-distance subexperiment used for the lower side of the matched-rate statement. The construction fixes the known rectangular geometry from Definition 27 and places many separated local perturbations along the bottom edge of the treatment boundary. Each perturbation changes the treatment-effect target at one boundary point by a calibrated amount, while the corresponding one-observation signed-distance laws remain close enough for the overlap direct-product argument in Lemma 1 to apply.
For every integer , there is a constant such that, for every and every , there are constants , with when , and a constant such that the following construction is available for every . Writing , there exist an integer , constants , points , laws indexed by , and probability laws for and , satisfying:
(Size and calibration.)
(Common geometry.) The support is , the treated region is , the control region is , the boundary is , the metric is Euclidean distance, and the kernel is . The points lie on the middle of the bottom edge, and the disks are contained in , pairwise separated by distance at least , and pairwise disjoint.
(Class membership.) For every , the law belongs to with the common geometry above.
(Bit structure.) For every and , The restriction of to depends on only through , and the restriction of outside is common across all .
(Signed-observation laws.) Conditional on , the one-observation signed-distance law at under is .
(Signal separation.) If two vertices agree in coordinate , then their treatment-effect targets at agree. Flipping coordinate changes the treatment-effect target at by exactly :
(Local comparison laws.) For each , the laws and have the same signed-distance marginal, agree outside the event , and obey the two Kullback–Leibler bounds
Throughout, a point of is written , and we abbreviate . The arm sign attached to the side intervals , of Definition 22 is written , as in the arm-mass clause there.
Two fixed smooth profiles. Fix once and for all a radially symmetric function of class with let be its radial profile, so that , and put
Fix also a nondecreasing function with for and for .
The constants. Set All four are positive, and because the bracket is positive, and, since every is nonnegative and appears in the bracketed sum,
Because , the required order-zero separation holds; it in fact holds for every .
The small-separation threshold. The map is continuous at with value , so there is such that every satisfies
The envelope threshold. The signed radial polynomial energy is coercive: there is such that, for both and every , with the basis of Definition 17.
Put so that .
Fix and ; in particular . Fix also .
Common geometry and grid. Take the fixed known geometry of Definition 20, with the uniform kernel of Definition 19; the interface is the one displayed in Definition 16. With put the balls being those of Definition 3.
Each has second coordinate and first coordinate of modulus at most , so it lies on the middle of the bottom edge of the treated rectangle, hence on .
Since gives and therefore and ,
The centers are equispaced with gap , so for using , which is where and are used.
Consequently the disks are pairwise disjoint, and each is contained in because and .
The vertex laws. For let be the law with the following two ingredients. First, has Lebesgue density supported on , given there by where the radial tilt and the horizontal direction cosine at are
Second, given the two potential outcomes are conditionally independent Bernoulli variables whose means are the arm regressions of Definition 22, both truncated to off , with and .
Because the disks are -separated while vanishes outside the unit ball, at most one summand is nonzero at any ; together with and this gives so the truncation is inactive on .
For the same reason the angular sum has modulus at most , whence
Class membership. We check the clauses of Definition 22 in turn, using throughout.
The support is the square , and is continuous on with
The control regression is constant, and the treatment regression is the restriction to of the globally function which satisfies the displayed envelope
The bump terms are controlled because the bandwidth is calibrated to the amplitude: for every the th derivative of one bump is bounded by , and since and give , while for ; at the bound is . The affine part contributes at most to the derivative maximum, because on and its only nonvanishing derivative is the constant , and at most to the Lipschitz maximum; and at most one bump is active at a time.
The conditional variances are the Bernoulli variances of the two means, so and they are continuous on .
Each takes values in , so the moment envelope holds with room to spare:
The sets are a Borel partition of with , and is a compact rectifiable curve, being the boundary of a square. Its size obeys the lower bound because contains the segment from to and the upper bound from a Lipschitz parametrization of the perimeter.
The metric is Euclidean and the kernel is ; the traces of closed Euclidean balls on form a VC class of index at most four by Lemma 4.
For the population Gram matrix, at every , both , every and every , because the density is at least and, in polar coordinates around , the arm- part of the window contains a quarter-disk on which the uniform kernel equals one, namely a sector of opening with one sign choice per coordinate determined by where sits on . A quarter, not a half, is what is available at every interface point: at the four corners of the arm- part of a small window is exactly a quarter-disk, and keeps that sector inside . Its angular measure is the second factor in the display. Combining this with the coercivity constant, , and gives the class threshold
For the small-bandwidth arm mass, at every , both , and every , since every point of , corners included, retains at least an eighth of the disk area in each arm.
Finally, for every , both , and every , the one-dimensional Hausdorff integral of over is finite and strictly positive, because that slice is a nondegenerate circular arc and the density lies in .
Hence , with the common geometry displayed above.
Bit structure. Put . Point reflection through maps the disk onto itself, preserves and negates , so the angular correction integrates to zero over ; the remaining background density is the constant . With the score projection of Definition 18,
Each bump and each angular correction indexed by vanishes off , and the disks are disjoint; therefore the restriction of to depends on only through , and the restriction of to the complement of is the same for every .
Conditional signed-observation laws. Every has and , so exactly when . Since , the signed distance of Definition 21 therefore reduces on to the vertically signed radius,
Using the conditional-law notation of Definition 1, define which is unambiguous by the one-bit locality just proved and is a probability law. Equivalently, the image of restricted to under equals .
Signal separation. On the treatment-effect curve of Definition 23 is and at only the th bump is active, with , so
Hence two vertices agreeing in coordinate have the same target at , and flipping coordinate moves that target by exactly the amplitude,
Common signed-distance marginal. The same reflection argument applies slice by slice: for every Borel , the set is invariant under reflection through , so the angular correction contributes nothing to its mass and which is the asserted equality of signed-distance marginals.
Localized comparison of the two bit laws. Under the outcome is Bernoulli given the signed distance, with success parameter in by the middle-half bound above. The two success parameters differ only at short positive radii: for every Borel ,
Indeed, enabling the th bit adds to the conditional mean on the upper half of the radial slice at radius and simultaneously tilts the design density by . For the cutoff is fully active, and , so ; since the affine part of the treatment regression contributes the factor , the tilt contributes per unit slice mass, and the mean of over the upper half circle of radius is , which cancels the added bump exactly. For the two conditional means differ by at most , and on the control side, where the signed distance is negative, both laws carry the common mean .
Consequently the restrictions of the two laws to the complement of the short positive window agree: writing , which is the asserted agreement outside .
The two divergence bounds. Two Bernoulli outcome laws over a common statistic, with success parameters in and setwise mean-mass discrepancy at most localized to , satisfy and the same bound with the two laws interchanged, the marginals being common.
Because , the localized mass obeys
Taking and substituting , so both directed divergences are at most .
Collecting the displayed size and calibration identities, the common geometry, class membership, the bit structure, the conditional signed-observation laws, the exact signal separation, and the three local comparison facts gives every clause of the statement for the constants and the threshold .
∎The constants in Lemma 6 separate geometry from signal size. The radius gives room for a perturbation of height , and the packing size supplies enough boundary coordinates to generate the logarithmic uniform-risk scale. The common rectangular geometry places every inside the same fixed-known-geometry subexperiment, while the bit structure makes each disk a coordinate of the hypercube. The local comparison clause supplies the testing input: after observations, the effective information in a coordinate is calibrated by the cell probability and the Kullback–Leibler bounds.
To obtain Theorem 3, set the hypercube signal to the risk scale . Since , the calibration in Lemma 6 gives , , and . The comparison laws then satisfy with the same order for the reverse divergence. This calibration keeps the coordinatewise overlap controlled while the maximum over the packed boundary points contributes the logarithmic factor. Applying Lemma 1 to the signed-distance observations converts the overlap into a lower bound for recovering the hypercube coordinates, and the signal-separation clause converts that testing loss into the interface sup-loss in Definition 26. Because all vertices lie in the fixed rectangular subexperiment, the conclusion applies both to from Definition 27 and to the full risk .
The matched statement follows by combining the lower and upper components already stated in the main text. Theorem 3 gives the positive lower constant for the signed-distance risk after the envelope threshold is met. Proposition 1 gives a uniform outer-expected upper bound of order for the explicit winsorized, Gram-stabilized local-polynomial estimator with bandwidth and winsorization level . Therefore Theorem 4 characterizes the displayed known-geometry signed-distance minimax risk on the normalization , with the estimator in Definition 28 attaining the upper side of the rate over .
Reproducibility Note
The displayed definitions and theorem implications have been machine checked, including the two unsigned converses, the signed fixed-geometry lower bound, and the assembly of the conditional signed upper and frontier results. In Proposition 1 and Theorem 4, distance identification, uniform first-order bias, and expected maximal control enter as hypotheses.
Proofs of the main results
Throughout, denotes one potential-outcome observation and its observed outcome, as in Definition 22; is the sample and its law. For we write so that , the norm being the one appearing in Lemma 5 and in Definition 30. Fix , and , and abbreviate the bandwidth and clipping schedules of the statement by as in Definition 8 and Definition 28. We write , , and for the objects of Definition 28, for the population Gram of Definition 22, for the population coefficient of Definition 24, and for the target of Definition 23.
The first input of Definition 33 fixes what is being estimated: it is what makes the boundary contrast approached by the two one-sided signed-distance conditional-mean limits, and the bias input of the same collection is stated relative to that same target through Definition 29. Everything below is quantitative and uses the bias and expected-maximal inputs, Lemma 5, and the envelopes of Definition 22.
Step 0: the constants. Let be the constant supplied by the uniform first-order bias input of the statement and put Let be the constant supplied by the expected maximal bounds of the statement, and let be the constant of Lemma 5. Define All four summands are positive because . This is the constant of the conclusion; the threshold is fixed at the end of Step 2.
Step 1: the four bandwidth calibrations. Let and write , . We claim First, for every , so and hence and . Raising the definition of to the fourth power gives the identity that drives every calibration,
The first claim is the exponent computation .
For the remaining three, take logarithms in the definition of : and, since forces , both are nonnegative and satisfy Dividing by and using the displayed identity, Taking square roots gives the second claim, because , and the fourth claim directly. For the third, factor out and use the previous display together with which holds because ; thus .
Step 2: activating the two assumed inputs along the frontier bandwidth. The bias input of the statement is asserted along antitone bandwidth sequences, and is antitone only from onward. Set therefore Then for every , the sequence is antitone, , and so the bias input applies to with and yields Since for every , there is such that
The expected maximal bounds of the statement are asserted along positive deterministic sequences tending to zero whose heavy-tail regime diverges. At , using and , because for , so the numerator’s polynomial factor dominates the logarithm. Hence there is such that, for every and every , with as in Definition 30. Only this first half of the expected maximal bounds is used; the stochastic control of the score is supplied instead by Lemma 5, whose bounded envelope is created by the winsorization built into Definition 28.
Finally , so there is with for , which is the small-bandwidth regime in which the population Gram floor, the arm-mass threshold and Lemma 5 are all available. Put and fix and for the rest of the proof. Write and ; by Step 1, and .
Step 3: a small Gram deviation activates the stabilization guard. Put We claim that on the event the guard of Definition 28 holds simultaneously at every arm and every interface point ; that is, Indeed, means exactly that every entry satisfies , and for any symmetric perturbation with entries bounded by , the middle step by .
The population Gram lower bound of Definition 22, valid because and , gives . Subtracting the previous display leaves , which is the guard. This is where the factor in the guard threshold is spent: it is exactly the room left by the population floor after the empirical perturbation is absorbed.
Step 4: the guarded coefficient error is the inverted score. Suppose the guard holds at . Then is invertible: if then the guard forces , so .
Moreover, for every , Indeed, applying the guard to the vector and then the Cauchy–Schwarz inequality, and dividing by when it is nonzero gives the claim; when it is zero the claim is trivial.
On the guarded branch , so subtracting and using the definition of the winsorized score in Definition 32, Selecting the intercept, bounding a coordinate by the Euclidean norm, and converting norms through ,
Step 5: the mean of the winsorized score is of order . Write for the centered supremum appearing in Lemma 5, so that for every and every , We claim the deterministic bound Fix , , and a coordinate , and write the single-observation weight with the signed distance of Definition 21, the uniform kernel of Definition 19, and the basis of Definition 17. Since vanishes off and for , with the ball of Definition 3; the equivalence is immediate from the definition of the signed distance, whose magnitude is . By Definition 32 the th coordinate of the winsorized score is the sample average , so the observations being independent and identically distributed and the summand integrable,
Split the integrand by inserting and removing the unwinsorized outcome, The first term integrates to zero: by the normal equations of Definition 24,
For the second term, the moment envelope of Definition 22 is converted into a deterministic envelope pointwise in the outcome. For and , every real satisfies If the left side vanishes; otherwise , so , using and then with .
Integrating this against the selected conditional law of given and applying the moment envelope of Definition 22 gives, for each arm and each , Note that it is this absolute conditional remainder, and not merely the difference of the two conditional means, that the argument requires, because the remainder is multiplied by the random weight before it is integrated.
Combining the weight bound, the consistency relation of Definition 22, which makes at most the sum of the two armwise remainders, and the density envelope together with the planar area bound , which give
we obtain
Multiplying by and taking the maximum over proves the claim, since the bound is free of , and .
Step 6: the pointwise interface loss. We claim that for every sample, Two preliminary facts hold at every . First, , and the derivative envelope of Definition 22 bounds each arm regression by , so
Consequently, since projecting onto an interval containing cannot increase the distance to , the clipping in Definition 28 is harmless: Second, and the definition of the bias ratio in Definition 29 give the deterministic first-order bias bound Now split on the design.
Case 1: . By Step 3 the guard holds at both arms and every , so Step 4 applies at both arms. Inserting the two population coefficients and using the triangle inequality, Bounding each summand by Step 4 and then each score by Step 5, which is The right-hand side no longer depends on , so taking the supremum over and adding the nonnegative Gram term gives the claim in this case.
Case 2: . Here the clipping alone controls the loss: by construction and , so On the other hand the case hypothesis and the value of give so the Gram term alone already dominates the loss, and the claim holds because the two remaining terms are nonnegative. This is the calibration that fixes the coefficient : it is the smallest multiple of that pays for the trivial envelope on the ill-conditioned event.
Step 7: integration. Take outer expectations under in Step 6. The outer expectation of Definition 14 is monotone, is subadditive on sums, returns a constant on a constant because is a probability measure, and satisfies for every finite constant . Applying these four properties in turn to the three summands of Step 6, where . The measurability, integrability and extended-arithmetic bookkeeping needed to pass from the pointwise inequality to this display is routine and is suppressed.
Step 8: substituting the two maximal bounds. The three terms are now evaluated at and using Step 1.
The deterministic term. By the first calibration , so
The score term. Since , and , Lemma 5 applies to at this bandwidth and clipping level and supplies the last inequality by the second and third calibrations of Step 1. Hence This is the step at which the winsorization level is paid for twice over, and in opposite directions: raising shrinks the deterministic term but inflates the envelope term . The choice is exactly the one that places both on the scale .
The Gram term. By Step 2 and the fourth calibration of Step 1,
Adding the three displays and recalling the definition of in Step 0, Every quantity in was fixed in Step 0, before and , so the bound is uniform over the law class, and taking the supremum over gives which is the assertion, with and as constructed.
∎Step 1 (constants). Fix and put . Let be the envelope threshold supplied by Lemma 6, so . Fix and , and let be the positive constants that lemma attaches to . Set Then , and . Since we have , and therefore gives the budget inequality
Step 2 (the logarithmic packing budget). We first record, as an eventual statement in , the calibration that will put the hypercube of Lemma 6 inside the range of Lemma 1. For the rate of Definition 8 satisfies , so the budget inequality above gives Also , so every integer with obeys For all sufficiently large we have , whence ; and, since , also . For such the last display yields Combining the two estimates, for all sufficiently large and every integer ,
The de-Poissonization remainder decays exponentially in , while decays only polynomially, so .
Since , we may therefore fix such that, for every , and the displayed logarithmic packing budget holds.
Step 3 (the hard family at scale ). Fix and a known-geometry point-indexed rule , and choose a family of Borel sections representing it as in Definition 25. Put , so . Applying Lemma 6 at separation produces an integer , a radius , the disk mass , centers on the middle of the bottom edge with pairwise disjoint disks , laws indexed by all carrying the same fixed rectangular geometry and conditional signed-distance laws with . In particular , because .
Step 4 (per-coordinate divergence budget). The Kullback–Leibler clause of Lemma 6, the identity , and give, for every , where the last inequality uses the logarithmic packing budget above and .
Step 5 (target values at the centers). For and let be the vertex with th coordinate and all other coordinates , and put Because the treatment-effect target at depends on the vertex only through its th coordinate, and because flipping that coordinate moves the target at by exactly (Lemma 6),
Step 6 (the Poissonized experiment). For a probability law on a standard Borel space and , write for the law of the finite marked configuration obtained by drawing points independently from , attaching to each point an independent mark uniform on , and listing the points in increasing order of their marks; the marks are independent of the points and almost surely distinct, so they only fix a canonical ordering. For a configuration write for its number of points and, when , write for its first points. Let be the common known geometry of the hypercube, which is the same tuple for every vertex because all share , the Euclidean metric and . Define the Poissonized value of the rule at a point by and the Poissonized maximum loss at a vertex by Both are measurable, each section being Borel.
Step 7 (independent blocks and their compressions). Split a configuration drawn from according to the partition of into the pairwise disjoint disks and their complement. Because thinning a marked Poisson configuration by a measurable partition produces independent marked Poisson configurations, the blocks are independent, , and is the corresponding marked Poisson configuration built from the normalized law of off at the complementary intensity. By the bit-structure clause of Lemma 6 the law of depends on only through , each disk having the same mass , while the law of is common across all vertices.
Compress the th block by replacing each of its points by its signed observation at under the common geometry, keeping its mark; the compressed blocks are the compressed coordinates required by Lemma 1: By the signed-observation clause of Lemma 6, the point law of the compressed th block at bit is exactly , so
Poisson tensorization gives equal to for the two unordered marked configurations, and mark-ordering is a measurable map, which cannot increase divergence; hence, with the per-coordinate divergence bound above,
Step 8 (direct-product testing bound). Let be independent uniform bits, so that the vertex is uniform on . The block product experiment first draws the common block and the cell blocks conditionally independently according to the bit-conditional laws above, and then superposes those blocks; at vertex , the superposed configuration has the canonical marked-Poisson law . Attach to coordinate the midpoint decoder The configuration feeding is obtained by superposing the compressed blocks , , the compressed ancillary block , and the already compressed block ; the raw th block therefore enters only through . Hence is a measurable function of that is unchanged when the raw coordinate vector is altered only in coordinate while is held fixed.
All hypotheses of Lemma 1 are thus met, and its Kullback–Leibler consequence with , fed by the compressed-coordinate divergence bound, gives for the block experiment the middle inequality because forces .
Step 9 (from decoding error to loss, and averaging). If the midpoint decoder is wrong at coordinate under the block product experiment, then the estimate is at least as close to the wrong value as to the true one, so the separation forces
After superposition, this is a statement under the canonical marked-Poisson law . Consequently, for every vertex , Averaging this inequality over the vertices and bounding the average of the right-hand side by its maximum selects a vertex with The left factor is the average block-product decoding-error probability bounded below in the direct-product display, transported through the superposition map; and by the definition of . Therefore
Step 10 (de-Poissonization). Write . The hypercube construction gives and the fixed rectangular geometry, so and each center lies on . Define the finite packing loss of the rule on a sample of size by which is measurable in because each section is Borel and is measurable.
On the event the Poissonized value at is the value of the rule on the retained prefix, because makes the two signed observations at coincide, and by the target identity above; hence
On the complementary event the Poissonized value is , so ; and since the two potential-outcome regressions of Definition 22 satisfy , the target of Definition 23 obeys
Conditionally on the first points of the configuration are i.i.d. from , and , so splitting the Poissonized expectation over the two events gives
By the choice of the second term is at most . Combining this upper bound with the preceding Poissonized lower bound and weakening to , Cancelling the finite common term yields
Step 11 (passage to the two risks). Each lies on , so pointwise in the sample The finite packing loss is measurable, so its integral is at most the outer expectation of the boundary supremum. Hence
Since , taking the supremum over the law class and then the infimum over gives, by Definition 26, Since also satisfies the fixed hard-geometry condition, the same two steps taken inside the restricted supremum of Definition 27 give
Step 12 (normalization and liminf). For we have , hence The two sequences are eventually equal, so their lower limits agree, which is the first asserted identity.
Dividing the two eventual risk inequalities by gives, for every , and the order property of under an eventual lower bound yields This proves the theorem.
∎Write .
By Theorem 2, applied with the same and , choose such that Set .
The two normalizations of the common-map risk agree eventually. Indeed, for all sufficiently large , and because is the reciprocal of . Taking liminfs of these eventually equal sequences gives
For each , the common-map risk dominates the point-indexed risk: Two facts combine to give this.
First, a common map in Definition 11 is a point-indexed rule in Definition 13: take the same Borel section at every boundary point, as recorded by Lemma 3. Hence the infimum defining runs over a superset of the rules defining .
Second, for a common-map rule the two expectation conventions coincide term by term. Fix and a common map , and write the boundary loss at the sample point as The map is jointly Borel: is measurable, is continuous in and Borel in , and is continuous on the compact set because it lies in the Hölder ball of Definition 4.
Consequently, for every level , is the projection onto the sample coordinate of a Borel subset of the product of the Polish space with the sample space, hence an analytic set.
Analytic sets are universally measurable, so every superlevel set of is measurable for the completion of the product law generated by ; therefore is measurable on that completed space and its completed integral coincides with its outer expectation,
Taking the supremum over in the last display and then the infimum over the larger point-indexed class in Definition 14 bounds above by the common-map risk in Definition 12, which is the displayed inequality. Dividing by the same nonnegative normalizing factor preserves this order:
Applying monotonicity of liminf to the preceding pointwise normalized-risk inequality gives Combining this with the choice of in the first step yields and the equality of normalizations from the second step gives the displayed conclusion.
Fix and , and write .
A finite packing lower bound, uniform over point-indexed rules. Let with be the constants of Lemma 2, and put Take beyond the threshold of that lemma, and let , , , the boundary points , the laws , and the values be the objects it supplies, with the common covariate support. Two of its clauses are used at once: , so ; and , because contains a relatively open neighbourhood of the support point inside , which therefore has positive covariate mass, and that mass is .
Fix a point-indexed rule and law-independent Borel sections representing it as in Definition 13.
The Poissonized product experiment. Put independent uniform binary coordinates on . Conditionally on , let be Poisson with mean , independent of an i.i.d. stream of observations from , and attach to each observation an independent mark uniform on ; list the resulting marked observations in increasing mark order. Split this configuration along the partition of into the cells and their complement, and write Splitting a Poisson configuration along a measurable partition makes the blocks independent conditionally on . For , is the marked Poisson configuration with intensity times the restriction of to , equivalently with mean and the conditional observation law on . The complement block is the marked Poisson configuration with intensity times the restriction of to ; when this complement has positive mass this is the corresponding conditional law with that mean, and when the complement has zero mass it is the zero-intensity Poisson configuration. By the equal-cell-mass, one-bit-locality and common-outside clauses of Lemma 2, the conditional law of depends on only through , and the conditional law of is common for all values of , including the zero-intensity case.
The compressions and the decoders. The compressed blocks required by Lemma 1 are written here, the plain letter being reserved for the covariate support; let so that records, for each atom of the th cell block, its outcome, its distance from , and its mark.
For coordinate , superpose in mark order the compressed own block , the images under of the other cell blocks , and the image under of the complement block ; call the result , and set
Each is a measurable -valued function of , and it is unchanged when the raw block is altered while is held fixed, because enters only through .
The coordinatewise divergence budget. Because for , the conditional law in Definition 1 is defined; write Then Indeed, under , is exactly the block of radii at most of the marked Poisson experiment generated by the one-observation outcome-distance law at for the single-bit vertex with th bit . Restricting to that block is a measurable map, so it cannot increase divergence, and a marked Poisson experiment of mean has divergence at most twice that of the corresponding -fold product experiment, which is the compressed-distance divergence bounded by in Lemma 2.
The direct-product step. The uniform bits , the conditionally independent blocks , the compressions , and the decoders satisfy all the hypotheses of Lemma 1. Applying that lemma with and using the previous display gives the last inequality because and give .
From decoding errors to loss, and choice of a vertex. If , then the nearest-value decoder has picked the wrong endpoint, so the triangle inequality and the signal separation give
Multiplying the previous two displays and averaging over the uniform prior on selects a vertex such that, writing for expectation in the mean- marked Poisson experiment under ,
De-poissonization. Put , so , , and every . Define the finite packing loss of the rule at a sample point of size by a measurable function of , being a finite maximum of compositions of the Borel sections with the measurable distance compressions of Definition 10.
On the event the first atoms in mark order are an exact i.i.d. sample from , and by construction is then the value of the rule at on that retained sample and ; hence the Poissonized maximum equals of the retained sample.
On the complementary event every equals the fallback value , so the maximum is at most , the envelope coming from the zeroth-order derivative bound in of Definition 4.
Since is Poisson with mean , a Chernoff bound for its lower tail gives , and therefore
Because , we may enlarge the threshold on so that .
Combining the last three displays and cancelling one copy of the finite quantity yields: for every point-indexed rule there are , an integer , and boundary points with
For such and , the finite maximum is bounded pointwise by the full boundary loss because each lies on : The left-hand side is measurable, so its ordinary expectation is bounded above by the outer expectation of the right-hand side, that is, by the quantity appearing in Definition 14. Taking the supremum over , and then the infimum over all rules in , gives a sample-size threshold , distinct from the Poisson count introduced in the Poissonized experiment, such that
For every , and because is the reciprocal of . Thus the two normalized sequences agree eventually, so monotonicity of in both directions gives
For every , the positivity of allows division in the preceding eventual lower bound: Taking preserves this eventual lower bound. With , the asserted equality of normalizations and the lower bound follow.
1. Apply Theorem 3 for the chosen order , and let be the envelope threshold it provides. Set Then . Fix , , and the three displayed hypotheses of the theorem. The choice of gives both and .
2. The converse theorem at supplies a constant such that That theorem also records the parallel lower bound for the fixed-geometry risk of Definition 27; only the full-risk statement just displayed is used below.
3. By Proposition 1, applied with and the distance-identification, first-order-bias, and expected-maximal-bound hypotheses, there are and such that, for every ,
Since eventually, division by and passage to the limsup give
4. The explicit rule is admissible. For every the frontier bandwidth is positive, so the winsorization level of Definition 28 is well defined, and the estimator is generated by the law-independent family of sections where is the stabilized coefficient of Definition 28 computed from the signed-distance sample . This family does not depend on the unknown member law, and each fixed- section is Borel in , because the empirical Gram , the uniform conditioning event , the matrix inverse on that event, the winsorized empirical moment , and the clipping map are all Borel functions of . Moreover, for every and every , the signed-distance sample built from the known geometry of Definition 20 is exactly of Definition 21, since the treatment indicator is the deterministic function of the covariate recorded in . Hence
Therefore the minimax risk of Definition 26 is bounded above by the risk of this explicit rule:
After normalizing by , this inequality holds eventually in , and monotonicity of limsup yields
5. Finally set Then . The lower bound from Step 2, the general inequality , and the two upper bounds from Steps 3 and 4 give and These are exactly the asserted frontier bounds.
∎References
- Cattaneo, Matias D. and Titiunik, Rocio and Yu, Ruiqi Rae (2026). Estimation and Inference in Boundary Discontinuity Designs: Distance-Based Methods. Journal of Econometrics. doi
- Thistlethwaite, Donald L. and Campbell, Donald T. (1960). Regression-Discontinuity Analysis: An Alternative to the Ex Post Facto Experiment. Journal of Educational Psychology. doi
- Hahn, Jinyong and Todd, Petra and van der Klaauw, Wilbert (2001). Identification and Estimation of Treatment Effects with a Regression-Discontinuity Design. Econometrica. doi
- Imbens, Guido W. and Lemieux, Thomas (2008). Regression Discontinuity Designs: A Guide to Practice. Journal of Econometrics. doi
- Lee, David S. and Lemieux, Thomas (2010). Regression Discontinuity Designs in Economics. Journal of Economic Literature. doi
- Cattaneo, Matias D. and Titiunik, Rocio (2022). Regression Discontinuity Designs. Annual Review of Economics. doi
- Cattaneo, Matias D. and Idrobo, Nicolás and Titiunik, Rocío (2019). A Practical Introduction to Regression Discontinuity Designs. Cambridge University Press. doi
- Cattaneo, Matias D. and Idrobo, Nicolas and Titiunik, Rocio (2023). A Practical Introduction to Regression Discontinuity Designs: Extensions. Cambridge University Press. arXiv
- Calonico, Sebastian and Cattaneo, Matias D. and Titiunik, Rocio (2014). Robust Nonparametric Confidence Intervals for Regression-Discontinuity Designs. Econometrica. doi
- Imbens, Guido W. and Kalyanaraman, Karthik (2012). Optimal Bandwidth Choice for the Regression Discontinuity Estimator. The Review of Economic Studies. doi
- Calonico, Sebastian and Cattaneo, Matias D. and Farrell, Max H. (2020). Optimal Bandwidth Choice for Robust Bias-Corrected Inference in Regression Discontinuity Designs. The Econometrics Journal. doi
- Imbens, Guido W. and Wager, Stefan (2019). Optimized Regression Discontinuity Designs. The Review of Economics and Statistics. doi
- Armstrong, Timothy B. and Kolesar, Michal (2020). Simple and Honest Confidence Intervals in Nonparametric Regression. Quantitative Economics. doi
- Kolesár, Michal and Rothe, Christoph (2018). Inference in Regression Discontinuity Designs with a Discrete Running Variable. American Economic Review. doi
- McCrary, Justin (2008). Manipulation of the Running Variable in the Regression Discontinuity Design: A Density Test. Journal of Econometrics. doi
- Keele, Luke J. and Titiunik, Rocio (2015). Geographic Boundaries as Regression Discontinuities. Political Analysis. doi
- Keele, Luke J. and Titiunik, Rocio and Zubizarreta, Jose R. (2015). Enhancing a Geographic Regression Discontinuity Design Through Matching to Estimate the Effect of Ballot Initiatives on Voter Turnout. Journal of the Royal Statistical Society: Series A (Statistics in Society). doi
- Keele, Luke J. and Titiunik, Rocio (2016). Natural Experiments Based on Geography. Political Science Research and Methods. doi
- Kendall, Emmett B. and Beck, Brenden and Antonelli, Joseph (2026). Robust Inference for Geographic Regression Discontinuity Designs: Assessing the Impact of Police Precincts. Journal of the Royal Statistical Society Series A: Statistics in Society. doi
- Papay, John P. and Willett, John B. and Murnane, Richard J. (2011). Extending the Regression-Discontinuity Approach to Multiple Assignment Variables. Journal of Econometrics. doi
- Reardon, Sean F. and Robinson, Joseph P. (2012). Regression Discontinuity Designs With Multiple Rating-Score Variables. Journal of Research on Educational Effectiveness. doi
- Wong, Vivian C. and Steiner, Peter M. and Cook, Thomas D. (2013). Analyzing Regression-Discontinuity Designs With Multiple Assignment Variables. Journal of Educational and Behavioral Statistics. doi
- Choi, Jin-young and Lee, Myoung-jae (2018). Minimum Distance Estimator for Sharp Regression Discontinuity with Multiple Running Variables. Economics Letters. doi
- Díaz, Juan D. and Zubizarreta, José R. (2023). Complex discontinuity designs using covariates: Impact of school grade retention on later life outcomes in Chile. The Annals of Applied Statistics. doi
- Sawada, Masayuki and Ishihara, Takuya and Kurisu, Daisuke and Matsuda, Yasumasa (2024). Local-Polynomial Estimation for Multivariate Regression Discontinuity Designs. . arXiv
- Merlano, Eugenio Felipe (2025). Boundary Estimation in the Regression-Discontinuity Design: Evidence for a Merit- and Need-Based Financial Aid Program. . arXiv
- Cattaneo, Matias D. and Titiunik, Rocio and Yu, Ruiqi Rae (2025). Boundary Discontinuity Designs: Theory and Practice. . arXiv
- Fan, Jianqing and Gijbels, Irene (1996). Local Polynomial Modelling and Its Applications. Chapman and Hall.
- Ruppert, David and Wand, Matthew P. (1994). Multivariate Locally Weighted Least Squares Regression. The Annals of Statistics. doi
- Stone, Charles J. (1982). Optimal Global Rates of Convergence for Nonparametric Regression. The Annals of Statistics. doi
- Tsybakov, Alexandre B. (2009). Introduction to Nonparametric Estimation. Springer. doi
- Mammen, Enno and Tsybakov, Alexandre B. (1995). Asymptotical Minimax Recovery of Sets with Smooth Boundaries. The Annals of Statistics. doi
- Cuevas, Antonio and Rodriguez-Casal, Alberto (2004). On Boundary Estimation. Advances in Applied Probability. doi
- Vapnik, Vladimir N. and Chervonenkis, Alexey Ya. (1971). On the Uniform Convergence of Relative Frequencies of Events to Their Probabilities. Theory of Probability and Its Applications. doi
- Pollard, David (1984). Convergence of Stochastic Processes. Springer. doi
- van der Vaart, Aad W. and Wellner, Jon A. (1996). Weak Convergence and Empirical Processes. Springer. doi
Comments on earlier versions
Anchored to: