Approximation theory · Neural networks
Sharp closure thresholds for two-hidden-layer ReLU networks
Abstract
We determine the minimum first hidden width needed to approximate the coordinate maximum by a fixed two-hidden-layer ReLU architecture with unrestricted parameters in L2. In dimension d ≥ 14, the threshold is κd = − ⌊d2/4⌋. The same value holds for ordered L-statistics whose weights have a nonzero discrete second difference. A fan-localization theorem assembles cellwise ridge formulas across a finite hyperplane arrangement, using a fixed finite second layer and output map. A complementary pair-transversal bound turns stable triple junctions into first-layer lower bounds. Its quantitative form is uniform over all weights, biases, and coefficients. The threshold proof is analytic for d ≥ 40. Smaller dimensions use stored tournament masks for block sizes 7 ≤ n ≤ 13 and exact integer existence bounds for 14 ≤ n ≤ 19. For the maximum, the critical construction has second width exp(O(d log d)) and converges in every finite Lp, with an explicit parameter–error rate. The common closure width is asymptotic to d2/4, whereas exact realization requires first width at least (1/2 − o(1))d2. At the critical first width, the constructed architecture therefore has zero approximation-error infimum with no finite-parameter minimizer.
Citation
Shunqi Lu. (2026). Sharp closure thresholds for two-hidden-layer ReLU networks (v1). Zenodo. https://doi.org/10.5281/zenodo.22728502