← All papers
First page of Benchmarking Recursive-Collapse Warning Claims Under Matched False-Positive Control

Benchmarking Recursive-Collapse Warning Claims Under Matched False-Positive Control

David Mullett

eess.SY May 29, 2026 · v2 cs.LG stat.ML
Uses a Lean artifact to specify the claim boundary of the benchmark framework; it does not verify telemetry, benchmark validity, or detector performance.
Recursive systems can enter collapse-like regimes – self-reinforcing amplification, persistent recursion, and narrowing diversity that mask accelerating internal degradation – before overt failure becomes visible. We introduce Loopzero, a claim-bounded benchmark framework for testing whether recursive failures follow a directional telemetry pattern: rising gain (G), recursive persistence (p), and declining diversity ($δ$). The claim boundary is specified in Lean; the Lean artifact does not verify real telemetry, benchmark validity, or detector performance. We evaluate the bridge on two frozen public-artifact benchmarks: a segmented public-markets benchmark (Volmageddon 2018, COVID MWCB 2020) and a MovieLens-25M offline deterministic recommender replay. Detectors are evaluated under a locked equal-false-positive contract (FP $\in$ [0.03, 0.07], pre-registered) so all configurations face the same alert budget. Neither tested standard comparators nor Loopzero's pre-registered quantile detector achieved an accepted operating point. Directional witness alignment held on both canonical benchmarks, with adjacent-horizon and row-level limitations disclosed. Digitized Shumailov et al. (2024) LLM training-loop trajectories are directionally consistent with the pattern; matched-FP evaluation in that domain is deferred. The contribution is a reproducible, falsifiable benchmark framework for evaluating recursive-collapse warning claims under an explicit alert-budget contract – non-acceptance reported as a first-class scientific outcome.

Recursive systems can enter collapse-like regimes, with amplification, persistent recursion, and narrowing diversity, before overt failure is visible. Warning claims about such regimes need falsifiable evaluation under a controlled alert budget.

Loopzero is a claim-bounded benchmark framework that tests whether recursive failures follow a directional telemetry pattern: rising gain, rising recursive persistence, and declining diversity. The claim boundary is specified in Lean. Detectors are evaluated under a pre-registered equal-false-positive contract (FP in [0.03, 0.07]). The evaluation uses two frozen benchmarks: a segmented public-markets benchmark (Volmageddon 2018, COVID MWCB 2020) and a MovieLens-25M offline recommender replay.

Neither the standard comparators nor Loopzero's pre-registered quantile detector reached an accepted operating point. Directional witness alignment held on both canonical benchmarks, with adjacent-horizon and row-level limitations disclosed. Digitized LLM training-loop trajectories from Shumailov et al. (2024) are directionally consistent with the pattern; matched-FP evaluation in that domain is deferred.