← All papers
First page of Verification with Transfer: Exact Information Frontiers and Their Price in Calls

Verification with Transfer: Exact Information Frontiers and Their Price in Calls

Hazar Yueksel

cs.LG Oct 8, 2026 · v1 cs.CR cs.IT stat.ML
Nearly all numbered results are formalized in Lean 4 with Mathlib and released as arXiv ancillary files, assuming two published results.
A verifier that accepts or rejects whole answers reveals little: under a flat prior over $k$-bit answers, zero error needs $2^k-1$ verifications. The usual remedy is to solve related source tasks, either all first, as a curriculum does, or interleaved with verification. We price this remedy in information and in calls. With an exact verifier, the least causal information that any interleaving of source calls and $n$ verifications needs to succeed with probability $s$ is a list rate-distortion function, attained by one observation before any verification. It lower-bounds the expected number of binary source calls, which designed sources meet within $1+\log_25$ calls for unique answers and within a logarithmic term in general, where no additive constant suffices. With an exact verifier and fixed sources, moving every call before the first verification preserves all hard caps on calls, although interleaving can save unboundedly many expected calls; under a noisy verifier, source-first protocols can lose unbounded factors in information and in error. For linear banks over $\mathbb{F}_2$, optimal accuracy has a closed form, and after a polynomial-time reduction the budget profile is computable in time $2^{O(h^2)}\operatorname{poly}(J,k+h)$ for $J$ sources and nuisance dimension $h$. In these banks, for zero error under a hard cap, the calls beyond the rounded-up information price are exactly those spent on nuisance. Every numbered result apart from two clauses about the planner is machine-checked in Lean 4, assuming two published results. Used as a ruler, the frontier shows a small transformer using all delivered bits at latent dimension $5$ and none at $11$ within fixed training budgets; in a test with predictions recorded before training, low XOR degree of the target bits did not suffice for their use.

A verifier that accepts or rejects whole answers gives little information per call. Under a flat prior over k-bit answers, zero error needs 2^k-1 verifications. The paper asks how much source-task information and how many source calls are needed to succeed with probability s, and when doing all source calls first, as a curriculum does, loses nothing.

The least causal auxiliary information over all interleaved protocols is characterized as a list rate-distortion function. The analysis then bounds expected and hard-capped call counts for designed and fixed sources, and compares exact with noisy verifiers. Linear task banks over F_2 get closed-form accuracy and a planner whose running time is exponential only in the nuisance dimension. Numbered results are machine-checked in Lean 4 with Mathlib, except two planner clauses.

Designed sources meet the information price within 1+log2 5 expected calls for unique answers. With an exact verifier and fixed sources, a source-first schedule preserves every hard cap on calls, but interleaving can save unboundedly many expected calls. Under noisy verification, source-first can lose unbounded factors in information and in error. Used as a ruler, the frontier shows a small transformer using all delivered bits at latent dimension 5 and none at 11 within fixed training budgets.