← All papers
First page of Counterfactual Probing for Parallel Unmasking with Hidden Forest Structure

Counterfactual Probing for Parallel Unmasking with Hidden Forest Structure

Ryotaro Kawata, Satoshi Hayakawa, Taiji Suzuki

cs.LG Sep 29, 2026 · v2 stat.ML
The manuscript's named mathematical claims were formalized in Lean 4 with AI assistance and kernel-checked; the formalization is released on GitHub.
Masked generative models offer parallel token prediction, but accurate parallel sampling must account for dependencies among tokens. When dependencies are unknown, finding safe batches also costs model evaluations. We study whether total evaluations, including discovery, can be sublinear in sequence length $N$; sublinear sequential depth then follows. We consider discrete distributions with hidden forest structure, accessed through a fixed approximate conditional oracle. Under explicit regularity conditions and uniform Hellinger error bounds, for any fixed target accuracy $\varepsilon\in(0,1/8]$ and sufficiently large $N$, our sampler achieves seed-averaged total-variation error at most $\varepsilon$, with total masked-state submissions and sequential depth both bounded by $\widetilde{O}(N^C \varepsilon^{-a})$ for constants $0<C<1$ and $a>0$. These guarantees use polynomial vocabulary size and an edge-response lower bound set by $N$ and $\varepsilon$. The sampler shares evaluations of hypothetical reveals across dependence tests to identify safe parallel batches without requiring full recovery of the hidden forest. A tunable parameter trades probing cost against irreversible commit rounds. In the same class, any admissible irreversible product-commit sampler attaining the same seed-averaged accuracy requires $Ω(N^c \varepsilon^b)$ counterfactual submissions or commit rounds in the worst case, for constants $c,b>0$.

Masked generative models can predict many tokens in parallel, but accurate parallel sampling must respect unknown dependencies among tokens, and discovering safe batches also costs model evaluations. The question is whether total evaluations, including discovery, can be sublinear in sequence length N at fixed accuracy.

Targets are discrete distributions factorizing over a hidden forest, accessed through a fixed approximate conditional oracle with uniform Hellinger error bounds. The sampler uses shared counterfactual probing: it compares singleton predictions under hypothetical token fillings, uses random colorings and majority votes to certify low-degree neighborhoods, peels high-degree vertices by singleton commits, and then commits forest centroids in parallel. A parameter d trades probing cost against irreversible commit rounds. The named mathematical claims were formalized and kernel-checked in Lean 4 with AI assistance.

For fixed ε ≤ 1/8, the sampler achieves seed-averaged TV error at most ε with total submissions and depth Õ(N^{2/3+1/(9s)}), which is o(N). A matching-style lower bound shows that any admissible irreversible product-commit sampler needs Ω(N^c ε^b) counterfactual submissions or commit rounds. Exact-oracle experiments with N from 8192 to 16384 are consistent with sublinear scaling.