Complex-Valued Phase-Coherent Transformer
Leona Hioki
cs.LG
May 11, 2026 · v3
TL;DR
Theorems on per-layer and all-layer phase coherence are machine-checked in Lean/Mathlib, the latter under an explicit premise bundle, with a public repository.
Abstract
Complex-valued Transformers have largely inherited softmax attention from real-valued architectures. However, row-normalised token competition is not necessarily aligned with phase-preserving computation. In this paper, we introduce the Phase-Coherent Transformer (PCT), which applies a real-valued, element-independent, smooth gate to L2-normalised complex query-key similarities. PCT replaces token competition with token-non-competing attention and is designed to preserve phase information across layers. Across mid-scale benchmarks spanning long-range memory, hierarchical long-range reasoning, positional retrieval, phase-based memory and superposition, and image classification, PCT shows strong generalisation across task categories. Under parameter-fair comparison, PCT consistently outperforms both the standard softmax Transformer and its direct complex-valued counterpart. Moreover, even on tasks traditionally considered difficult for complex-valued neural networks, such as NIAH and LRA-Text, PCT remains competitive with Multiscreen, the strongest real-valued NN baseline in our comparison. Experiments introducing gates that deliberately violate the PCT conditions show that the design is not incidental: smooth gates that preserve negatively aligned phase components remain strong, whereas gates that delete such components collapse on long-range retrieval, and gates whose outputs become excessively large suffer clear performance degradation. PCT also shows no depth-related accuracy collapse across the tested depth range. These results support introducing multi-layer phase-coherent structure into attention as a promising design principle for achieving generalisation in complex-valued Transformers.
Problem
Complex-valued Transformers usually inherit softmax attention. Its row normalisation forces tokens to compete for attention mass, which may conflict with preserving phase information across layers.
Approach
The Phase-Coherent Transformer (PCT) applies a real-valued, element-independent, smooth gate (a sigmoid) to L2-normalised complex query-key cosine scores, so tokens do not compete for attention mass. A four-condition framework (C1–C4) on the gate is linked to per-layer and all-layer phase coherence. Theorem 1 is formalized in Lean without sorry. Theorem 2 is machine-checked in Lean under an explicit premise bundle that packages lemmas not yet formalized. PCT is compared under parameter-fair conditions against real and complex softmax, sigmoid and screening baselines.
Results
PCT matches or exceeds the baselines across copy memory, NIAH, ListOps, FFT-MNIST, phase-memory and multi-pitch tasks, and shows no depth-related collapse. Gates that delete anti-phase components collapse on long-range retrieval. A complex screening model with a phase-coherent recurrence solves Path-X. #print axioms on the Lean theorem reports only the standard axioms.
| Task | r_softmax | r_screen | c_softmax | PCT |
|---|
| Copy d=2000 | 0.10 | 1.00 | 0.08 | 1.00 |
| NIAH L=2048 | 0.00 | 1.00 | 0.00 | 1.00 |
| ListOps mid L1024 | 0.146 | 0.698 | 0.104 | 0.854 |
| FFT-MNIST | 0.32 | 0.43 | 0.39 | 0.45 |
Headline accuracy (subset)