← All papers
First page of Co-Linguistics: AI-augmented Theory Construction in Linguistics

Co-Linguistics: AI-augmented Theory Construction in Linguistics

Emmanuel Chemla, Benjamin Spector, Alexandros Kalomoiros, Philippe Schlenker

cs.CL Sep 29, 2026 · v1
As an illustration, Claude re-formalizes geometric sign-language semantics and checks the equivalence of simplified restatements in Lean 4.
LLMs have been studied in recent linguistics as potential models of humans' linguistic abilities. Here we discuss an entirely different use of AI, namely as a co-scientist, to help construct and assess linguistic theories (we refer to the result as "Co-Linguistics"). Since the 1960s, linguistics has developed theories that are in principle mathematically formalizable, often in the language of formal language theory or model theory. The AI revolution in mathematics will thus have consequences in linguistics-but with an essential twist: proving new theorems is rarely the linguist's goal. Rather, one seeks to find the best set of axioms to derive empirical statements. AI could accelerate research by making existing theories fully explicit, by comparing competing theories, and more ambitiously, by proposing new theories (in machine learning, this relates to "program induction"). It will also help assess theories by accelerating the identification and test of crucial predictions, thanks to unparalleled access to data (in machine learning, this relates to "active learning"). While the cycle from theory evaluation to theory construction may give rise to recursive and possibly autonomous improvement of linguistic theories, humans remain central: linguists provide scientific directions and evaluate theories conceptually, and experimental participants are needed to assess empirical predictions that are outside the reach of LLMs.

Linguistic theories are in principle mathematically formalizable, but in linguistics the goal is rarely proving new theorems. Instead, linguists look for the best axioms from which empirical statements can be derived. The paper asks how AI can serve as a co-scientist in building and assessing such theories.

The authors set out a program they call 'Co-Linguistics'. In it, AI makes existing theories explicit, compares competing theories, proposes new ones (related to program induction), and identifies crucial predictions to test (related to active learning). Untrusted LLM theory generators are coupled with trusted proof assistants. As an illustration, Claude re-formalized the geometric part of a semantics for sign-language classifiers and checked the equivalence of simplified restatements in Lean 4.

Claude produced two equivalent, simpler geometric restatements of the classifier semantics and verified the equivalences in Lean. The authors argue that humans remain central, both for setting scientific direction and conceptual evaluation and as experimental participants.