What an odour descriptor corpus can and cannot measure: valence, attenuation, and the ceiling of the public record
Stylianos Kampakis, Fabio Rovai
cs.LG
Sep 11, 2026 · v1
TL;DR
Twenty supporting theorems on conditional testing, crosswalk recoverability, capacity bounds and majority aggregation are machine-checked in Lean 4 with Mathlib.
Abstract
Machine olfaction trains on pooled public descriptor corpora, but whether a shared descriptor word measures the same thing across corpora has not been tested, nor has the ceiling of what any of them can measure. We audit four corpora from Pyrfume. Conditioning on the molecule makes McNemar's test the exact conditional test of the corpus effect. Corpora disagree heterogeneously across descriptors ($I^2 = 80\%$) and non-uniformly with labelling breadth ($z = 17.2$), so no single offset repairs pooling. Median tetrachoric agreement is 0.795 against median $κ$ of 0.413: sources largely concur on which molecules deserve a word and differ on how readily they apply it. Of 109 descriptors with an estimable effect, 36 show large differential functioning on the ETS scale. Against a human panel's reliability, Morgan fingerprints with the full RDKit descriptor block reach 32.9% of achievable; adding every label from two merged corpora reaches 33.9%. The gap does not close with model capacity, encoding choice, more molecules, or more words. The missing variance is valence. One pleasantness rating per molecule reaches 54.6% of achievable (57.1% on an independent older instrument). Valence recovered from descriptors ($ρ= 0.457$) yields only 16.6%, so it must be measured. Five raters exceed structure plus the full descriptor record; fifteen to twenty saturate. We release a descriptor crosswalk and twenty machine-checked theorems.
Problem
Machine olfaction pools public odour descriptor corpora. It has not been tested whether a shared descriptor word measures the same thing across corpora, or how much of human perceptual similarity those corpora can capture at all.
Approach
Four Pyrfume corpora (Dravnieks, Keller, Leffingwell, GoodScents) are audited. Molecule-conditioned McNemar tests, random-effects meta-analysis and differential item functioning analysis separate differences in criterion from differences in construct. Representations are scored as a share of the reliability ceiling of a human panel. Twenty supporting theorems are machine-checked in Lean 4 against Mathlib with only the standard axioms: exactness of the conditional test, crosswalk recoverability, a 2^k capacity bound, and three-rater majority aggregation.
Results
Corpora disagree heterogeneously across descriptors (I^2 = 80%), and 36 of 109 descriptors show large differential functioning. Structure plus every descriptor label reaches about 34% of achievable similarity, while a single pleasantness rating per molecule reaches 54.6%. Valence is identified as the missing variance, and five raters already exceed structure plus the full descriptor record.
| Predictors | ρ | Share of achievable |
|---|
| Structure and every descriptor label | 0.303 | 33.6% |
| Pleasantness alone | 0.492 | 54.6% |
| Intensity alone | 0.130 | 14.5% |
| Familiarity alone | 0.105 | 11.7% |
Share of achievable perceptual similarity by predictor