Mutual Evaluation and Supervision without Peers
Zachary Robertson
cs.GT
Sep 17, 2026 · v1
cs.IT
TL;DR
Mutual-evaluation game theory results (annotation representation, envelopes, regret, equilibrium transfer) are formalized in a Lean 4 repository.
Abstract
This article introduces mutual evaluation of a replicable task worker and a critic that incentivizes truthful reporting, both modeled as strategic agents. The critic chooses a finite-valued rule that induces an evaluation score on joint report laws. Their common payoff is analyzed through regret relative to the unrestricted critic envelope. The critic rule is distinct from the evaluation score. This class enables a peer-free information elicitation mechanism using conditionally independent replications of a worker on the same task. This replication-loop mechanism implements a type-agreement payoff using same-task replications and new-task samples. In contrast to the peer-prediction and scoring-rule literature, implementations are shown that produce unbiased Pearson and Shannon information scores without requiring peers, a ground-truth reference, or likelihood-ratio estimation. A valid binary critic also can be represented by shared finite type annotations of worker returns. One runtime restriction is that the number of required replicas is random and can depend on the critic rule. Other timing effects, such as commitment and reoptimization, yield distinct incentives, connecting the framework to variational peer prediction. This mechanism class illustrates why strategic considerations matter for both critic and worker agents.
Problem
Peer prediction and proper scoring rules incentivize truthful reporting but require peer workers or observed ground-truth outcomes. The question is whether an incentive mechanism can elicit truthful information without peers or ground truth.
Approach
A mutual-evaluation game is defined between a replicable task worker and a strategic critic choosing a finite-valued rule that induces an evaluation score on joint report laws. A replication-loop mechanism uses conditionally independent same-task and fresh-task replications, with first-match waiting times giving unbiased Pearson and Shannon information scores. Payoffs are analyzed via a critic envelope and regret, and incentive properties (weak robustness, equilibrium transfer) are proven. The theoretical results are formalized in a Lean 4 repository.
Results
Valid binary critics correspond to finite type annotations of returns (Theorem 1). Value envelopes equal the true task mutual information, attained by literal-agreement critics, and critic regret equals discarded information (Theorem 3). Under weak robustness, truthful reporting with an optimal critic forms a Nash equilibrium (Theorem 5); expected sample counts are finite on finite alphabets despite unbounded waiting times.