← All papers
First page of EdgeBench: Unveiling Scaling Laws of Learning from Real-World Environments

EdgeBench: Unveiling Scaling Laws of Learning from Real-World Environments

Deyao Zhu, Xin Zhou, Shengling Qin, Xuekai Zhu, Hangliang Ding, Shu Zhong, Zixin Wen, Zhonglin Xie, Chenhui Gou, Linxuan Ren, Yueyang Wang, Junfeng Zhong, Rui Liu, Tian Gao, Yangguang Lin, Jingyuan Zhang, Maojia Song, Xuan Qi, Jinhong Wu, Chenyang Zhang, Yinzhu Piao, Ziru Niu, Hongbin Lin, Lingxiang Meng, Peng Tang, Chengyao Tang, Shanyu Wu, Huanyu Zheng, Yu Liu, Liya Zhu, He Wang, Ming Ding, Ziyu Wan, Hao Liu, Sibo Wang, Haotian Zhu, Xintian Zhang, Nan Chai, Yipeng Liu, Panhao Lai, Sihang Yuan, Zixin Su, Ge Zhang, Wangchunshu Zhou, Yantao Du, Wenhao Huang, Guang Shi

cs.CL Jul 6, 2026 · v1 cs.LG
Includes a formal-math task family where agents fill proof obligations in Lean 4 projects such as the Carleson and FLT-regular formalizations.
Pretraining scaling laws reveal that model capability improves predictably with data and compute. But learning from real world environments after deployment remains far less understood. Analyzing roughly 38,000 hours of agent interaction with the environment across 134 real world tasks, we find, to the best of our knowledge, the first evidence that overall performance during environment learning follows a log-sigmoid scaling law with remarkably high precision, reaching R^2 = 0.998. Across model generations, we also find that agent learning speed roughly doubles every three months. This discovery stems from EdgeBench, a suite of 134 real world tasks with ultra-long horizons, spanning scientific discovery, software engineering, combinatorial optimization, professional knowledge work, formal mathematics, and interactive games. Each task sustains at least 12 hours of continuous agent operation under rich, multilevel feedback, and is built through substantial expert effort. We publicly release 51 tasks and our full evaluation framework to accelerate the study of how agents learn from real world experience.

Pretraining scaling laws are well studied, but it is unclear whether agents learning from real-world environments after deployment improve in a predictable way. Existing benchmarks give little feedback or allow only short horizons, so they cannot measure within-run learning.

EdgeBench is a suite of 134 ultra-long-horizon tasks across six families: scientific discovery, software engineering, combinatorial optimization, professional work, formal mathematics, and games. Each task supports at least 12 hours of agent operation with multilevel feedback. The formal-math tasks require completing proof obligations in Lean 4 projects, including real/functional analysis, Carleson's theorem, and Fermat's Last Theorem for regular primes, with transitive dependency checking. Five frontier models were run in three 12-hour trials per task, and the resulting performance trajectories were fitted with several functional forms.

Figure 2 : EdgeBench task taxonomy. 134 real world tasks across six capability families, with feedback channels designed to support within-run improvement. Recorded human expert effort estimates: mean 57.2 h.

Over roughly 38,000 hours of interaction, aggregate performance follows a log-sigmoid curve in interaction time with R^2 = 0.998. Agent learning speed roughly doubles every three months across model generations. Claude Opus 4.8 leads at 12 hours with 51.3, followed by GPT-5.5 at 48.4; 51 tasks are publicly released.

Functional formRMSE
Log-Sigmoid0.390
Log-Probit0.398
Log-Gompertz0.402
Weibull CDF0.404
Log-Linear0.717
Fit quality of functional forms for environment-learning curves