← All papers
First page of SURGE: On the Potential of Large Language Models as General-Purpose Surrogate Code Executors

SURGE: On the Potential of Large Language Models as General-Purpose Surrogate Code Executors

Bohan Lyu, Siqiao Huang, Zichen Liang

cs.LG Feb 16, 2025 · v5 cs.CL
One benchmark subset contains Lean 4 proofs generated by Goedel-Prover on Lean-Workbook and labeled by Lean compiler verification; LLMs predict verification outcomes.
Neural surrogate models are powerful and efficient tools in data mining. Meanwhile, large language models (LLMs) have demonstrated remarkable capabilities in code-related tasks, such as generation and understanding. However, an equally important yet underexplored question is whether LLMs can serve as surrogate models for code execution prediction. To systematically investigate it, we introduce SURGE, a comprehensive benchmark with $1160$ problems covering $8$ key aspects: multi-language programming tasks, competition-level programming problems, repository-level code analysis, high-cost scientific computing, time-complexity-intensive algorithms, buggy code analysis, programs dependent on specific compilers or execution environments, and formal mathematical proof verification. Through extensive analysis of $21$ open-source and proprietary LLMs, we examine scaling laws, data efficiency, and predictive accuracy. Our findings reveal important insights about the feasibility of LLMs as efficient surrogates for computational processes. The benchmark and evaluation framework are available at https://github.com/Imbernoulli/SURGE.

It is unclear whether large language models can act as general-purpose surrogate code executors, predicting program behavior without running the code.

SURGE is a benchmark of 1160 problems in eight subsets. The subsets cover multi-language code, competition problems, repository-level code, scientific computing, time-complexity-heavy algorithms, buggy code, environment-dependent programs, and formal proof verification. The formal-language subset uses Goedel-Prover to generate Lean 4 proofs on Lean-Workbook, with correctness checked by the Lean compiler; models must predict the verification result. Twenty-one open-source and proprietary LLMs are evaluated under zero-shot, chain-of-thought, and few-shot prompting.

Figure 3: The Construction of SURGE employs 4 methodologies: 1. Iterative Refactor, 2. Repository Sampling, 3. Manual Implementation, and 4. Inference & Verification.

Claude-3.5-Sonnet achieves the highest zero-shot average (51.59), followed by DeepSeek-V3 (42.53) and GPT-4o (40.03). Performance varies widely across subsets, which shows substantial room for improvement in LLM-based execution prediction.

ModelFLAvg.
Claude-3.5-Sonnet17.9251.59
DeepSeek-V332.4642.53
GPT-4o21.9940.03
Qwen-Max29.8932.84
Zero-shot average scores on SURGE (selected models)