Towards Certificate-Driven Software Porting: A Self-Improving Agentic Harness for Scientific Program Optimization
Piyush Jha, Aishik Ghosh, Vijay Ganesh
cs.AI
Sep 28, 2026 · v1
cs.PL cs.SE
TL;DR
Lean is used in the evaluation harness to verify GPU work-assignment contracts, drawing on Physlib/HepLean and Mathlib; checked rules are translated into executable code.
Abstract
The upgrade and rewriting of large scientific codebases has traditionally been a major challenge. While evolutionary search with large language models (LLMs) can port and accelerate legacy code, repair feedback in prompts alone does not prevent subsequent candidates from repeating the same errors. We introduce Certificate-Driven Evolutionary Search (CDES), which extends evolutionary search with enforceable restrictions derived from failed candidates, recorded as certificates of assumptions, checker evidence, and justified restrictions. Its control logic enforces these restrictions through rejection, backtracking, and targeted repair while preserving compatible edits. We apply CDES to CPU-to-GPU translation of two particle-simulation functions from the Geant4 toolkit, evaluated with a harness that goes beyond unit tests to combine formal checks, numerical comparisons, physics checks, and GPU safety tests. Generated implementations achieve 13.78x and 23.54x function-level speedups over CPU code, including data conversion and transfers; for one function, GPU throughput exceeds an expert implementation by 14.9%, reaching 16.1% when complementary components are combined. In an ablation over execution settings, certificate feedback increases the fraction of candidates passing required correctness checks from 55% to 90%.
Problem
Porting and accelerating large legacy scientific codebases such as Geant4 to GPUs is difficult. In LLM-driven evolutionary search, repair feedback given only in prompts does not stop later candidates from repeating the same errors.
Approach
Certificate-Driven Evolutionary Search (CDES) extends ShinkaEvolve-style evolutionary search. Failed candidates are recorded as certificates that hold assumptions, checker evidence, and justified restrictions. Control logic enforces these restrictions through rejection, backtracking, and targeted repair. The evaluation harness combines Lean proofs of GPU assignment contracts (supported by Physlib/HepLean and Mathlib), CBMC, differential tests, physics checks, and GPU safety tests.
Results
On the Geant4 Compton and Fermi functions, generated CUDA code is 13.78x and 23.54x faster than the CPU versions. For Compton, GPU throughput is 14.9% above the expert Celeritas implementation, and a hybrid of the two reaches 16.1%. Certificate feedback raises the share of candidates passing required correctness checks from 55% to 90%.
| Function | Geant4 CPU | CDES CUDA | Speedup |
|---|
| Compton | 30.162 | 2.189 | 13.78x |
| Fermi | 340.090 | 14.447 | 23.54x |
Function-level speedups (ms)
| Metric | Without certificates | With certificates |
|---|
| Passing correctness checks | 55% | 90% |
| Passing with >10% gain | 50% | 80% |
| Passing with >20% gain | 30% | 35% |
| Best gain on withheld inputs | 66.5% | 68.6% |
Ablation of certificate feedback