A machine-learning experiment can run successfully while producing an invalid result. A leaked test label, disconnected gradient, or unwired evaluation flag may survive shallow checks and waste subsequent research iterations. RankEvolve studies whether runtime controls and cross-review between coding-agent products can reduce those failures.
The framework targets generative ranking models. It combines a declared experimental procedure, an execution graph containing complete coding-agent products, and a persistent record of hypotheses, patches, metrics, and failed experiments.
The runtime controls the procedure
An Executable Operating Protocol specifies phases, dependencies, gates, branches, and loops. Supported structural markers compile into a state machine that governs which step runs next. Natural-language instructions within a phase are still interpreted by the agent; the system does not claim they acquire formal semantics.
Human control remains explicit. In the reported deployment, an operator authorizes compute, reviews proposals, and can veto an experiment. Runtime enforcement can preserve process constraints, but it does not guarantee that generated code is semantically correct.
The cross-review layer composes complete products such as Claude Code and Codex rather than treating every worker as a bare model call. Adapters provide tasks, repository snapshots, tool permissions, budgets, and prior artifacts. Bounded review-and-repair loops let one product inspect another's patch.
Correctness is evaluated separately from ranking gains
ExecML uses hidden executable oracles to check patches against regression tests, task behavior, scientific-safety requirements, and evaluator integrity. Each of the HSTU and LitGPT repositories contributes ninety-six private tasks.
The authors report that heterogeneous composition raises all-oracle execution accuracy from the best matched-budget single-product baseline of 45.8% to 62.5%. The paired improvement is 16.7 percentage points, with a reported 95% confidence interval from 6.6 to 26.7 points. A pre-specified LitGPT transfer split reports a 12.5-point improvement.
The remaining 10.4% silent critical-defect rate is equally important. Better execution accuracy is not equivalent to eliminating runnable patches that invalidate an experiment.
What the twelve-iteration deployment demonstrates
On MovieLens-20M, the reported HSTU deployment reaches NDCG@10 of 0.2192 for LARGE and 0.1948 for BASE, respectively 4.48% and 2.80% above published anchors. These are historical best-checkpoint endpoints, not independent-seed estimates. The largest LARGE endpoint comes from an added genre side feature, rather than a newly invented modeling primitive.
The record includes rejected approaches and leakage incidents. Full-test metrics are distinguished from faster subset evaluations, which otherwise produced misleading apparent gains. Preserving those distinctions is part of the framework's evidence-management contribution.
Its knowledge layer records negative results and makes relevant lessons available to later steps. However, the study does not isolate that layer with a memory-on/off ablation. The paper supports a measured benefit from product-level cross-review in the evaluated settings, not a guarantee for arbitrary long-horizon research or a claim that every component's individual benefit has been established.
Comments (0)
to join the discussion
No comments yet
Be the first to share your thoughts!