Choosing a base model for agentic post-training can be expensive when the selection test requires a working agent. Before They Can Solve proposes measuring the checkpoint at a verified decision point inside a successful coding trajectory instead.
The paper starts with post-trained agents' successful runs, replays every code-changing step and reruns the repository tests. The decisive step is the first action whose cumulative patch turns failure into success. That action supplies a target grounded in the task's executable verifier.
The checkpoint starts with a real trajectory prefix
Untuned base checkpoints often struggle with tool-call formatting and multi-step harness interaction. End-to-end failure can therefore hide whether a model supports a useful repair action. The proposed probes give it the preceding interaction and repository state, removing the need to operate the whole agent from a cold start.
The authors call the decisive action golden given its prefix. It is different from the human reference patch: a real agent produced it, and the same tests used to define task success certify it. Actions before that point may contain exploration or incidental edits, while later actions may change code after the issue is already resolved.
Three probes examine different kinds of support
Decisive-Action BPB measures the model's likelihood for the certified action, normalized by bytes so different tokenizers can be compared. Lower bits per byte means the checkpoint assigns more probability to that action. The scoring masks formatting inserted by the chat template.
Patch MCQ places the golden action among alternatives rejected by the same verifier. The model chooses an option using answer-letter likelihood, with reordered choices to reduce positional bias. This tests discrimination between resolving and non-resolving actions in a shared context.
Prefix-conditioned pass@K samples continuations from the decisive point and accepts any continuation the verifier passes. It tests the checkpoint's ability to generate a functionally correct action, allowing repairs beyond the recorded golden text.
Rank agreement is evidence for selection, not a guaranteed training outcome
Across ten public base/post-trained pairs, the probes closely track downstream SWE-bench Verified performance. For example, the paper reports a Spearman rank correlation of 0.964 for decisive-action BPB built from held-out DeepSWE trajectories. Rank correlation describes checkpoint ordering; it is not a predicted solve rate for each future model.
The paper distinguishes cross-benchmark evidence from probes constructed using the same benchmark as the downstream target. It also compares public post-trained scores where available with evaluations the authors ran where necessary. Changes in post-training procedure or model–harness pairing remain relevant to interpreting the cohort.
The Turbo Harness study's fixed execution model and task-specific runtime edits illustrate why checkpoint potential and the eventual agent system need separate evaluation. A promising base checkpoint does not specify the software around it.
This approach depends on successful trajectories and a verifier that captures task success. Supplying a near-solution prefix measures local support for useful behavior, rather than showing the base model can discover the entire solution unaided. It offers a screening tool for expensive training decisions, with end-to-end validation still needed after training.
Comments