A coding agent can improve a model's main score while damaging another behavior that the application needs. Binqian Xu and colleagues address that problem in ConflictGuide, an automated research method that adds feedback about competing behaviors to the code-editing loop.
The authors report task-error reductions of up to 28% and conflict-related error reductions of up to 14% compared with scalar-only AutoResearch across five model families. Those are the largest reported reductions across the evaluated settings, rather than uniform gains for every task. The proposal depends on identifying a specific conflict and measuring both behaviors with qualified probes.
A single task score can hide a regression
The paper examines settings with different sources of tension. An uncertainty-aware classifier must fit familiar examples while retaining useful distance-aware feature geometry. An echo-state network balances memory retention with nonlinear processing. An image compressor can lower the representation rate while losing fine structure.
A scalar task metric tells an agent whether an edit improved the selected objective. It does not reveal which competing behavior changed or whether the edit resolved the underlying conflict. ConflictGuide adds read-only probe metrics that make those effects visible alongside the code changes and task gains.
The authors organize potential conflicts in a literature-grounded taxonomy and use evidence about the target model, training objective and evaluation setting to choose a supported mechanism. The skill abstains when that evidence does not establish two desirable behaviors coupled by a concrete mechanism. Before evolution, calibration and paired evaluation reject unsupported probes.
Broad exploration comes before conflict-aware refinement
Stage I follows scalar-guided AutoResearch, giving the coding agent room to explore task-oriented changes. Stage II introduces the qualified behavioral measurements. The stage boundary is fixed before evolution in the evaluated protocol; it is not an unspecified agent decision to switch whenever a plateau appears.
The feedback also affects which edits survive. Substantial task improvements follow the main retention path. An edit with a marginal positive gain or eligible tie needs enough probe-measured conflict alleviation to qualify through the auxiliary path. Negative task gains are never retained.
In matched-budget exploratory continuations from common branch points, the share of proposals improving both behaviors rose from 7.2% to 13.1%. Joint degradations fell from 29.0% to 20.7%. Those proposal-level proportions explain how the feedback changed search, but they are separate from the final task-error reductions.
Evaluate the edits again after the search
The five settings cover scientific operator learning, uncertainty-aware classification, node classification, reservoir computing and learned image compression. For formal evaluation, the authors transfer the evolved source modifications and retrain each selected winner from scratch under matched settings. Several tests change datasets or evaluation conditions rather than reusing only the search metric.
This protocol helps distinguish a promising search trajectory from an improvement that survives retraining. It also makes the limits of adoption clear: an engineering team needs probes that measure its own behavioral conflict, a defined retention threshold and held-out checks. Adding extra numbers to an agent's prompt without validating what they measure would not reproduce the method tested in the paper.
Comments (0)
to join the discussion
No comments yet
Be the first to share your thoughts!