Babak Barazandeh and colleagues propose OnTrack, a monitor for an agent's execution while the run is still unfolding. It compares steps and their dependencies with successful reference trajectories when those references are available.
The paper distinguishes monitoring from judging a completed log. Its decisions can warn a user or halt an action before more execution occurs, but the available checks depend on the evidence and tool schemas supplied to the monitor.
Build a graph while dependencies are still incomplete
OnTrack represents events as nodes and connects them when one step uses an artifact from another. A search result may only become relevant when a later edit consumes the discovered file.
That creates a streaming problem: a new step can appear disconnected before its dependency becomes visible. The method treats structure as provisional, phases in its influence and inserts dependencies discovered later.
The reference comparison also considers which reference steps the agent could have reached. Comparing an early prefix against every step of a completed solution would penalize unfinished work even when the agent is progressing correctly.
The authors warm-start the alignment solver instead of resolving it from scratch after every event. They report roughly a millisecond per step for the streaming update, with a converged solve before an intervention.
Fewer inputs mean narrower safeguards
With reference runs and tool schemas, OnTrack can assess deviation from a plan and check declared prerequisites for irreversible actions. With schemas alone, it retains those prerequisite checks plus diagnostics for loops, stalls and repeated calls.
With only an event stream, the monitor retains self-referential diagnostics. It no longer has a reference plan to assess or a schema-based gate to enforce.
Dependency extraction is another constraint. Deterministic matching can recover links when artifacts have traceable identifiers, but misses dependencies hidden inside shell execution or unrecorded context. The paper's batch-equivalence statement requires online dependencies to match those the batch extractor would find.
These conditions prevent interpreting OnTrack as a complete safety guarantee for any agent or tool environment.
Savings concern failing runs in the evaluated corpus
The authors evaluate SWE-bench trajectories. Based on the first eight steps, their method improves failure-versus-success ranking by 0.057 AUROC over content-similarity approaches.
An added abort policy saves about 18% of compute otherwise spent on runs heading toward failure. This is not a claim of an 18% reduction in total production costs across successful and failing work.
The paper reports that 83% of interrupted runs were heading toward failure, described as five out of six aborts being correct. Interruptions therefore also carry a risk of stopping work that would succeed.
OnTrack offers a concrete way to use execution structure during a run. Its practical evaluation depends on reference quality, dependency visibility and the intervention policy, alongside the cost of a missed failure or an unnecessary halt.
Comments