Back to AI Research

AI Research

DepGPO follows command dependencies to assign terminal-agent training credit

Key Takeaways

  • The method traces writes and supporting reads back from verifier-inspected resources instead of crediting every step equally.
  • A terminal agent might inspect a file, try a repair, overwrite it and then make a different change that passes the task.
  • Rewarding the entire attempt equally gives little information about which commands contributed to the checked result.
  • Yu Li and colleagues propose using execution dependencies to allocate that training signal.
  • Their [Dependency-Aware Group Policy Optimization paper](https://arxiv.org/abs/2610.03634) introduces DepGPO.

A terminal agent might inspect a file, try a repair, overwrite it and then make a different change that passes the task. Rewarding the entire attempt equally gives little information about which commands contributed to the checked result. Yu Li and colleagues propose using execution dependencies to allocate that training signal.
Their Dependency-Aware Group Policy Optimization paper introduces DepGPO. The method constructs a command graph from execution traces, starts at resources read by the task verifier and follows the relevant writes and supporting reads backward.

Trace the resources that verification inspects

The graph captures file reads and writes as well as reuse of values printed to standard output. File writes are tracked at line level where before-and-after contents are available; reads are recorded at file level, which does not reveal exactly which lines a command used.
For a write, credit reflects the fraction of written lines that reach verifier-inspected resources directly or through qualifying later operations. Overwritten material does not automatically remain relevant. Reads can receive credit when their printed information supports a later write connected to the checked outcome, with contributions decaying over dependency distance.
The paper also handles ports and processes inspected by the verifier. These cannot use the same line-level representation as a file, so the method uses separate resource rules rather than pretending all command effects have identical granularity.

Dependencies distribute reward; they do not determine success

DepGPO keeps the trajectory's group-relative advantage, which says whether an attempt performed better or worse than others sampled for the same task. Dependency-derived factors change how much of that signal each step receives.
A command connected to verification can therefore receive a stronger positive or negative update. Relevance does not itself establish that the action was correct. The authors normalize factors across generated tokens so that redistribution does not change the overall advantage scale.
If tracing provides no useful differentiation, the method falls back to uniform trajectory-level assignment. A zero group-relative advantage also remains zero. DepGPO does not manufacture success information when every sampled attempt has the same binary outcome.

The reported experiment covers selected terminal tasks

The study evaluates Qwen3.5-9B and Qwen3.6-27B with SETA and TMAX training data and tests the resulting policies on Terminal-Bench 2.0 and 2.1. Within each configuration, the controlled comparison starts from the same supervised checkpoint and selected task pool.
That pool excludes all-pass and all-fail groups under binary rewards. The selection is part of the experimental setup and should remain visible when interpreting training gains. Evaluation estimates pass@1 from five independent attempts per task and repeats the full evaluation three times.
The authors report improved task performance and training stability, but the captured source ends in the experimental setup before complete outcome tables. Franklin cannot supply a supported percentage-point leaderboard from that excerpt.
The method offers a concrete use for terminal traces beyond debugging: allocate outcome supervision according to recorded command dependencies. It remains an attribution mechanism based on available traces and verifier reads, not a proof that every connected command was necessary or that it improves all coding agents.

Comments