Summarizing an agent's context can save space while removing information needed for its next actions. A new analysis of TRACE's paired-replay corpus asks whether the agent's recent behavior can predict which compression boundaries will hurt.
The study examines 590 harness-triggered compaction boundaries in AppWorld. Its main finding is modest: the available pre-boundary history only weakly predicts additional errors or repeated calls after compaction. The authors cannot determine whether their best trigger beats a token-budget rule at matched retention from the released data.
Compare the same boundary with two context representations
TRACE re-executes a recorded prefix in a fresh environment, then runs the agent with either its pre-compaction context or the recorded summary plus retained last turn. It measures the burden of the next actions, including calls that error and calls repeated after their results were already available.
The primary outcome counts the difference over the next three actions. Of 590 boundaries, 344 have a positive difference in that union of errors and refetches.
These are local execution measurements, not final task-success results. A repeated call can contribute burden even if the agent eventually completes its assignment. Conversely, a blocked compaction could postpone a problem to a later boundary that the local measure does not price.
Avoid reading a phase proxy as a causal explanation
One initial test compares boundaries preceded by a write with read-only prefixes. Its uncertainty interval is too wide to provide the prespecified informative answer.
A later diagnostic shows that 368 of the 494 write-prefixed boundaries contain only login or session writes. The label often measures how far a trajectory has progressed, rather than a change to task-relevant application data. After adjusting for measured phase variables, an apparent substantive-write difference approaches zero.
The study's distinction between stored information and appropriate later use also connects to FAME's counterfactual examination of memory-induced beliefs. TRACE intervenes on context representation and observes actions; FAME uses model representations to examine whether remembered concepts adapt to changed scenarios.
Read the trigger results with their limits
The best extension-protocol history trigger reaches held-out AUROC 0.66, compared with 0.72 for a same-boundary replicate yardstick. AUROC measures ranking discrimination across thresholds, not a percentage of compactions repaired.
A frozen, interpretable logistic trigger avoids about 21% of positive-burden boundaries while retaining about 84% of compaction opportunities. It beats random selection in avoided boundary count at that operating point, but not in total positive burden mass. The comparison is post hoc.
The corpus uses one model as agent and compressor with a 4,096-token window, and all boundaries are harness-triggered. Released observables omit ordered actions, token counts and summary text. Without token counts, preserving a share of opportunities cannot establish equal savings against a token-budget baseline.
The analysis supports better measurement and richer released corpora before adopting a history-based scheduling policy. It does not supply a verified rule that makes compaction safe across coding agents or long-running workflows.
Comments