What the paper is about
Long horizon language model agents continually accumulate reasoning history, increasing context length and inference cost even after earlier decisions have been executed and observed. Unlike static Chain of Thought compression, removing historical reasoning can change future actions and the resulting interaction trajectory. We study when such reasoning can be safely forgotten. We propose Interaction Aware Compression for Long Horizon Reasoning (ICLR), a training free online method that ranks reasoning blocks using frozen proxy entropy while preserving actions, tool calls, and observations. On 260 WorkBuddyBench tasks, ICLR improves average reward from 0.699 to 0.718, while reducing input, output, and cache read tokens by 25.5%, 14.4%, and 33.3%, respectively. Ablations reveal trajectory amplification, where local reasoning deletion produces nonlinear changes in total computation by altering subsequent interaction. Representation probing, activation patching, and controlled trajectory analyses further suggest that historical reasoning becomes more replaceable once task relevant derived state has been reliably externalized into code, files, tool outputs, or environmental feedback. These results characterize agent reasoning as dynamic working state rather than permanent interaction history. The same ai evaluation question is explored in What Should We Ask Next? Retrieval-Aware..., which adds a research perspective.
What it covers
When Can Agents Forget Their Reasoning? ICLR for Long-Horizon Agent Context Compression Mingxuan Wang Affiliation: TierFlow Team Email: mailto:[email protected]@ruc.edu.cn Fei Luo Affiliation: TierFlow Team Bo Wang Affiliation: TierFlow Team Guorun Yao Affiliation: TierFlow Team Yinglong Guo Affiliation: TierFlow Team Chao Ning Affiliation: TierFlow Team Hongyue Chen Affiliation: TierFlow Team Yanbiao Ma Affiliation: Gaoling School of Artificial Intelligence, Renmin University of China Jungong Han Affiliation: Tsinghua University
- Corresponding authors. Abstract Long horizon language model agents continually accumulate reasoning history, increasing context length and inference cost even after earlier decisions have been executed and observed. Unlike static Chain of Thought compression, removing historical reasoning can change future actions and the resulting interaction trajectory. We study when such reasoning can be safely forgotten. We propose Interaction Aware Compression for Long Horizon Reasoning (ICLR) , a training free online method that ranks reasoning blocks using frozen proxy entropy while preserving actions, tool calls, and observations. On 260 WorkBuddyBench tasks, ICLR improves average reward from 0.699 0.699 to 0.718 0.718 , while reducing input, output, and cache read tokens by 25.5 % 25.5% , 14.4 % 14.4% , and 33.3 % 33.3% , respectively. Ablations reveal trajectory amplification, where local reasoning deletion produces nonlinear changes in total computation by altering subsequent interaction. Representation probing, activation patching, and controlled trajectory analyses further suggest that historical reasoning becomes more replaceable once task relevant derived state has been reliably externalized into code, files, tool outputs, or environmental feedback. These results characterize agent reasoning as dynamic working state rather than permanent interaction history. Figure 1 : Reasoning utility and cost are state dependent. (a,b) Scores and token usage vary across reasoning settings and task difficulty. (c) Online interventions show a nonmonotonic reward–efficiency tradeoff. See Appendices J and F . 1 Introduction Large language models are increasingly used in complex agent tasks, where reasoning, action, and environmental feedback form a repeated interaction loop 1 , 2 , 3 . As a task progresses, earlier reasoning is repeatedly included in later requests, increasing context length and inference cost 4 , 5 , 6 . Meanwhile, new observations continuously change the task state. This raises a central question for agent context management: once reasoning has produced an action and the environment has returned feedback, does that reasoning still need to remain in the active context? Figure 1 motivates this question from two observations. The most effective reasoning setting varies with task difficulty, and more reasoning does not consistently yield higher scores. Our online interventions also show a nonmonotonic relation among reasoning removal, reward, and input reduction. Together, these observations suggest that the value of reasoning depends on the current state rather than on reasoning quantity alone. Prior work has shown that explicit Chain of Thought reasoning contains substantial redundancy 7 , 8 , 9 , and existing methods can shorten intermediate reasoning while preserving answer quality 9 , 10 , 11 , 12 . These settings are largely static because a completed reasoning trace can be modified and then evaluated against the same final answer. Agent reasoning is different. Changing reasoning at step t t may alter the next tool call, observation, and subsequent decisions. Reasoning compression for agents therefore intervenes in a closed loop decision process rather than simply modifying text 13 , 14 , 15 . Reasoning can move across different forms of state. A plan, constraint, or intermediate conclusion may exist only in the reasoning trace, but become available through tool calls, code, files, or environmental feedback. Context management and memory systems use persistent or external state to reduce dependence on the active context 16 , 17 , 18 , 19 . We therefore hypothesize that the future necessity of reasoning depends on the current agent state and on whether task relevant information has already been externalized. We propose Interaction Aware Compression for Long Horizon Reasoning (ICLR) , a training free compression method for long horizon agents. After each interaction step, a frozen proxy model estimates the predictive entropy of reasoning blocks and preferentially compresses low entropy blocks 9 . The compressed history is directly used by the next agent request. Across long horizon agent tasks from multiple domains, ICLR improves reward from 0.699 0.699 to 0.718 0.718 , while reducing input, output, and cache read tokens by 25.5 % 25.5% , 14.4 % 14.4% , and 33.3 % 33.3% . Our ablations reveal that local deletion and total system computation are not proportional. Small changes to reasoning history can produce much larger changes in token consumption, while removing all reasoning does not necessarily produce the shortest trajectory because local interventions alter later decisions. We refer to this effect as trajectory amplification . To understand which reasoning should remain available, we further study its future necessity . Using representation probing techniques 20 , 21 , 22 , we find that the hidden state at the current context boundary contains decodable information associated with later rederivation, additional reasoning, and replanning, reaching an AUROC of 0.844 0.844 under strict task level nested evaluation. Activation patching provides complementary evidence 23 : replacing deleted history representations with their full history counterparts progressively restores the perturbed next token distribution, with KL Repair increasing from 0.047 0.047 at layer 4 to 0.459 0.459 at layer 20, 0.724 0.724 at layer 28, and 0.978 0.978 at layer 31. Trajectory analysis further suggests an internalization and externalization gap . When task specific derived state remains available only in reasoning, future behavioral risk is 67.3 % 67.3% , compared with 36.2 % 36.2% after that state has been externalized. Controlled interventions support the same interpretation: keeping only the latest reasoning preserves the full history rule score, deleting it reduces the score, clearing reasoning after relevant code has been written leaves the final score unchanged, and hiding observed environmental feedback causes the agent to recover the same information through an additional interaction. Together, these findings suggest that historical reasoning is most valuable when it remains the only reliable carrier of task relevant derived state. As that state is externalized, the original reasoning becomes increasingly replaceable. Our contributions are threefold. (1) Training free online reasoning compression for agents. We introduce ICLR, which continuously compresses accumulated reasoning inside a real agent interaction loop while preserving actions, tool calls, and observations. (2) Trajectory level characterization of agent reasoning compression. Through full benchmark evaluation and systematic ablations, we show that reasoning compression can alter the execution trajectory, producing a nonlinear relation between local deletion and total system computation. (3) Analysis of future reasoning necessity. We study historical reasoning through hidden state probing, activation patching, trajectory analysis, and information externalization. The results support a view of agent reasoning as dynamic working state whose future value depends on where task relevant information is stored. 2 Related Work 2.1 Efficient Reasoning and Chain of Thought Compression Extended Chain of Thought reasoning has become a central mechanism for mathematical reasoning, code generation, and complex decision making, but longer reasoning traces also increase token usage, latency, and inference cost 7 , 8 , 24 , 25 . A broad line of work therefore studies how much explicit reasoning is actually necessary, including adaptive reasoning budgets, early stopping, concise reasoning, and direct pruning of intermediate reasoning steps 9 , 10 , 11 , 26 , 12 . These studies collectively suggest that reasoning length and task performance are not simply proportional. A complete reasoning trace often mixes critical deductions with repeated checks, confirmations, and procedural elaboration. Identifying which parts can be removed without harming downstream behavior has therefore become an important direction in efficient reasoning. Recent work further estimates redundancy at the level of individual reasoning steps. Step Entropy 9 ranks reasoning steps according to uncertainty in the predictive distribution and shows that many low entropy steps can be removed from mathematical reasoning traces while preserving final answer accuracy. CONCISE 11 , TokenSkip 12 , and related compression methods similarly reduce explicit reasoning by identifying steps or tokens that contribute less to the final solution. These methods mainly optimize reasoning within a single completed generation. The decision to remove a step is therefore evaluated against a fixed downstream answer rather than an evolving sequence of actions and observations. Our setting differs because historical reasoning remains part of the agent’s future decision state. Removing a reasoning block can change the next action, which changes the next observation and all later decisions. The relevant objective is therefore not only whether compressed reasoning still supports the same answer, but whether repeated compression remains effective throughout a closed loop interaction. 2.2 Context Management and Memory for Language Model Agents The growth of interaction history has motivated context management methods for long horizon agents 27 , 28 , 29 , 30 , 31 . Token level prompt compression methods such as LLMLingua 32 , LongLLMLingua 33 , LLMLingua-2 34 , context compression 35 , and gist tokens 36 reduce redundant context. Agent oriented approaches manage evolving histories through relevance estimation, summarization, adaptive pruning, explicit context operations, or external memory 37 , 38 , 39 , 40 . Representative examples include ACON 4 , PACE 5 , SWE Pruner 14 , Self Compact 13 , Self GC 41 , Sculptor 15 , ContextBudget 6 , Context as a Tool 42 , ARC 43 , and structured context eviction 44 . A complementary line of work uses persistent or external memory to reduce dependence on the active context. MemGPT 16 , SAM 17 , CoMem 45 , ACM 18 , Mem1 46 , proactive memory extraction 47 , and transferable agent memory 48 preserve or recover information outside the immediate interaction window. These systems ask what information should remain accessible. We instead ask whether reasoning that has already produced an action still needs to remain after its consequences have been observed. This distinction follows from the information flow of an agent. Reasoning combines available evidence into plans and intermediate conclusions, while actions and observations can move the same task relevant information into code, files, tool outputs, or environmental feedback 19 , 42 , 43 . Historical reasoning may therefore change from being the only carrier of useful derived state to being redundant with information already stored elsewhere. Prior work on hidden state probing 20 , 21 , 22 , 49 and activation interventions 23 provides tools for studying this transition. We use them to examine whether the current model state predicts future reasoning reuse and whether reasoning becomes more replaceable after task relevant state is externalized. 3 Methodology Figure 2 : Overview of ICLR. At each completed agent step, newly generated reasoning is partitioned into blocks while actions, tool calls, and observations are preserved. A frozen proxy scorer estimates token level predictive entropy for each block. Low entropy blocks are removed before the compressed history is written back for the next agent request. Figure 2 summarizes the online compression process. ICLR operates inside the interaction loop rather than on a completed trajectory. It reduces accumulated reasoning while preserving the external records that define the evolving task state. We first describe online compression and entropy based selection, then the analyses used to study when historical reasoning remains useful. 3.1 Online Reasoning Compression We consider a long horizon agent that repeatedly reasons, acts, and receives environmental feedback. Let the t t th interaction step be S t = ( R t , A t , O t + 1 ) , S_{t}=(R_{t},A_{t},O_{t+1}), (1) where R t R_{t} is the generated reasoning, A t A_{t} the resulting action or tool call, and O t + 1 O_{t+1} the returned observation. Before the k k th model request, the interaction history is ℋ k = ( P , S 1 , … , S k − 1 ) , \mathcal{H}{k}=(P,S{1},\ldots,S_{k-1}), (2) where P P contains the system prompt, tool definitions, and task instruction. Static Chain of Thought compression modifies a completed reasoning trace before evaluating its final answer. Online compression instead changes the state from which future actions are generated. For the k k th request, ICLR induces ℋ k → ℋ ~ k → A k → O k + 1 → ℋ k + 1 . \mathcal{H}{k}\rightarrow\widetilde{\mathcal{H}}{k}\rightarrow A_{k}\rightarrow O_{k+1}\rightarrow\mathcal{H}{k+1}. (3) A local compression decision can therefore affect later tool selection, observations, recovery, and termination. The first agent request is uncompressed. Before each subsequent request, ICLR compresses only the reasoning from the just completed interaction step while preserving its action, tool call, and observation. The compressed reasoning is written back into persistent history and is not restored, so all later model calls operate on the modified trajectory. Local text removal is not equivalent to total computation saved. A small change to reasoning history can alter later decisions and therefore change both the number and content of future requests. Local deletion should thus be distinguished from total computation over the resulting trajectory. 3.2 Entropy Guided and Context Conditioned Pruning For each completed reasoning trace, we partition the reasoning into blocks, R t = ( B 1 , B 2 , … , B N ) . R{t}=(B_{1},B_{2},\ldots,B_{N}). (4) The current implementation uses double newline boundaries as a deterministic segmentation rule and excludes empty blocks. Compression is applied only to reasoning. Actions, tool arguments, tool outputs, observations, and final answers remain unchanged. This design prevents the compression policy from directly deleting evidence that has already entered the external interaction record. Proxy entropy. We score each reasoning block using a frozen proxy language model 9 . Given scorer context c t c_{t} , let z ( c t ) z(c_{t}) denote the logits predicting the next token. The predictive distribution is p θ ( v ∣ c t ) = softmax ( z ( c t ) ) v . p_{\theta}(v\mid c_{t})=\operatorname{softmax}!\left(z(c_{t})\right){v}. (5) The token level predictive entropy is h t = − ∑ v ∈ 𝒱 p θ ( v ∣ c t ) log 2 p θ ( v ∣ c t ) , h{t}=-\sum_{v\in\mathcal{V}}p_{\theta}(v\mid c_{t})\log_{2}p_{\theta}(v\mid c_{t}), (6) and the score of block B i B_{i} is the mean entropy of its tokens, H ( B i ) = 1 | B i | ∑ t ∈ B i h t . H(B_{i})=\frac{1}{|B_{i}|}\sum_{t\in B_{i}}h_{t}. (7) Lower entropy indicates greater predictive certainty under the frozen proxy model. We use this proxy entropy to rank reasoning blocks for compression. The acting agent is DeepSeek V4 Flash, while the scorer is a frozen Qwen3.5 9B model. Context conditioned ranking. A reasoning block is not scored in isolation. The scorer receives the system prompt, tool definitions, accumulated interaction history, and the current completed step in their original causal order. The score can therefore be viewed as a context dependent quantity, H ( B i ∣ ℋ t ) , H(B_{i}\mid\mathcal{H}{t}), (8) so the same reasoning text may receive a different score under a different trajectory state. This is important in agent settings because the relevance of earlier reasoning can change after new actions and observations have modified the task state. The full context scorer has a 42 42 K token input limit. If this limit is exceeded, pruning is skipped for that step rather than truncating the history observed by the acting agent. Block selection. Given N N reasoning blocks, ICLR removes the lowest entropy fraction, 𝒟 t = BottomK ⌊ ρ N ⌋ ( { H ( B i ) } i = 1 N ) , ρ = 0.8 . \mathcal{D}{t}=\operatorname{BottomK}{\lfloor\rho N\rfloor}\left({H(B{i})}{i=1}^{N}\right),\qquad\rho=0.8. (9) Each block in 𝒟 t \mathcal{D}{t} is replaced with [SKIP] , while all retained blocks preserve their original order and text. When N = 1 N=1 , no block is removed. Because the rule operates on blocks rather than individual tokens, the realized token reduction varies across interaction steps. We use a fixed compression ratio rather than tuning a task specific threshold. This choice creates a consistent intervention across trajectories and allows us to study whether agents can repeatedly discard a large fraction of historical reasoning during real interaction. Intervention variants. We construct three additional variants for controlled comparison. Step Only Entropy scores only the reasoning from the current completed step and removes earlier trajectory context from the scorer. Delete All Thinking removes the complete targeted reasoning trace while preserving actions and observations. Random r r removes a deterministic random subset of reasoning tokens with r ∈ { 5 % , 10 % , 20 % , 40 % } . r\in{5%,10%,20%,40%}. (10) All variants follow the same online write back protocol. The modified history is therefore consumed by every subsequent real agent request. These interventions separate the effects of context conditioning, selection strategy, and deletion strength. 3.3 Future Necessity Probing The compression policy determines which reasoning is removed, but it does not explain which reasoning will matter again later. We therefore study whether the current model state contains information associated with future reasoning reuse before that later behavior occurs. Following standard representation probing methods 20 , 21 , 22 , we analyze the model state available at each compression boundary. We use a frozen Qwen3.5 9B model as a representation sensor. For sample i i , layer l l , and pooling rule p p , let 𝒵 i ( l ) \mathcal{Z}{i}^{(l)} denote the hidden states associated with the current context boundary. We define the pooled representation as z i ( l , p ) = Pool p ( 𝒵 i ( l ) ) . z{i}^{(l,p)}=\operatorname{Pool}{p}\left(\mathcal{Z}{i}^{(l)}\right). (11) A linear probe then predicts a behavioral target y i y_{i} , p ^ i = σ ( w ⊤ z i ( l , p ) + b ) . \widehat{p}{i}=\sigma\left(w^{\top}z{i}^{(l,p)}+b\right). (12) The target records whether the later trajectory contains rederivation, additional reasoning, or explicit replanning. We use this behavioral target as an operational measure of future reasoning reuse. The probe characterizes this signal from the current representation. To prevent task leakage, evaluation uses strict task level nested splits. Layer selection, pooling choice, feature standardization, and PCA are selected using training tasks only. The outer evaluation folds contain held out tasks. This protocol asks whether the current representation contains information that generalizes across tasks rather than information specific to trajectories already seen during probe selection. Full extraction and evaluation details are reported in Appendix G . 3.4 Mechanistic Interventions for Reasoning Persistence Representation probing tests whether future reasoning reuse can be predicted from the current state. We complement this analysis with interventions that examine how historical reasoning affects model predictions and how its role changes after task relevant information becomes available outside the reasoning trace. Activation patching. We first test whether representations conditioned on historical reasoning participate in the next token prediction of the frozen sensor model 23 . For sample i i , let P i full P_{i}^{\rm full} denote the next token distribution under the full reasoning history and P i del P_{i}^{\rm del} the corresponding distribution after historical reasoning is removed. At layer l l , we replace the prompt boundary residual under the deleted history condition with the residual from the full history condition. The resulting distribution is denoted by P i , l patch P_{i,l}^{\rm patch} . Let d i = D KL ( P i full ∥ P i del ) . d_{i}=D_{\rm KL}\left(P_{i}^{\rm full}|P_{i}^{\rm del}\right). (13) For samples with d i > 0 d_{i}>0 , we measure distributional recovery as Repair i ( l ) = 1 − D KL ( P i full ∥ P i , l patch ) d i . \operatorname{Repair}{i}(l)=1-\frac{D{\rm KL}\left(P_{i}^{\rm full}|P_{i,l}^{\rm patch}\right)}{d_{i}}. (14) Larger values indicate that the patched representation more closely recovers the full history next token distribution. The intervention is performed only on the frozen sensor model and does not modify the online ICLR policy. State externalization. We next examine whether reasoning becomes more replaceable after task specific derived state has been written into a persistent external carrier. Let 𝒰 i \mathcal{U}{i} denote the unexternalized task specific derived state at compression boundary i i . We define G i = 𝕀 ( 𝒰 i ≠ ∅ ) . G{i}=\mathbb{I}\left(\mathcal{U}{i}\neq\varnothing\right). (15) Here, derived state refers to an intermediate plan, relation, constraint, synthesis, or task specific conclusion rather than a direct copy of the task instruction or a raw observation. External carriers include code, files, tool outputs, and explicit environmental feedback. The indicator G i G{i} separates states in which reasoning remains the only available carrier of task relevant derived information from states in which that information has already been materialized elsewhere. This allows us to study whether future reasoning reuse changes as information moves from internal reasoning into persistent external state. Controlled state interventions. Finally, we construct fixed state continuations that modify the available reasoning or feedback while keeping the surrounding task state unchanged. The interventions compare retaining or deleting recent reasoning, clearing reasoning after relevant code has been written to disk, and retaining or hiding previously observed environmental feedback. These interventions test whether the next decision still depends on information available only in the reasoning trace. Together, probing, activation patching, state externalization analysis, and controlled interventions provide complementary views of reasoning persistence. The full protocols are reported in Appendices H and I . 4 Experiments We evaluate ICLR on WorkBuddyBench using DeepSeek V4 Flash as the task executing agent and a frozen Qwen3.5 9B model as the proxy entropy scorer. WorkBuddyBench contains long horizon tasks from Code, Office, Security, and Web domains. We report official task reward together with input tokens, output tokens, cache read tokens, and wins, ties, and losses. The full benchmark contains 260 tasks, consisting of 80 Code, 50 Office, 60 Security, and 70 Web tasks. A fixed 80 task set with 20 tasks from each domain is used for ablation analysis. 4.1 Overall Performance We first compare ICLR with representative context management baselines on the fixed Pilot40 task set used by prior WorkBuddyBench context management experiments. Table 1 includes simple sliding window baselines, periodic summarization, and representative context management methods including PACE, SelfCompact, ACON Core, SAM, and SWE Pruner. This comparison focuses on the tradeoff between task quality and token consumption. ICLR directly prunes low entropy reasoning from the online interaction history using a frozen scorer, providing a lightweight alternative to summary based or external memory based context management. Table 1 : Comparison with context management methods on WorkBuddyBench Pilot40. All methods use the same fixed 40 tasks. Scores are multiplied by 100 100 . Parenthesized values show changes from the DeepSeek-V4-Flash base agent. Method Code Office Sec. Web Avg. Tokens ↓ \downarrow Base and Ours DeepSeek-V4-Flash 72.9 81.6 44.6 71.0 67.5 1.89M ICLR (Ours) 74.1 (+1.2) 81.1 (-0.5) 64.1 (+19.5) 62.0 (-9.0) 70.3 (+2.8) 2.02M (+6.9%) Simple Context Baselines Sliding Window ( K = 5 K=5 ) 66.8 (-6.1) 67.4 (-14.2) 32.3 (-12.3) 59.0 (-12.0) 56.4 (-11.2) 2.95M (+56.3%) Sliding Window ( K = 10 K=10 ) 74.4 (+1.4) 73.1 (-8.6) 40.1 (-4.5) 61.0 (-10.0) 62.1 (-5.4) 1.98M (+5.1%) Sliding Window ( K = 20 K=20 ) 79.3 (+6.4) 85.1 (+3.5) 40.3 (-4.3) 70.0 (-1.0) 68.7 (+1.1) 2.02M (+7.0%) Periodic Summary ( n = 3 n=3 ) 65.4 (-7.5) 67.7 (-14.0) 26.3 (-18.3) 82.7 (+11.7) 60.5 (-7.0) 2.21M (+17.2%) Periodic Summary ( n = 5 n=5 ) 83.6 (+10.7) 75.7 (-5.9) 40.0 (-4.5) 81.3 (+10.3) 70.2 (+2.6) 2.37M (+25.5%) API-based Methods PACE 5 70.4 (-2.6) 55.8 (-25.9) 43.5 (-1.0) 64.0 (-7.0) 58.4 (-9.1) 1.10M (-41.6%) LLMLingua-2 34 64.6 (-8.3) 73.3 (-8.3) 40.1 (-4.5) 70.0 (-1.0) 62.0 (-5.5) 2.40M (+27.3%) SelfCompact 13 76.8 (+3.9) 79.6 (-2.0) 36.2 (-8.4) 71.0 (0.0) 65.9 (-1.6) 1.57M (-17.0%) ACON-Core 4 71.8 (-1.1) 68.6 (-13.1) 47.9 (+3.4) 65.0 (-6.0) 63.3 (-4.2) 1.40M (-25.7%) Self-GC 41 63.5 (-9.5) 71.7 (-10.0) 43.7 (-0.9) 71.0 (0.0) 62.5 (-5.1) 1.66M (-12.4%) LRE 50 62.6 (-10.3) 63.4 (-18.2) 27.1 (-17.4) 68.0 (-3.0) 55.3 (-12.2) 2.29M (+21.3%) CoMem 45 56.4 (-16.5) 59.9 (-21.8) 2.5 (-42.1) 58.0 (-13.0) 44.2 (-23.3) 0.77M (-59.5%) SAM 17 67.6 (-5.3) 81.6 (0.0) 47.2 (+2.7) 67.0 (-4.0) 65.9 (-1.7) 1.35M (-28.7%) SWE-Pruner 14 63.5 (-9.5) 72.6 (-9.0) 40.9 (-3.7) 65.0 (-6.0) 60.5 (-7.1) 1.50M (-20.5%) Sculptor 15 51.2 (-21.7) 81.6 (0.0) 35.2 (-9.3) 69.0 (-2.0) 59.3 (-8.3) 1.29M (-31.5%) Released-policy Method ACM 18 38.7 (-34.2) 55.9 (-25.7) 6.2 (-38.4) 34.0 (-37.0) 33.7 (-33.8) 2.36M We next evaluate ICLR on the complete 260 task benchmark. Table 2 follows the score and token presentation used by prior WorkBuddyBench context management evaluations. Average reward increases from 0.699 0.699 to 0.718 0.718 . Detailed token accounting is reported separately in the appendix. The method obtains 102 / 76 / 82 102/76/82 wins, ties, and losses. These results show that aggressive online pruning of reasoning can reduce repeated processing of historical context without sacrificing aggregate task performance across domains. Domain-level results and component-wise token changes are reported in Appendix E . Table 2 : Full benchmark system performance and token usage. Scores are multiplied by 100 100 . Green and red indicate favorable and unfavorable changes from the base system. Model Code Office Sec. Web Avg. Tokens ↓ \downarrow DeepSeek-V4-Flash 77.0 81.8 47.8 72.1 69.9 2.69M ICLR (Ours) 70.6 (-6.4) 77.6 (-4.2) 70.7 (+22.9) 69.9 (-2.2) 71.8 (+1.9) 2.19M (-18.5%) 4.2 Ablations and Trajectory Amplification Figure 3 : Trajectory amplification. Token savings do not track deletion. We next study whether the observed effect depends on entropy based selection, access to full history during scoring, or simply reducing the amount of reasoning. Table 3 reports results on the fixed 80 task ablation set. ICLR reaches a reward of 0.738 0.738 , compared with 0.627 0.627 for the baseline. Step only entropy reaches 0.711 0.711 , while deleting all thinking reaches 0.718 0.718 . Random deletion also changes task performance, but none of the reported random variants reaches the reward of full history entropy. These results suggest that reasoning history contains substantial redundancy while also indicating that context conditioned ranking provides useful information beyond deletion alone. Appendix F reports the full intervention and token summaries used for this analysis. Table 3 : Ablation results on the fixed 80 task WorkBuddyBench set. All variants use the same fixed 80 tasks, with 20 Code, 20 Office, 20 Security, and 20 Web tasks. Scores are multiplied by 100 100 . Parenthesized values show changes from the Base Agent. Method Code Office Sec. Web Avg. Tokens ↓ \downarrow Base and Ours Base Agent 69.09 78.94 44.17 58.50 62.67 3.00M ICLR (Ours) 63.85 (-5.24) 81.14 (+2.20) 72.03 (+27.86) 78.00 (+19.50) 73.76 (+11.09) 2.49M (-17.0%) Reasoning Removal Controls Delete All Thinking 69.20 (+0.11) 82.04 (+3.10) 67.41 (+23.24) 68.50 (+10.00) 71.78 (+9.11) 1.11M (-62.9%) Random 5% 59.81 (-9.28) 81.12 (+2.18) 64.17 (+20.00) 66.83 (+8.33) 67.98 (+5.31) 1.40M (-53.5%) Random 10% 55.08 (-14.01) 82.65 (+3.71) 71.91 (+27.74) 60.61 (+2.11) 67.56 (+4.89) 1.19M (-60.3%) Random 20% 67.23 (-1.86) 76.92 (-2.02) 64.82 (+20.65) 61.50 (+3.00) 67.62 (+4.95) 1.69M (-43.8%) Random 40% 61.35 (-7.74) 76.17 (-2.77) 61.37 (+17.20) 71.62 (+13.12) 67.63 (+4.96) 1.47M (-51.0%) Scoring Context Ablation Step Only Entropy 60.48 (-8.61) 80.58 (+1.64) 73.10 (+28.93) 70.38 (+11.88) 71.14 (+8.47) 0.98M (-6 The same ai evaluation question is explored in The Delegation Blind Spot, which adds a research perspective. as detailed in the full paper on Arxiv The same large language models question is explored in Emergent Collusion in Long-Horizon LLM Agent..., which adds a research perspective.
Comments (0)
to join the discussion
No comments yet
Be the first to share your thoughts!