What the paper is about
Tool observations dominate the context of software-engineering agents, making long interaction histories costly to maintain. Existing context compression methods can discard information needed by later actions, while adapting agents to soft-token representations can compromise their original behavior. To reduce context while preserving action-critical information and agent behavior, we combine Latent Observations, Hard Actions (LOHA), a context layout that separates compressed history from text needed for exact reference, with Anchored Context Distillation (ACD), a training method that enables latent reading while constraining behavioral drift. LOHA compresses older tool observations into soft tokens while retaining the agent's own turns and the last K observations in text, providing compact access to historical information and exact access to recent content. To enable the agent to use this representation, ACD distills the base model's full-text predictions into the latent view while anchoring its behavior on plain-text inputs to the same base model. On SWE-bench Verified, K=3 reduces context per call by 43% for Qwen3-4B and 57% for SWE-Master-4B-RL, with resolve rates of 12.1% and 21.8% versus 14.5% and 27.5% for their uncompressed bases. A single-run recency sweep reaches 14.4% and 23.0% at K=8, with larger windows generally favoring task performance over compression. Under a 32K-token limit, Qwen3 with K=3 resolves 21.1% of a 199-instance subset versus 11.1% for the same adapted agent using full text. In concurrent single-GPU serving, it achieves 1.9 times that full-text agent's instance throughput. The same large language models question is explored in What Should We Ask Next? Retrieval-Aware..., which adds a research perspective.
What it covers
Compress What You See, Not What You Say: Anchored Context Distillation for Latent-Observation Software Engineering Agents Zhensheng Zou Guoqing Wang Dan Hao Affiliation: Peking University Abstract Tool observations dominate the context of software-engineering agents, making long interaction histories costly to maintain. Existing context compression methods can discard information needed by later actions, while adapting agents to soft-token representations can compromise their original behavior. To reduce context while preserving action-critical information and agent behavior, we combine Latent Observations, Hard Actions (LOHA), a context layout that separates compressed history from text needed for exact reference, with Anchored Context Distillation (ACD), a training method that enables latent reading while constraining behavioral drift. LOHA compresses older tool observations into soft tokens while retaining the agent’s own turns and the last K K observations in text, providing compact access to historical information and exact access to recent content. To enable the agent to use this representation, ACD distills the base model’s full-text predictions into the latent view while anchoring its behavior on plain-text inputs to the same base model. On SWE-bench Verified, K = 3 K{=}3 reduces context per call by 43% for Qwen3-4B and 57% for SWE-Master-4B-RL, with resolve rates of 12.1% and 21.8% versus 14.5% and 27.5% for their uncompressed bases. A single-run recency sweep reaches 14.4% and 23.0% at K = 8 K{=}8 , with larger windows generally favoring task performance over compression. Under a 32K-token limit, Qwen3 with K = 3 K{=}3 resolves 21.1% of a 199-instance subset versus 11.1% for the same adapted agent using full text. In concurrent single-GPU serving, it achieves 1.9 × 1.9\times that full-text agent’s instance throughput. 1 Introduction Large language models are increasingly used as software-engineering agents. To address a GitHub issue, an agent uses tools to inspect the repository, modify source code, and test its changes, refining the patch based on the results ( Jimenez et al., 2024 ; Yang et al., 2024a ; Xia et al., 2025 ; Wang et al., 2025 ) . These interactions produce tool observations, such as file contents, search results, stack traces, and test logs, that accumulate in the agent’s context. In CoderForge trajectories ( Ariyak et al., 2026 ) , tool observations account for 69% of trajectory tokens on average. Existing approaches reduce context through observation masking ( Lindenbauer et al., 2025 ) , history summarization ( Packer et al., 2023 ; Kang et al., 2026 ) , sub-task folding ( Sun et al., 2025 ; Ye et al., 2025 ) , or selective pruning of tool outputs ( Wang et al., 2026a ; Wang et al., 2026b ; Ren et al., 2026 ; Chen et al., 2026 ) . These methods shorten interaction histories, but deciding what to discard is difficult when later actions may depend on details in earlier observations. Coding agents also require more than a semantic summary: a str_replace edit, for example, requires an exact match to the source text. Information omitted during compression may therefore have to be retrieved again before the agent can act. Source-code compressors ( Shi et al., 2025 ; Wang et al., 2024 ) exploit program structure to guide compression, but are less suited to the search results, stack traces, and test logs that also occupy a substantial portion of an agent’s context. Soft-token compression ( Mu et al., 2023 ; Ge et al., 2024 ; Li et al., 2025 ; Li et al., 2026 ) offers an alternative by encoding text into a shorter sequence of continuous embeddings that the decoder reads directly. For coding agents, however, a useful compressed representation must support both understanding past observations and reproducing specific details for tool calls. These requirements are distinct: retaining enough information to reason about a file does not necessarily preserve the exact strings needed to edit it. Adopting soft-token representations also requires the decoder to learn a new input representation while retaining its existing ability to reason, invoke tools, and modify code. This raises two closely related questions: how should an agent combine compressed history with access to exact text, and how can it learn to use that history without compromising its ability to act? We address these questions with a framework that combines context compression with behavior-preserving adaptation (Figure 1 ). Its two components, Latent Observations, Hard Actions (LOHA) and Anchored Context Distillation (ACD), determine how the agent’s history is represented and how the agent learns to use it. LOHA follows a simple principle: compress what the agent sees, not what it says . It represents older tool observations as soft tokens while keeping the agent’s own turns, the system prompt, and the task description in ordinary text. To support actions that require exact strings, it also retains the K K most recent observations in raw text. This layout provides a compact representation of historical observations while preserving direct access to recent content for precise tool calls. ACD then enables the agent to read the compressed observations while limiting changes to its existing behavior. It combines two complementary objectives: a distillation term aligns predictions from the latent view with those of the frozen base model given the full text, while an anchor term constrains the adapted model’s predictions on full-text inputs to remain close to the same base model. Together, LOHA and ACD address the representation and adaptation challenges of using compressed observations in coding agents. Figure 1: LOHA retains recent text and compresses older observations; ACD adds latent reading with a full-text behavioral anchor. Only the adapter, reading patch, and normalization gains are trained. We evaluate LOHA + ACD on SWE-bench Verified with Qwen3-4B-Instruct and its RL-fine-tuned derivative, SWE-Master-4B-RL. At their full context windows, K = 3 K{=}3 reduces context per call by 43% and 57% , with resolve rates of 12.1% and 21.8% versus 14.5% and 27.5% for their uncompressed bases. In the single-run recency sweep, K = 8 K{=}8 resolves 14.4% and 23.0%, versus 13.6% and 22.2% at K = 3 K{=}3 , with more text retained and higher compute cost (Table 3 a). On SWE-Master, K = 3 K{=}3 also outperforms masking (17.6%). Under a 32K-token limit, Qwen3 with K = 3 K{=}3 resolves 21.1% of a 199-instance subset versus 11.1% for the adapted full-text agent. Concurrent single-GPU serving achieves 1.9 × 1.9\times that agent’s instance throughput. Our contributions are:
• A context layout for latent-observation coding agents. LOHA combines soft-token representations of older tool observations with uncompressed agent turns and a recent-observation window, providing compact historical context alongside exact text for actions involving recent content.
• Anchored Context Distillation for behavior-preserving adaptation. ACD combines cross-view distillation with a self-anchoring objective to teach latent reading while limiting behavioral drift. Our experiments show that better read-back performance alone does not ensure better agent behavior and identify self-anchoring as an effective way to balance the two.
• An empirical evaluation of the performance–efficiency trade-off. We compare task success and context usage on both an instruction-tuned agent and an RL-fine-tuned agent, and examine context limits and serving throughput. The results characterize the performance cost of compression and the settings in which it provides practical benefits. 2 Approach Our framework combines two components: Latent Observations, Hard Actions (LOHA), which organizes the agent’s context into compressed observations and uncompressed text, and Anchored Context Distillation (ACD), which trains the agent to use this representation while limiting changes to its existing behavior. We first define the context layout, then describe the student and its objectives, and finally explain training, inference, and cost accounting. 2.1 Problem Setup An issue-resolution trajectory consists of a system prompt s s containing tool schemas, an issue description u u , assistant turns a t a_{t} , and tool observations o t o_{t} : τ = ( s , u , a 1 , o 1 , … , a T ) \tau=(s,u,a_{1},o_{1},\ldots,a_{T}) . Assistant turns contain reasoning and tool calls; observations contain information returned by the environment, such as file contents, search results, and test logs. We use an LCLM-based compression pipeline ( Li et al., 2026 ) , denoted by ℰ \mathcal{E} , comprising an encoder and an adapter that maps its output to the decoder’s input space. At the default 16 × 16\times ratio, n n text tokens become approximately n / 16 n/16 continuous embeddings ( soft tokens ). The decoder processes these alongside ordinary text-token embeddings ( hard tokens ). 2.2 Latent Observations, Hard Actions LOHA compresses older tool observations while retaining the system prompt, issue description, and all assistant turns in text. A recent-observation window also remains uncompressed, providing exact text for actions involving recently accessed content. Rendering. Each observation remains in its original position within the decoder’s native tool-response envelope. For an observation selected for compression, its body is wrapped in memory markers, encoded in 1,024-token windows, and replaced by soft tokens. The surrounding envelope remains in hard tokens, preserving its association with the tool call; no additional header is introduced. Unlike pruning, this represents the observation through an encoder rather than selecting text spans, though compression does not guarantee verbatim reconstruction. The recency policy (Hard-Last- K K ). An edit often follows a file view, and a test command may quote an identifier from a recent failure report. LOHA therefore uses recency to retain raw text, rather than predicting which observation the next action will quote. Given m m observations, the body of o i o_{i} is represented as r K ( o i ∣ m ) = { o i , i > m − K , ℰ ( o i ) , otherwise . r_{K}(o_{i}\mid m)=\begin{cases}o_{i},&i>m-K,\ \mathcal{E}(o_{i}),&\text{otherwise}.\end{cases} For K > 0 K>0 , each new observation enters as text and becomes eligible for compression when it leaves the window. Older observations remain available through their latent representations; exact content can be retrieved again using the existing tools. The window size is selected at inference without retraining: increasing K K retains more exact text but reduces compression for a fixed history. We use K = 3 K=3 for the main comparisons and evaluate K ∈ { 0 , 1 , 2 , 3 , 4 , 8 , ∞ } K\in{0,1,2,3,4,8,\infty} , spanning fully latent to full-text observations (Section 3.3 ). 2.3 Anchored Context Distillation ACD adapts a base agent to read latent observations while regularizing its predictions on ordinary text. We apply the same recipe to Qwen3-4B-Instruct-2507 ( Qwen Team, 2025 ) and SWE-Master-4B-RL, each serving as the base decoder for its own student and as its frozen teacher. Rather than adopting the continually pre-trained LCLM decoder directly, we use its weight difference from the base decoder to initialize a low-rank reading patch. The training objective covers two input views. On the latent view, the student learns to match the teacher’s predictions from the corresponding full-text context. On the full-text view, the student is regularized toward the teacher’s original predictions. These objectives address latent reading and behavioral preservation within the same adapted model. Figure 2: (A) Projection-weight and normalization-gain changes from Qwen3 to LCLM. (D) Unit-normalized latent/text probe accuracy and mean norm ratio. Full measurements: Appendix E . Design motivation. Soft tokens introduce a new input representation, so the decoder must learn to use them while retaining its existing tool-use policy. The LCLM decoder supplies a trained reading initialization, but adopting it directly would also replace the base agent’s decoder. A low-rank patch provides a limited set of parameters through which to transfer and adapt that capability. Both latent and full-text inputs use the adapted decoder, so a reading objective can also change its predictions on ordinary text. ACD therefore adds a full-text anchor that directly regularizes those predictions. Figure 2 summarizes diagnostics of the decoder weights and latent inputs; quantitative measurements and initialization probes are reported in Appendix E . 2.3.1 The Student The student contains a frozen LCLM encoder, a trainable MLP adapter, and the base decoder augmented with a low-rank reading patch. Decoder normalization gains are also trainable and initialized from the base model; its remaining parameters are frozen. For each projection W 0 W_{0} , we define Δ W = W LCLM − W 0 \Delta W=W_{\mathrm{LCLM}}-W_{0} and parameterize the adapted weight as W θ = W 0 + A B W_{\theta}=W_{0}+AB , with rank ( A B ) ≤ ρ \operatorname{rank}(AB)\leq\rho . The factors A A and B B are initialized so that their product equals the rank- ρ \rho truncated singular value decomposition of Δ W \Delta W . The patch covers all 252 decoder projections. At the default rank ρ = 64 \rho=64 , it has 132M parameters, and the adapter has 9.2M. We denote the trainable patch, adapter, and gains by θ \theta , the student by p θ p_{\theta} , and the unmodified teacher by p 0 p_{0} . Teacher predictions are computed before student training and cached. The student uses the same adapted decoder in both views; the compression encoder and adapter are used only when constructing latent inputs. 2.3.2 Two Views, Two Terms Each trajectory is rendered as a full-text view x H x^{\mathrm{H}} , with all observations in hard tokens, and a latent view x L x^{\mathrm{L}} , with all observation bodies compressed. Let 𝒥 \mathcal{J} index corresponding assistant tokens and c j H c_{j}^{\mathrm{H}} , c j L c_{j}^{\mathrm{L}} their preceding contexts. Predictions are aligned by assistant-token identity, since compression changes absolute positions. We also append read-back questions to latent examples. Let ℛ \mathcal{R} index their answer tokens, with target y r y_{r} and preceding context c r L c_{r}^{\mathrm{L}} . These positions directly supervise recovery of content from compressed observations. The objective is ℒ = ℒ distill + λ ℒ anchor \mathcal{L}=\mathcal{L}{\mathrm{distill}}+\lambda\mathcal{L}{\mathrm{anchor}} , where ℒ distill \displaystyle\mathcal{L}{\mathrm{distill}} = 1 N distill [ ∑ j ∈ 𝒥 D KL ( p 0 ( ⋅ ∣ c j H ) ∥ p θ ( ⋅ ∣ c j L ) ) − ∑ r ∈ ℛ log p θ ( y r ∣ c r L ) ] , \displaystyle=\frac{1}{N{\mathrm{distill}}}\Bigg[\sum_{j\in\mathcal{J}}D_{\mathrm{KL}}!\left(p_{0}(\cdot\mid c_{j}^{\mathrm{H}}),\middle|,p_{\theta}(\cdot\mid c_{j}^{\mathrm{L}})\right)-\sum_{r\in\mathcal{R}}\log p_{\theta}(y_{r}\mid c_{r}^{\mathrm{L}})\Bigg], (1) ℒ anchor \displaystyle\mathcal{L}{\mathrm{anchor}} = 1 N anchor ∑ j ∈ 𝒥 D KL ( p 0 ( ⋅ ∣ c j H ) ∥ p θ ( ⋅ ∣ c j H ) ) . \displaystyle=\frac{1}{N{\mathrm{anchor}}}\sum_{j\in\mathcal{J}}D_{\mathrm{KL}}!\left(p_{0}(\cdot\mid c_{j}^{\mathrm{H}}),\middle|,p_{\theta}(\cdot\mid c_{j}^{\mathrm{H}})\right). (2) Here N distill = | 𝒥 | + | ℛ | N_{\mathrm{distill}}=|\mathcal{J}|+|\mathcal{R}| and N anchor = | 𝒥 | N_{\mathrm{anchor}}=|\mathcal{J}| count supervised positions within a batch. We use forward KL by default; alternative directions are evaluated in Appendix B.4 . Learning from the latent view. The distillation term encourages the student to make predictions from compressed observations that are consistent with the base model’s predictions from the original text. Matching the teacher’s distribution provides supervision beyond the single action recorded in the trajectory. The read-back component adds explicit supervision for extracting literal content from compressed observations. Its answers are substrings of the source observations and account for approximately 3% of supervised positions. The construction and filtering of these questions are described in Appendix D . Anchoring on the full-text view. The anchor term regularizes the adapted model on inputs that contain no latent observations. Cross-view distillation alone does not directly constrain predictions on these full-text inputs, even though both views use the same adapted decoder. Because the full-text view bypasses the compression pipeline, the anchor term updates the reading patch and normalization gains but has no gradient with respect to the compression adapter. The latent-view loss updates all three components. Separately normalizing the two terms makes their relative weight explicit through λ \lambda , rather than letting it depend on the numbers of supervised tokens in the two views. We use λ = 1 \lambda=1 in all reported runs. Choice of teacher. Both terms use the original base model as their teacher, but for different purposes: it provides the full-text prediction target for latent reading and the reference distribution for behavioral preservation. Using an external model for the anchor would instead encourage the student to adopt that model’s full-text predictions. Self-anchoring directly expresses our objective of retaining the starting agent’s behavior while adding latent-reading capability. The distillation teacher need not be identical to the anchor; Appendix B.4 examines their roles separately. Estimating the divergences. Since the teacher is frozen and the training trajectories are fixed, its predictions can be computed once. At each supervised assistant position, we cache the teacher’s top- k logits k_{\mathrm{logits}} token log-probabilities, with k logits = 64 k_{\mathrm{logits}}=64 , together with the remaining probability mass. During training, we evaluate each divergence over the cached tokens and one additional bucket containing the rest of the vocabulary. This coarsened divergence matches the probability mass assigned to the tail but does not constrain its internal distribution. Before numerical safeguards, it is a lower bound on the full-vocabulary KL in either direction. Implementation details and approximation measurements are in Appendices D and E.5 . 2.4 Training and Inference Training. We select the shortest successful CoderForge trajectory per task ( Ariyak et al., 2026 ) , yielding 35,140 trajectories; ablations use the first 3,514. Trajectories are rendered in both views and packed into 204,800-token sequences. We train for one epoch with AdamW, using a learning rate of 10 − 4 10^{-4} for the reading patch and normalization gains and 5 × 10 − 5 5\times 10^{-5} for the adapter. Training is offline: predictions are evaluated on recorded trajectory prefixes rather than on interactions generated by the student. Cached targets avoid a teacher forward pass during optimization, while gradients propagate through the student decoder to its trainable components. A plain-text tool-calling probe checks editing behavior and tool-call validity beyond the training losses. All reported arms are evaluated regardless of probe outcomes. Further optimization and probe details are in Appendix D . Inference. The adapted agent uses the scaffold’s native function-calling interface and ordinary tools. Before each model call, LOHA renders the history with the selected K K ; observations shorter than 128 characters remain in hard tokens regardless of age. Training thus uses the fully latent and full-text endpoints, whereas mixed-window inference combines older latent observations with recent raw text. Cost accounting. We distinguish representation length from the cost of an executed trajectory. Encoded observation bodies use approximately one soft token per 16 source tokens. The context-level reduction is smaller because envelopes, short observations, prompts, assistant turns, and the recent window remain in text. For a fixed trajectory, we compare the compressed representation with the full-text rendering of the same history. Separately, we measure cumulative decoder input over actual rollouts, whose lengths and tool calls may differ across conditions. Advancing the recent window changes an existing prefix and invalidates its cache from the first modified observation. The current server also re-encodes observations on each call. We therefore report encoder work, prefix reuse, and serving throughput separately from context length (Section 3.4 ); metric definitions are in Appendix A . 3 Experiments We evaluate task performance and context cost on two base agents, Qwen3-4B-Instruct-2507 and SWE-Master-4B-RL, its RL-fine-tuned derivative. Both use the same LOHA + ACD recipe, with each agent anchored to its own unmodified model. We sweep the recency window for both families; further Qwen3 experiments examine context limits, anchor identity, and serving efficiency. Setup. We evaluate both model families on 499 instances of SWE-bench Verified ( Jimenez et al., 2024 ) using OpenHands ( Wang et al., 2025 ) 0.62.0. Qwen3 uses a 262K-token context window and a 200-iteration cap; SWE-Master uses 131K tokens and 100 iterations. Conditions within each family share the scaffold and decoding settings, with one attempt per instance per seed. The uncompressed base, observation masking, and the two ACD conditions use two or three seeds; pruning baselines use one run. Full protocols are in Appendix A . Comparisons. The main comparisons use LOHA + ACD at K = 3 K{=}3 , the uncompressed base agent, and the same adapted model using full text ( K = ∞ K{=}\infty ). The latter separates adaptation from the effect of context compression. Baselines are observation masking ( Lindenbauer et al., 2025 ) , SWE-Pruner ( Wang et al., 2026a ) , Self-Prune, and LongCodeZip ( Shi et al., 2025 ) , all using the same recent-observation window and compression threshold. Metrics. We report resolve rate, context tokens per call, decoder input tokens per trajectory, and serving throughput. The whole-trajectory compression ratio compares the plaintext equivalent of an agent’s own contexts with the tokens it actually reads. Encoder work, prefix-cache reuse, and estimated compute are reported separately; metric definitions and paired-test results are in Appendices A and B . 3.1 Full-Window Results At K = 3 K{=}3 , LOHA reduces context per call by 43% for Qwen3 and 57% for SWE-Master, with a performance cost in both families (Tables 1 and 2 ). The recency sweep includes K = 8 K{=}8 as a larger-window alternative (Table 3 a). Table 1: Qwen3-4B-Instruct-2507 (262K context, 200 iterations). Resolved: out of 499, mean ± \pm standard deviation; pruners use single runs. Cost definitions: Appendix A . Condition resolved calls in/call dec in cached enc/aux tok ratio dec PF enc/aux PF Uncompressed anchor 72.3 ± \pm 3.1 (14.5%) 37.9 43.6K 1.65M 96.9% — 1.00 × 1.00\times 1.82 — LOHA + ACD (ours) plain text ( K = ∞ K{=}\infty ) 66.0 ± \pm 2.8 (13.2%) 48.0 50.9K 2.45M 95.2% — 1.00 × 1.00\times 4.24 — Hard-Last-3 (default) 60.5 ± \pm 0.7 (12.1%) 60.9 24.8K 1.51M 92.5% 1.86M 2.24 × 2.24\times 2.10 1.64 Reduction baselines (same window and threshold) Masking ( K = 3 K{=}3 ) ( Lindenbauer et al., 2025 ) 61.0 ± \pm 5.3 (12.2%) 60.6 20.1K 1.22M 91.5% — 2.39 × 2.39\times 1.69 0 SWE-Pruner ( Wang et al., 2026a ) 67 (13.4%) 48.5 29.1K 1.41M 94.0% 11.8K 1.47 × 1.47\times 1.83 0.01 Self-Prune 67 (13.4%) 50.4 38.8K 1.96M 94.3% 17.6K 1.34 × 1.34\times 2.72 0.13 LongCodeZip ( Shi et al., 2025 ) 71 (14.2%) 43.5 38.3K 1.67M 96.1% 13.6K 1.01 × 1.01\times 1.62 0.04 Table 2: SWE-Master-4B-RL (131K context, 100 iterations). Reporting follows Table 1 ; each family uses its own uncompressed reference. Condition resolved calls in/call dec in cached enc/aux tok ratio dec PF enc/aux PF Uncompressed teacher 137.0 ± \pm 5.7 (27.5%) 85.0 51.8K 4.40M 97.2% — 1.00 × 1.00\times 3.89 — LOHA + ACD (ours) plain text ( K = ∞ K{=}\infty ) 121.5 ± \pm 6.4 (24.3%) 78.9 50.1K 3.95M 92.3% — 1.00 × 1.00\times 6.77 — Hard-Last-3 (default) 109.0 ± \pm 7.1 (21.8%) 81.5 22.2K 1.81M 85.0% 2.45M 2.35 × 2.35\times 4.34 2.16 Masking ( K = 3 K{=}3 ) ( Lindenbauer et al., 2025 ) 88.0 ± \pm 1.0 (17.6%) 86.6 18.9K 1.64M 82.3% — 2.62 × 2.62\times 4.22 0 SWE-Pruner ( Wang et al., 2026a ) 120 (24.0%) 82.8 36.3K 3.00M 86.9% 57.9K 1.36 × 1.36\times 7.75 0.07 Self-Prune 109 (21.8%) 76.4 33.8K 2.58M 92.0% 78.8K 1.47 × 1.47\times 4.13 0.57 LongCodeZip ( Shi et al., 2025 ) 120 (24.0%) 79.3 48.6K 3.85M 91.5% 66.6K 1.00 × 1.00\times 7.60 0.17 Qwen3-4B-Instruct. LOHA + ACD resolves 12.1% of instances against 14.5% for the uncompressed base and 13.2% for the adapted full-text agent. Mean context per call falls from the base agent’s 43.6K to 24.8K tokens. Masking reaches a similar resolve rate (12.2%) with 20.1K tokens per call. The pruning baselines resolve 13.4–14.2% while retaining 29–39K tokens per call. SWE-Master-4B-RL. LOHA + ACD resolves 21.8% against 27.5% for the uncompressed base and 24.3% for the adapted full-text agent, while reducing mean context per call from 51.8K to 22.2K tokens. Masking drops to 17.6%; retaining a latent history therefore preserves more task performance on this base. SWE-Pruner and LongCodeZip resolve 24.0%, and Self-Prune matches LOHA’s 21.8%, with all three retaining larger contexts (34–49K tokens per call). Trajectory cost. Per-call savings translate differently across the two agents. Qwen3 with LOHA makes more calls than its base (60.9 versus 37.9), so decoder input falls only from 1.65M to 1.51M tokens per trajectory. SWE-Master makes a similar number of calls (81.5 versus 85.0), and decoder input falls from 4.40M to 1.81M. The representation ratios, measured against each agent’s own contexts rendered in full text, are 2.24 × 2.24\times and 2.35 × 2.35\times , respectively. 3.2 Performance under Context Limits Figure 3: Qwen3: (a) resolve rates on 199 instances: single runs at 32K/64K, seed means at 262K. (b) Overflow and resolved counts at 32K. (c) Single-GPU throughput at concurrency 16. Protocols and paired tests: Appendix B . For Qwen3, we repeat the comparison with 64K- and 32K-token context limits on a 199-instance subset. A trajectory that exceeds the limit is stopped and its patch is graded. At 64K, the four conditions resolve 31–36 instances. At 32K, Hard-Last-3 resolves 42, compared with 22 for the adapted full-text agent, 26 for masking, and 31 for the base agent (Figure 3 a). The gains over the adapted full-text agent and masking are supported by paired McNemar tests ( p = 0.0002 p{=}0.0002 and p = 0.0025 p{=}0.0025 , respectively). Masking exceeds the 32K limit less often than LOHA (23 versus 36 trajectories) but resolves fewer tasks (Figure 3 b). Fitting within the context window is therefore only part of the benefit: the results also favor retaining a compressed history over deleting older observations. Full counts and paired comparisons are in Appendix B.2 . 3.3 Ablation Studies We vary the recency window for both model families and the distillation teacher for Qwen3. The teacher ablation uses rank-64 students trained on 3,514 trajectories. Table 3: Recency and teacher ablations. (a) 499 instances per model; dec PF: estimated decoder PFLOPs per trajectory. (b) 199 random instances with the teacher supplying both targets, followed by a separate 20-task diagnostic listing teacher/anchor pairs. Res.: resolved; patch: non-empty patches; stuck: loop terminations. Single runs except the 199-task base mean. Details: Appendices C and B.4 . (a) Recency-window sweep Qwen3 SWE-Master K K res. dec PF res. dec PF ∞ \infty 80 1.47 128 2.99 8 72 2.10 115 8.37 4 76 2.08 108 4.49 3 68 1.67 111 3.74 2 65 1.40 103 2.93 1 47 1.21 88 1.89 0 37 1.18 51 0.93 (b) Distillation teacher Full text Hard-Last-3 Teacher res. patch stuck res. patch stuck Uncompressed base 34.3 — — — — — Self (Qwen3-4B) 28 140 43 29 127 50 Qwen3-30B-A3B 15 68 136 14 59 141 SWE-Master-4B-RL 0 22 59 0 22 68 20 tasks: distillation teacher / anchor Uncompressed base 10/20 14 2 — — — Self / self 11/20 14 3 11/20 16 5 30B / 30B 7/20 8 10 6/20 8 9 30B / self 12/20 15 5 9/20 13 6 Recency-Window Sweep. We sweep K ∈ { 0 , 1 , 2 , 3 , 4 , 8 } K\in{0,1,2,3,4,8} and full text while holding each trained student fixed (Table 3 a). Larger windows generally improve resolve rate while retaining more observations in text. Both models lose performance at K = 1 K{=}1 and again at K = 0 K{=}0 ; a separate 20-task diagnostic finds that increasing patch rank does not repair the fully latent policy (Appendix B.3 ). In this single-run sweep, increasing K K from 3 to 8 raises resolved counts from 68 to 72 on Qwen3 (13.6% to 14.4%) and from 111 to 115 on SWE-Master (22.2% to 23.0%). Estimated decoder compute rises from 1.67 to 2.10 PF and from 3.74 to 8.37 PF per trajectory, respectively. The trend is not strictly monotonic: Qwen3 resolves 76 tasks at K = 4 K{=}4 . Thus, K = 3 K{=}3 is a compact default, while K = 8 K{=}8 offers higher resolve rates at greater text retention and compute cost. On 358 literals quoted by later actions, the released LCLM decoder achieves exact-recitation rates of 20.4%, 15.6%, and 15.1% at 4 × 4\times , 8 × 8\times , and 16 × 16\times compression, respectively (Appendix E.6 ). The low fidelity even at 4 × 4\times supports retaining recent source text for exact tool arguments. Distillation Teacher. The 199-task block of Table 3 b uses the distillation teacher for both targets. Under full text and Hard-Last-3, the self-teacher re The same large language models question is explored in The Delegation Blind Spot, which adds a research perspective. as detailed in the full paper on Arxiv The same large language models question is explored in AutoViewMem, which adds a research perspective.
Comments (0)
to join the discussion
No comments yet
Be the first to share your thoughts!