Back to AI Research

AI Research

Retrieval-Augmented Skill Optimization via Cross-Ha... | AI Research

Key Takeaways

  • What the paper is about An agent skill is a reusable, actionable natural-language artifact that guides an agent to perform a task effectively under a given h...
  • An agent skill is a reusable, actionable natural-language artifact that guides an agent to perform a task effectively under a given harness.
  • Recent studies have explored the optimization of agent skills, contributing to a growing collection of publicly available skills spanning diverse tasks, domains, and harnesses.
  • To address this, we propose \textbf{Retrieval-Augmented Skill Optimization (RASO)}, a framework that leverages an external skill corpus as prior knowledge throughout skill optimization.
  • RASO retrieves relevant knowledge from existing skills and adapts it to the target task and harness via Cross-Harness Adaptation, accounting for mismatches in both domain and harness.
Paper AbstractExpand

An agent skill is a reusable, actionable natural-language artifact that guides an agent to perform a task effectively under a given harness. Recent studies have explored the optimization of agent skills, contributing to a growing collection of publicly available skills spanning diverse tasks, domains, and harnesses. Despite millions of publicly shared skills, existing skill optimization methods largely overlook this accumulated knowledge, instead relying solely on expensive agent rollouts to iteratively refine skills for a target task. To address this, we propose \textbf{Retrieval-Augmented Skill Optimization (RASO)}, a framework that leverages an external skill corpus as prior knowledge throughout skill optimization. RASO retrieves relevant knowledge from existing skills and adapts it to the target task and harness via Cross-Harness Adaptation, accounting for mismatches in both domain and harness. RASO comprises two complementary stages: \textbf{Retrieval-Augmented Skill Initialization (RASI)} constructs a knowledge-grounded initial skill without requiring agent rollouts, while \textbf{Retrieval-Augmented Skill Update (RASU)} iteratively refines the skill by retrieving external knowledge guided by execution feedback. Across four agent benchmarks and two models, extensive experiments show that RASO consistently outperforms baselines without retrieval-augmented skill initialization and updating.

What the paper is about

An agent skill is a reusable, actionable natural-language artifact that guides an agent to perform a task effectively under a given harness. Recent studies have explored the optimization of agent skills, contributing to a growing collection of publicly available skills spanning diverse tasks, domains, and harnesses. Despite millions of publicly shared skills, existing skill optimization methods largely overlook this accumulated knowledge, instead relying solely on expensive agent rollouts to iteratively refine skills for a target task. To address this, we propose \textbf{Retrieval-Augmented Skill Optimization (RASO)}, a framework that leverages an external skill corpus as prior knowledge throughout skill optimization. RASO retrieves relevant knowledge from existing skills and adapts it to the target task and harness via Cross-Harness Adaptation, accounting for mismatches in both domain and harness. RASO comprises two complementary stages: \textbf{Retrieval-Augmented Skill Initialization (RASI)} constructs a knowledge-grounded initial skill without requiring agent rollouts, while \textbf{Retrieval-Augmented Skill Update (RASU)} iteratively refines the skill by retrieving external knowledge guided by execution feedback. Across four agent benchmarks and two models, extensive experiments show that RASO consistently outperforms baselines without retrieval-augmented skill initialization and updating.

What it covers

Retrieval-Augmented Skill Optimization via Cross-Harness Adaptation Jaewon Chu Affiliation: Korea University Email: [email protected] Ji Soo Lee Affiliation: KAIST Email: [email protected] Jihwan Park Affiliation: KAIST Dohwan Ko Affiliation: Korea University Jeehye Na Affiliation: KAIST Seunghun Lee Affiliation: KAIST Taehoon Lee Affiliation: KAIST Minseo Yoon Affiliation: KAIST Minseok Joo Affiliation: Korea University Yunyang Xiong Affiliation: Meta AI Hyunwoo J. Kim † † thanks: Corresponding author Affiliation: KAIST Abstract An agent skill is a reusable, actionable natural-language artifact that guides an agent to perform a task effectively under a given harness. Recent studies have explored the optimization of agent skills, contributing to a growing collection of publicly available skills spanning diverse tasks, domains, and harnesses. Despite millions of publicly shared skills, existing skill optimization methods largely overlook this accumulated knowledge, instead relying solely on expensive agent rollouts to iteratively refine skills for a target task. To address this, we propose Retrieval-Augmented Skill Optimization (RASO) , a framework that leverages an external skill corpus as prior knowledge throughout skill optimization. RASO retrieves relevant knowledge from existing skills and adapts it to the target task and harness via Cross-Harness Adaptation, accounting for mismatches in both domain and harness. RASO comprises two complementary stages: Retrieval-Augmented Skill Initialization (RASI) constructs a knowledge-grounded initial skill without requiring agent rollouts, while Retrieval-Augmented Skill Update (RASU) iteratively refines the skill by retrieving external knowledge guided by execution feedback. Across four agent benchmarks and two models, extensive experiments show that RASO consistently outperforms baselines without retrieval-augmented skill initialization and updating. 1 Introduction Large language models (LLMs) are widely deployed as agents within execution harnesses that define available tools, file access, and scoring procedures ( Yao et al., 2023 ; Yang et al., 2024b ) . In these settings, performance depends not only on the parametric knowledge of the underlying model but also on the procedural policies governing task execution ( Wang et al., 2024 ; Wang et al., 2025 ) , commonly referred to as agent skills ( Anthropic, 2025 ; Li et al., 2026c ) . An agent skill is a reusable, actionable natural-language artifact that specifies how an agent should accomplish tasks under a given harness. Unlike policies encoded in model weights, agent skills are expressed as text, making them readily inspectable, auditable, and transferable across models without modification. To automatically construct effective agent skills, skill optimization has recently attracted growing attention ( Wang et al., 2026 ; Xia et al., 2026 ; Ding et al., 2026 ; Chen et al., 2026 ) . Existing methods iteratively refine skills based on agent experience collected through training-task rollouts under a given harness. Through such refinement, millions of skills have been publicly shared ( Destefanis et al., 2026 ) , collectively encoding procedural knowledge accumulated across diverse tasks, models, and harnesses. Despite this extensive knowledge, existing methods largely rely on the agent’s own experience to optimize each skill ( Yang et al., 2026 ; Alzubi et al., 2026 ; Ni et al., 2026 ; Tang et al., 2026 ) , requiring costly rollouts for iterative refinement. As a result, leveraging an external skill corpus as a prior for automatic skill optimization remains largely unexplored. To this end, we propose Retrieval-Augmented Skill Optimization (RASO) , a skill optimization framework that leverages an external skill corpus as prior knowledge. As in Fig. 1 , RASO comprises two complementary stages, Retrieval-Augmented Skill Initialization (RASI) and Retrieval-Augmented Skill Update (RASU) . RASI leverages prior knowledge from an external skill corpus to construct an effective initial skill directly from task and harness descriptions, without using any expensive agent rollouts. Upon observing a failure during task execution, RASU identifies the missing knowledge underlying the failure and retrieves relevant content from the corpus to address it. However, retrieved skills are originally written for specific tasks and harnesses, potentially encoding domain-specific assumptions or referencing tools unavailable in the target environment. To address this domain and harness mismatch, RASO introduces Cross-Harness Adaptation, a shared operation employed by both RASI and RASU. Rather than directly incorporating retrieved content, this operation abstracts away source-domain-specific instruction and re-expresses the underlying procedure in terms of the objects, commands, and units supported by the target harness. Figure 1: Comparison of RASO to prior skill optimization. Existing skill optimization methods (left) either generate an initial skill from the LLM alone or directly reuse retrieved skills, then refine them mainly from execution feedback. However, retrieved skills often come from mismatched harnesses; on SpreadsheetBench 92.9% originate from a different harness (right, top). RASO (middle) addresses this by retrieving relevant external skill knowledge and adapting it to the target domain and execution harness through Cross-Harness Adaptation , supporting both skill initialization (RASI) and skill update (RASU). Overall, our RASO results in substantially stronger performance than both direct retrieval and retrieval-free optimization methods. We evaluate RASO through extensive experiments on four agent benchmarks, OfficeQA ( Opsahl-Ong et al., 2026 ) , SpreadsheetBench ( Ma et al., 2024 ) , ALFWorld ( Shridhar et al., 2021 ) , and WebShop ( Yao et al., 2022 ) , with two LLMs, Qwen-3.5-9B ( Qwen Team, 2026 ) and GPT-5.6-Luna ( OpenAI, 2026 ) . Specifically, RASI synthesizes strong initial skills that outperform both retrieved skills and retrieval-free initialization without requiring agent rollouts, providing a strong initialization for subsequent optimization. Moreover, RASO with RASU consistently outperforms strong skill optimization algorithms, including TextGrad ( Yuksekgonul et al., 2025 ) , GEPA ( Agrawal et al., 2026 ) , SkillOpt ( Yang et al., 2026 ) , and WikiSkill ( Tang et al., 2026 ) . Our contributions are as follows:

• We propose Retrieval-Augmented Skill Optimization (RASO) , a framework that leverages an external skill corpus as prior knowledge for both skill initialization and update. Through its two complementary stages, RASI and RASU , RASO grounds skill construction in retrieved procedural knowledge rather than relying on the optimizer’s parametric knowledge.

• We introduce Cross-Harness Adaptation, a shared operation that adapts retrieved procedural knowledge to the vocabulary of the target domain and harness. This enables knowledge transfer across diverse domains and harnesses without requiring domain- or harness-matched skills in the corpus.

• Across four benchmarks and two models, we demonstrate that RASI improves performance without requiring agent rollouts, while RASU further refines the skill by retrieving the missing knowledge guided by execution feedback. Our analysis further shows that these gains are associated with adapting diverse external knowledge rather than relying on particular source documents. 2 Related Works Agent Skills. Agent skills are reusable textual documents that provide procedural guidance for accomplishing tasks within an execution harness ( Anthropic, 2025 ) . Represented as text rather than model parameters, skills are easy to inspect, edit, share, and reuse across models without retraining. Prior work has made this knowledge explicit through executable skill libraries ( Wang et al., 2024 ) , reusable workflows ( Wang et al., 2025 ) , reflections and insights ( Shinn et al., 2023 ; Zhao et al., 2024 ) , or reasoning and procedural memories ( Ouyang et al., 2026 ; Fang et al., 2026 ) , while large public collections such as GitSkills ( Destefanis et al., 2026 ) now make this knowledge available at scale. However, not every skill is useful: its benefit depends on both the procedural knowledge it contains and how well that knowledge matches the target task and harness. Accordingly, two main approaches have emerged: learning procedural knowledge from the agent’s own execution experience and reusing knowledge from external sources . Skill Optimization from Execution Experience. A major line of work improves agent skill from its own experience, treating the skill as a textual decision variable optimized with execution feedback on the target task. Methods that optimize prompts or contexts using LLM-generated feedback ( Pryzant et al., 2023 ; Yang et al., 2024a ; Chu et al., 2026 ; Zhang et al., 2026 ) are directly applicable to skill optimization, most notably TextGrad ( Yuksekgonul et al., 2025 ) , which backpropagates textual feedback, and GEPA ( Agrawal et al., 2026 ) , which evolves prompts by reflecting on execution traces. Skill-specific optimizers follow the same recipe: SkillOpt ( Yang et al., 2026 ) edits a skill from rollout trajectories and accepts only edits that improve validation performance, WikiSkill ( Tang et al., 2026 ) compiles agent experience into persistent knowledge, and others distill or refine skills from trajectories ( Ni et al., 2026 ; Chen et al., 2026 ; Moll et al., 2026 ; Wang et al., 2026 ; Alzubi et al., 2026 ; Ding et al., 2026 ) . However, these methods primarily optimize skills from observed execution experience, limiting their exploration of procedural knowledge beyond what can be inferred from the agent’s own rollouts. In contrast, our framework supplements execution feedback with procedural knowledge retrieved from an external skill corpus throughout optimization. Skill Construction from External Knowledge. Procedural knowledge that an agent cannot infer from its own experience often already exists: in shared skills, documentation, and the web, motivating a growing line of work that draws on such external knowledge. Retrieval-based methods, such as SkillRouter ( Zheng et al., 2026 ) , select relevant skills from large libraries through improved retrievers ( Li et al., 2026a ; Miao et al., 2026 ) or by organizing libraries into graphs and execution structures ( Meng et al., 2026 ; Liu et al., 2026 ; Fu et al., 2026 ; Li et al., 2026b ; Xia et al., 2026 ) , with dedicated benchmarks for skill retrieval ( Kang et al., 2026 ; Su et al., 2026 ) . Since retrieved skills may refer to different tasks, tools, or actions, other methods adapt external experience to the target interface ( Tang et al., 2025 ) or compile external resources into reusable skills ( Pan et al., 2026 ; Yan et al., 2026 ) . However, these approaches either require repeated retrieval and adaptation for each task instance during test time or use external knowledge only during skill initialization, without further leveraging it as target-task experience accumulates. In contrast, we leverage external knowledge throughout the skill optimization process, using it both to construct the initial skill and to further improve the skill during iterative updates. 3 Problem Formulation We consider a frozen language model ℳ \mathcal{M} acting as an agent through an execution harness h h , which defines the available tools, file access, and observation interface. A skill s s is a natural-language artifact provided as the agent’s context to guide task completion under a given harness. Executing the agent on a task instance x x with skill s s yields a trajectory h ⁡ ( ℳ , x , s ) h(\mathcal{M},x,s) , which is associated with a reward r ⁡ ( ⋅ ) ∈ [ 0 , 1 ] r(\cdot)\in[0,1] by a benchmark-specific evaluator. When a reference answer is available, the evaluator compares the agent’s final output against it, and otherwise uses the environment’s native success criterion. We refer to each agent execution and its corresponding evaluation as a rollout . Given disjoint task splits 𝒟 train \mathcal{D}{\mathrm{train}} , 𝒟 val \mathcal{D}{\mathrm{val}} , and 𝒟 test \mathcal{D}{\mathrm{test}} , candidate skills are constructed from rollouts on 𝒟 train \mathcal{D}{\mathrm{train}} and selected based on performance on 𝒟 val \mathcal{D}{\mathrm{val}} as: s ⋆ = arg ​ max s ⁡ 𝔼 x ∼ 𝒟 val ​ [ r ⁡ ( h ⁡ ( ℳ , x , s ) ) ] . s^{\star}=\argmax{s}\mathbb{E}{x\sim\mathcal{D}{\mathrm{val}}}\left[r\left(h(\mathcal{M},x,s)\right)\right]. (1) The selected skill s ⋆ s^{\star} is then evaluated on the held-out 𝒟 test \mathcal{D}{\mathrm{test}} . Since both ℳ \mathcal{M} and h h remain fixed throughout optimization, the only decision variable is the natural-language skill s s . We include the construction of the initial skill itself in the skill optimization problem, rather than assuming that an initial skill is externally supplied. We therefore define skill optimization to encompass both skill initialization , which constructs an initial skill from task and harness descriptions without requiring any rollouts, and skill update , which refines the skill using results from training rollouts. 4 RASO: Retrieval-Augmented Skill Optimization Figure 2: Overview of RASO. RASI (top, without rollouts) generates retrieval queries from the task and harness descriptions, retrieves the top- K K skill sections per query from an external corpus, and adapts them into grounded lessons via Cross-Harness Adaptation to synthesize the initial skill s 0 s{0} . RASU (bottom, with rollouts) starts from s 0 s_{0} , executes the current skill, and generates retrieval queries from rollout results. The retrieved knowledge is adapted into grounded lessons via Cross-Harness Adaptation and incorporated into skill.md to iteratively refine the skill. Each subsequent iteration performs new rollouts using the committed skill. In this section, we introduce Retrieval-Augmented Skill Optimization (RASO) , a framework that evolves agent skills through two complementary stages: skill initialization and skill update, illustrated in Figure 2 . Both stages share a common knowledge retrieval and adaptation mechanism but differ in the rollout evidence available to guide skill optimization. Section 4.1 describes the shared mechanism, which retrieves procedural knowledge from a large-scale skill corpus ( Destefanis et al., 2026 ) and adapts it to the target task and harness through cross-harness adaptation . Section 4.2 introduces Retrieval-Augmented Skill Initialization (RASI) , which constructs an initial skill using only the target task, harness description, and retrieved knowledge, without requiring agent rollouts. Section 4.3 introduces Retrieval-Augmented Skill Update (RASU) , which leverages execution feedback and retrieved knowledge to iteratively refine the skill. 4.1 Skill Retrieval and Cross-Harness Adaptation We first perform fine-grained, section-level skill retrieval from an external skill corpus to acquire relevant prior knowledge. We then introduce Cross-Harness Adaptation, a mechanism that bridges the gap between source and target domains and harnesses by transforming retrieved knowledge into actionable guidance tailored to the target task and harness. Skill retrieval and adaptation constitute a shared pipeline used by both RASI (Section 4.2 ) and RASU (Section 4.3 ). Section-level skill retrieval. We first divide skill documents into heading-delimited sections to enable fine-grained retrieval. Since external skill documents are developed for diverse tasks and workflows, retrieving entire documents may introduce irrelevant content and favor documents with similar overall objectives over those containing relevant procedural sections. Section-level retrieval instead enables us to identify relevant procedural knowledge while improving the signal-to-noise ratio of retrieved content. Given the resulting section-level corpus 𝒞 \mathcal{C} , a query-generation agent formulates a query q q and retrieves the top- K K most relevant sections using BM25 ( Robertson and Zaragoza, 2009 ) , i.e. , 𝒮 q = BM25 ​ ( q , 𝒞 , K ) \mathcal{S}{q}=\text{BM25}(q,\mathcal{C},K) . Cross-Harness Adaptation. We employ an adaptation agent, denoted by 𝒜 adaptation \mathcal{A}{\mathrm{adaptation}} , to adapt retrieved knowledge to the target task and harness. Let c i c_{i} denote a requirement needing external knowledge to be resolved, such as a specific task procedure in RASI or a textual gradient in RASU. Given c i c_{i} , the task description T T , the harness description H H , and the corresponding top- K K retrieved sections 𝒮 q i \mathcal{S}{q{i}} , the agent produces a concise, actionable lesson ℓ i \ell_{i} : ℓ i = 𝒜 adaptation ​ ( c i , T , H , 𝒮 q i ) . \ell_{i}=\mathcal{A}{\mathrm{adaptation}}\left(c{i},T,H,\mathcal{S}{q{i}}\right). (2) Here, the lesson ℓ i \ell_{i} serves as a refined knowledge snippet that directly guides the agent to handle the requirement c i c_{i} within the target task and harness. To ensure that each lesson addresses the given requirement and remains valid within the target harness, the adaptation process follows three principles: (1) remove domain-specific nouns and omit procedures without counterparts in the target harness, (2) focus exclusively on requirement c i c_{i} , excluding unrelated issues, and (3) preserve specific claims about tool or parameter behavior only when corroborated by H H , prioritizing correctness within the target harness over potentially inaccurate specificity. Consequently, each lesson is expressed using the objects, commands, and units of the target task and harness. 4.2 Retrieval-Augmented Skill Initialization (RASI) We introduce Retrieval-Augmented Skill Initialization (RASI) , which constructs an initial skill for a target task under a given harness without requiring agent rollouts. Performed once at the beginning of skill optimization, RASI comprises four sequential steps: (1) procedure and query generation, (2) section-level skill retrieval, (3) Cross-Harness Adaptation, and (4) skill initialization. Procedure and query generation. Given the task description T T and harness description H H , the query-generation agent 𝒜 query ​ _ ​ init \mathcal{A}{\mathrm{query_init}} generates a set of requirement-query pairs ( c i , q i ) (c{i},q_{i}) : { ( c i , q i ) } i = 1 M = 𝒜 query ​ _ ​ init ​ ( T , H ) , {(c_{i},q_{i})}{i=1}^{M}=\mathcal{A}{\mathrm{query_init}}(T,H), (3) where M M denotes the number of generated pairs. Here, each c i c_{i} represents a specific task procedure or harness constraint ( e.g., multi-turn budget management), and q i q_{i} is the retrieval query created to search for external skills that address c i c_{i} . Section-level skill retrieval and Cross-Harness Adaptation. Given the generated requirement-query pairs, RASI applies the shared retrieval and adaptation pipeline described in Section 4.1 . For each query q i q_{i} , BM25 retrieves the top- K K relevant sections 𝒮 q i \mathcal{S}{q{i}} from the corpus 𝒞 \mathcal{C} . The adaptation agent 𝒜 adaptation \mathcal{A}{\mathrm{adaptation}} then transforms these sections into a grounded lesson ℓ i = 𝒜 adaptation ​ ( c i , T , H , 𝒮 q i ) \ell{i}=\mathcal{A}{\mathrm{adaptation}}(c{i},T,H,\mathcal{S}{q{i}}) . The resulting lessons form ℒ init = { ℓ 1 , … , ℓ M } \mathcal{L}{\mathrm{init}}={\ell{1},\dots,\ell_{M}} , which are used for subsequent skill initialization. Skill initialization. Finally, a skill-initializer agent 𝒜 skill ​ _ ​ init \mathcal{A}{\mathrm{skill_init}} synthesizes the initial skill s 0 s{0} from the task description T T , harness description H H , identified procedures and requirements { c i } i = 1 M {c_{i}}{i=1}^{M} , and grounded lessons ℒ init \mathcal{L}{\mathrm{init}} . The agent integrates each lesson ℓ i \ell_{i} into the execution step corresponding to c i c_{i} , yielding: s 0 = 𝒜 skill ​ _ ​ init ​ ( T , H , { c i } i = 1 M , ℒ init ) . s_{0}=\mathcal{A}{\mathrm{skill_init}}(T,H,{c{i}}{i=1}^{M},\mathcal{L}{\mathrm{init}}). (4) By incorporating retrieved and adapted knowledge before environment interaction, RASI provides a knowledge-grounded initial skill for subsequent optimization without consuming search rollouts. 4.3 Retrieval-Augmented Skill Update (RASU) Here, we introduce Retrieval-Augmented Skill Update (RASU) that iteratively refines the current skill s t s_{t} using execution feedback from agent rollouts. Complementing the rollout-free initialization of RASI, RASU identifies specific failure modes observed in agent trajectories and retrieves relevant external knowledge to address them. Given trajectories generated by the execution agent using s t s_{t} , each RASU iteration comprises four sequential steps: (1) textual gradient and query generation, (2) section-level skill retrieval, (3) Cross-Harness Adaptation, and (4) skill update. Textual gradient and query generation. We first sample a minibatch of tasks from 𝒟 train \mathcal{D}{\mathrm{train}} and execute agent rollouts using the current skill s t s{t} . A gradient-and-query generator agent 𝒜 query ​ _ ​ update \mathcal{A}{\mathrm{query_update}} then analyzes the failed trajectories to identify failure mode. For each failure modes, the agent generates a textual gradient δ i \delta{i} based on its parametric knowledge and a targeted retrieval query q i q_{i} to acquire relevant external knowledge: { ( q i , δ i ) } i = 1 N = 𝒜 query ​ _ ​ update ​ ( T , H , s t , { τ j } ) , {(q_{i},\delta_{i})}{i=1}^{N}=\mathcal{A}{\mathrm{query_update}}(T,H,s_{t},{\tau_{j}}), (5) where N N denotes the number of generated query and gradient pairs, and { τ j } {\tau_{j}} denotes the rollout trajectories. Section-level skill retrieval and Cross-Harness Adaptation. Given the failure-driven retrieval queries, RASU applies the shared retrieval and adaptation pipeline described in Section 4.1 . This process yields a set of grounded lessons ℒ update = { ℓ 1 , … , ℓ N } \mathcal{L}{\mathrm{update}}={\ell{1},\dots,\ell_{N}} for subsequent skill update. Skill update. Finally, a skill-updater agent 𝒜 skill ​ _ ​ update \mathcal{A}{\mathrm{skill_update}} generates a candidate skill s t + 1 s{t+1} by integrating the current skill s t s_{t} , textual gradients { δ i } i = 1 N {\delta_{i}}{i=1}^{N} , and trajectory-grounded lessons ℒ update \mathcal{L}{\mathrm{update}} : s t + 1 = 𝒜 skill ​ _ ​ update ​ ( s t , { δ i } i = 1 N , ℒ update ) . s_{t+1}=\mathcal{A}{\mathrm{skill_update}}(s{t},{\delta_{i}}{i=1}^{N},\mathcal{L}{\mathrm{update}}). (6) The updater refines the current skill s t s_{t} to candidate skill s t + 1 s_{t+1} based on the textual gradients and grounded lessons. The candidate skill s t + 1 s_{t+1} is accepted only if it outperforms s t s_{t} on the validation set 𝒟 val \mathcal{D}{\text{val}} , with s t s{t} retained otherwise. This pipeline is repeated for a fixed number of iterations, progressively refining the skill through execution feedback and retrieved prior knowledge. 5 Experiment Table 1: Agent performance with GPT-5.6-Luna ( OpenAI, 2026 ) and Qwen-3.5-9B ( Qwen Team, 2026 ) . We separate skill initialization from skill update; initialization-only rows report performance at 0 rollouts. RFSI denotes retrieval-free skill initialization. Results are mean ± \pm standard error over three random seeds. Benchmark Model Method OfficeQA Spreadsheet ALFWorld WebShop Skill initialization (without rollout) No skill 11.44 ± \pm 0.70 32.98 ± \pm 0.12 64.43 ± \pm 1.38 43.56 ± \pm 1.01 SkillRouter 11.44 ± \pm 0.97 33.21 ± \pm 0.94 55.97 ± \pm 1.14 42.20 ± \pm 0.37 RFSI 40.11 ± \pm 0.58 44.40 ± \pm 2.74 69.40 ± \pm 1.55 43.89 ± \pm 1.17 RASI 45.74 ± \pm 1.40 49.17 ± \pm 1.52 72.64 ± \pm 1.63 45.06 ± \pm 0.40 Skill update (with rollout) TextGrad 43.80 ± \pm 1.27 49.64 ± \pm 3.09 70.65 ± \pm 0.66 42.32 ± \pm 3.13 GEPA 43.99 ± \pm 1.40 54.53 ± \pm 1.46 71.14 ± \pm 1.08 45.54 ± \pm 0.35 SkillOpt 45.54 ± \pm 2.05 57.02 ± \pm 3.60 72.64 ± \pm 0.90 45.17 ± \pm 0.23 WikiSkill 41.47 ± \pm 1.18 47.14 ± \pm 3.98 71.14 ± \pm 0.66 44.40 ± \pm 0.64 GPT-5.6-Luna RASO 49.03 ± \pm 1.85 63.33 ± \pm 2.69 74.13 ± \pm 1.00 46.61 ± \pm 0.11 Skill initialization (without rollout) No skill 33.14 ± \pm 1.21 28.81 ± \pm 0.60 33.09 ± \pm 0.66 16.27 ± \pm 0.25 SkillRouter 34.11 ± \pm 0.51 23.45 ± \pm 0.48 31.34 ± \pm 1.29 9.42 ± \pm 0.47 RFSI 34.89 ± \pm 2.93 27.74 ± \pm 2.56 42.04 ± \pm 2.87 12.54 ± \pm 2.56 RASI 40.50 ± \pm 0.97 30.48 ± \pm 0.78 47.76 ± \pm 3.53 23.27 ± \pm 0.98 Skill update (with rollout) TextGrad 36.82 ± \pm 1.59 25.95 ± \pm 0.63 41.04 ± \pm 2.28 10.96 ± \pm 4.44 GEPA 36.43 ± \pm 1.97 24.05 ± \pm 1.56 43.53 ± \pm 1.99 12.80 ± \pm 2.20 SkillOpt 37.21 ± \pm 1.21 29.52 ± \pm 2.21 43.03 ± \pm 3.17 13.36 ± \pm 1.70 WikiSkill 37.40 ± \pm 3.03 29.76 ± \pm 1.80 42.79 ± \pm 2.93 13.43 ± \pm 2.01 Qwen-3.5-9B RASO 42.25 ± \pm 1.52 31.55 ± \pm 0.31 51.00 ± \pm 2.45 24.73 ± \pm 0.41 We evaluate RASO using two target LLMs: GPT-5.6-Luna ( OpenAI, 2026 ) and Qwen-3.5-9B ( Qwen Team, 2026 ) . We refer to the model that executes tasks as the target model and the model that generates or updates skill text as the optimizer model. Unless otherwise specified, we use the same model for both roles across all skill optimization methods. For skill retrieval, we use GitSkills ( Destefanis et al., 2026 ) , an external skill corpus spanning diverse domains and harnesses. Our evaluation covers four benchmarks with diverse interaction settings: OfficeQA ( Opsahl-Ong et al., 2026 ) , SpreadsheetBench ( Ma et al., 2024 ) , ALFWorld ( Shridhar et al., 2021 ) , and WebShop ( Yao et al., 2022 ) . For skill initialization, we compare RASI against three strategies: (1) No Skill, where the agent operates without an initialized skill, (2) SkillRouter ( Zheng et al., 2026 ) , which retrieves a relevant skill from the external corpus, and (3) Retrieval-Free Skill Initialization (RFSI), where the target LLM generates an initial skill directly from the task and harness descriptions without access to the external skill corpus. For iterative skill optimization, we compare RASO against TextGrad ( Yuksekgonul et al., 2025 ) , GEPA ( Agrawal et al., 2026 ) , SkillOpt ( Yang et al., 2026 ) , and WikiSkill ( Tang et al., 2026 ) . All experiments are conducted with three random seeds, and we report the mean performance over seeds. 5.1 Main Results Skill initialization. We first evaluate the quality of skills produced by different initialization methods, i.e. , before any subsequent agent rollout or skill optimization, for both GPT-5.6-Luna and Qwen-3.5-9B in Table 1 . Across both models and all four benchmarks, RASI consistently achieves the highest performance. With GPT-5.6-Luna, RASI improves over Retrieval-Free Skill Initialization (RFSI) by +5.63 on OfficeQA (45.74 vs. 40.11), +4.77 on Spreadsheet (49.17 vs. 44.40), +3.24 on ALFWorld (72.64 vs. 69.40), and +1.17 on WebShop (45.06 vs. 43.89). The margin is substantially larger over the No Skill and SkillRouter baselines, particularly on OfficeQA, where both achieve only 11.44. This advantage also holds for the smaller Qwen-3.5-9B backbone, where RASI outperforms RFSI by +5.61 on OfficeQA, +2.74 on Spreadsheet, +5.72 on ALFWorld, and +10.73 on WebShop. SkillRouter, which retrieves external skills without Cross-Harness Adaptation, underperforms even the No Skill baseline on several benchmarks, which suggests direct reuse can introduce irrelevant or mismatched procedural knowledge when the retrieved content is not adapted to the target task and harness. Overall, these results show that combining external knowledge retrieval with LLM-based adaptation, as in RASI, yields a more effective initialization than either retrieval alone or retrieval-free LLM-based skill generation, providing a stronger skill initialization for subsequent optimization. Skill Update We compare RASO with existing skill optimization methods in Table 1 . Following the conventional skill optimization setup, existing methods start from RFSI, where an initial skill is LLM-generated without rollouts or access to an external corpus, and subsequently refine it using their respective rollout-based optimization procedures. RASO overall outperforms existing skill optimization methods across both backbones. For GPT-5.6-Luna, RASO improves over the strongest competing method by +3.49 on OfficeQA (49.03 vs. 45.54), +6.31 on SpreadsheetBench (63.33 vs. 57.02), +1.49 on ALFWorld (74.13 vs. 72.64), and +1.07 on WebShop (46.61 vs. 45.54). Similarly, with Qwen-3.5-9B, RASO improves over the strongest competing method by +4.85 on OfficeQA (42.25 vs. 37.40), +1.79 on SpreadsheetBench (31.55 vs. 29.76), +7.47 on ALFWorld (51.00 vs. 43.53), and +11.30 on WebShop (24.73 vs. 13.43). These results show that RASO benefits from both retrieval-augmented initialization and subsequent retrieval-augmented skill refinement. 5.2 Analysis Table 2: Ablation on the initialization (Init.) and update components with GPT-5.6-Luna. Method Benchmark Init. Update OfficeQA Spreadsheet RFSI RFSU 40.70 ± \pm 1.67 51.67 ± \pm 2.03 RFSI RASU 47.56 ± \pm 1.85 61.07 ± \pm 2.15 RASI RFSU 45.93 ± \pm 1.58 58.45 ± \pm 1.78 RASI RASU 49.03 ± \pm 1.85 63.33 ± \pm 2.69 Table 3: Effect of adaptation with GPT-5.6-Luna. RASU used RASI with adaptation as an initial. Benchmark Method The ai search story also surfaces in Stanford AI discovery identifies natural weight..., adding another angle. as detailed in the full paper on Arxiv The ai search story also surfaces in Google AI Releases TimesFM 3 for..., adding another angle.

Comments (0)

No comments yet

Be the first to share your thoughts!