What the paper is about
Agent performance depends on both reasoning ability and the environment in which it acts. We study test-time AI-for-AI, asking how a Builder can learn to construct better execution environments for a Target while both models' weights remain fixed. To make the Builder's experience reusable, we introduce Meta-Skill: principles specifying when support is needed and what resources to provide. The Builder learns these principles from Target's execution feedback on the development set, then uses the frozen skill bank to construct harnesses for unseen tasks. Across Harness-Bench and NewtonBench, full-bank meta-skills improve macro-average performance by 8.95 percentage points over no-skill construction, and 12.02 points over direct delivery of the same bank to the Target. These results highlight the value of translating experience into executable support. Gains when the same model serves both roles further suggest a path to system level self-improvement through learning to build better environments. The ai agents story also surfaces in AI Agents Going Rogue Renew Calls..., adding another angle.
What it covers
Learning Meta-Skills for Agent Harness Design in Test-Time AI4AI Cheng Qian Affiliation: Apodex Affiliation: University of Illinois Urbana Champaign Kunlun Zhu Affiliation: Apodex Affiliation: University of Illinois Urbana Champaign Beibin Li Affiliation: Apodex Zhenhailong Wang Affiliation: Apodex Heng Ji Affiliation: Apodex Affiliation: University of Illinois Urbana Champaign Abstract Agent performance depends on both reasoning ability and the environment in which it acts. We study test-time AI-for-AI, asking how a Builder can learn to construct better execution environments for a Target while both models’ weights remain fixed. To make the Builder’s experience reusable, we introduce meta-skills : principles specifying when support is needed and what resources to provide. The Builder learns these principles from Target’s execution feedback on the development set, then uses the frozen skill bank to construct harnesses for unseen tasks. Across Harness-Bench and NewtonBench, full-bank meta-skills improve macro-average performance by 8.95 percentage points over no-skill construction, and 12.02 points over direct delivery of the same bank to the Target. These results highlight the value of translating experience into executable support. Gains when the same model serves both roles further suggest a path to system level self-improvement through learning to build better environments. † † footnotetext: Corresponding author: [email protected] . We also gratefully acknowledge Mr. Tianqiao Chen for the guidance and support throughout this work. The project code is released at https://github.com/qiancheng-apodex/MetaSkill-AI4AI . 1 Introduction An AI agent’s performance depends on both its reasoning ability and the environment in which it acts [ 11 ] . Consider a capable PhD student who spends days repeating experiments because configurations and results are scattered across scripts and notes. An effective advisor can address this bottleneck by establishing a shared experiment log, reproducible tools, and a clear validation workflow, helping the student devote more effort to scientific judgment. This analogy suggests a complementary direction for improving AI agents: learning how to provide the support that makes their existing capabilities more effective [ 3 , 13 ] . AI-for-AI (AI4AI) at test-time offers a route to this goal by enabling AI systems to design agent programs, workflows, and harnesses [ 1 , 15 , 3 ] . This form of AI4AI involves two complementary roles: a Builder , which designs and provides support, and a Target , which uses that support to solve tasks. We study this relationship through harness construction, where the Builder creates a harness comprising instructions, resources, and executable mechanisms for the Target. Just as an advisor learns from a PhD student’s progress to refine their guidance, the Builder can learn from the Target’s execution to improve its support. Building on recent work on meta-level learning [ 13 ] , we therefore ask: how can the Builder turn the outcomes of its own harnesses into reusable knowledge for supporting future tasks? To make this experience reusable, we distinguish the knowledge needed by the Builder from task skills used by the Target . Task skills describe how to perform a task [ 9 ] ; the Builder instead needs principles for designing the support that helps the Target perform it. We call these reusable support principles meta-skills . Empirically, each meta-skill shuold specify when support is needed, what capability or resource to provide , and how the Target should use it while retaining responsibility for judgment. In this paper, we connect the learning and implementation of these principles through a construction–execution–reflection loop. Starting with an empty meta-skill bank, the Builder constructs harnesses for development tasks, reviews the resulting Target execution records and scores, and adds or revises principles based on the observed evidence. After skill learning, the bank is frozen and guides fresh harness construction for each held-out task. Learning thus changes the Builder’s external knowledge and the support it constructs, while both the Builder’s and Target’s model weights remain fixed. We evaluate three Targets on Harness-Bench and NewtonBench under fixed Target execution budgets. The full-bank meta-skill Builder achieves a 65.31% macro-average score, exceeding the no-skill Builder by 8.95 percentage points and full-bank independently learned Target skills by 10.93 points. It also outperforms direct delivery of the same bank to the Target in all six settings. These comparisons suggest that meta-skills can help the Builder become a more effective advisor: they guide support design, while Target skills guide task execution. In our setting, teaching the Builder to translate experience into executable support can therefore be more beneficial than directly teaching the Target additional task skills. Because meta-skills encode reusable principles, we further examine how they develop and whether they remain useful beyond the setting in which they were learned. Through analysis, we discover that repeated reflection can strengthen the guidance, and transfer studies suggest possible meta-skill reuse across Builders and Targets. The Builder and Target can also be instances of the same model: across three such settings, meta-skills improve scores by 18.71 points on average over no-skill construction. These results suggest that a model can improve its own execution by learning to build better support for itself, establishing harness design as a promising route to system-level self-improvement. Looking ahead, AI4AI broadens the goal of agent learning: agents can learn both to solve problems and to create the conditions for others to succeed. Improving how Builders provide guidance and resources may be as consequential as improving how Targets use them. Better agents may begin with better advisors. Figure 1 : Overview of our framework. Rather than teaching the Target directly, our framework equips the Builder with meta-skills learned from Target execution feedback. These reusable principles guide the Builder in constructing harnesses for new test tasks. 2 Related Work AI for AI in Agentic Systems. AI-for-AI uses AI systems to automate model development and agent design. For model development, AIDE uses an LLM agent to write and refine machine-learning pipelines, while MLE-Dojo provides environments for training and evaluating agents performing this engineering work [ 2 , 6 ] . For test-time agent design, the optimization target becomes the system surrounding a fixed model: Automated Design of Agentic Systems and AFlow both focus on searching agent programs and workflows [ 1 , 15 ] , while Darwin Godel Machine evolves its own agent code, and Meta-Harness searches harness implementations using execution feedback [ 14 , 3 ] . Within this direction, strong-to-weak harness construction uses a stronger Builder to supply executable support for a weaker Target [ 5 ] . We build on this setting by learning reusable construction principles from Target executions during offline skill learning, then freezing them to guide fresh harness construction for each test task, with both models’ weights fixed. Agent Skills and Meta-Skills. Agent skills preserve procedural knowledge for reuse: Voyager stores executable routines, while ExpeL and Agentic Context Engineering accumulate textual insights and contextual playbooks [ 7 , 17 , 16 ] . Methods for improving these resources include SkillRL, which co-evolves skills and policies through reinforcement learning, and Evo-Harness, which compiles execution experience into transferable solver skills [ 10 , 9 ] . SkillsBench complements these methods by measuring skill utility with curated packages and deterministic verifiers [ 4 ] . At the meta level, learned knowledge guides agent improvement itself: Meta Context Engineering learns skills for constructing context files and code, while MetaSkill-Evolve evolves policies for improving task skills [ 13 , 8 ] . Within this direction, our Builder learns those principles from its own harness outcomes and implements them as task-specific environments for a separate Target , specifying when support is needed, what capabilities to supply, and the Target’s responsibilities. 3 Problem Setting and Method Problem Formulation. Let ℬ \mathcal{B} and 𝒯 \mathcal{T} be fixed Builder and Target models. Each benchmark provides a baseline environment H 0 H_{0} with native capabilities, an evaluator r r , and a Target execution budget C x C_{x} . A harness policy ℋ \mathcal{H} maps each public task input x x to an environment H x = ℋ ( x ) H_{x}=\mathcal{H}(x) that augments H 0 H_{0} with executable support. The objective is to maximize expected test performance: J ( ℋ ) = 𝔼 x ∼ 𝒟 test , , τ ∼ 𝒯 ( ⋅ ∣ x , H x ) [ r ( x , τ ) ] , cost ( τ ) ≤ C x . J(\mathcal{H})=\mathbb{E}{x\sim\mathcal{D}{\mathrm{test}},,\tau\sim\mathcal{T}(\cdot\mid x,H_{x})}[r(x,\tau)],\qquad\operatorname{cost}(\tau)\leq C_{x}. (1) The Builder learns the policy through Target executions on development instances and is evaluated on held-out test instances. The budget C x C_{x} covers only Target execution, excluding Builder computation for harness construction. Method Overview. Instead of directly improving the Target , we teach the Builder to learn what support the Target needs and provide it through a harness. The Builder’s experience accumulates in an external bank of meta-skills , which records reusable principles for designing support across tasks. These principles guide the harness construction, while both the Builder and Target model weights remain fixed throughout learning. Table 1: Overview of the harness refinement space, including neutral defaults in H 0 H_{0} , Builder-editable mechanisms, and illustrative examples. The underlying benchmark tools, evaluation budget, and scoring rules remain fixed. Instructions Memory Context Composed tools Execution control Verification & recovery Workspace Initial H 0 H_{0} Generic prompt Empty store Default history Native tools Default loop Native rules Empty scratch Builder ℬ \mathcal{B} refines what? Specific task guidance Record format; storage and retrieval rules History-selection rules Tool definitions; call sequences Execution phases; tool visibility Submission checks; failure handling Initial files; setup actions Example Implement “Draft before polishing” Recent trial buffer Budgeted history window Cross-file auditor Hide failing tools Audit-gated submission Templates; validator scripts 3.1 Learning Meta-Skills from Target Execution During Development Definition of Meta-Skill. A meta-skill s = ( when , provide , use ) s=(\mathrm{when},\mathrm{provide},\mathrm{use}) contains three fields: when identifies observable conditions that call for support; provide specifies the capability or resource the environment should supply; and use explains how the Target should employ that support and which judgments remain its responsibility. Skill Learning Workflow. Skill learning begins with an empty skill bank S 0 S_{0} . For each development task input x x , the Builder constructs a separate harness using the neutral environment H 0 H_{0} , public component interfaces, and its complete current bank. The Target executes within the harness, producing a public execution record and a benchmark-native development score. The same Builder then updates its skill bank by reviewing the current bank, its generated programs, and public execution feedback F ( e ) F(e) . This process will iterate until a fixed budget of development-set passes is reached, as formalized below: H x j = ℬ ( H 0 , x , S j ) , e x j = 𝒯 ( x ; H x j ) , S j + 1 = Revise ℬ ( S j , { ( H x j , F ( e x j ) ) } x ∈ G j ) . \displaystyle H_{x}^{j}=\mathcal{B}(H_{0},x,S_{j}),\qquad e_{x}^{j}=\mathcal{T}(x;H_{x}^{j}),\qquad S_{j+1}=\operatorname{Revise}{\mathcal{B}}!\left(S{j},{(H_{x}^{j},F(e_{x}^{j}))}{x\in G{j}}\right). (2) Skill Bank Update. At the end of each skill-learning step, the Builder reviews its generated harnesses and the Target’s execution feedback to identify an observed error or recurring burden that better support could address, then compares the resulting lesson with the current bank. It revises an existing meta-skill when the evidence corrects or strengthens its guidance, adds a new meta-skill for a distinct reusable support need, or keeps the bank unchanged when no update is justified. Each batch permits at most one addition or revision, which must cite supporting evidence from that batch to keep the learned guidance grounded in observed behavior. 3.2 Building New Harness for Every Task at Test-Time Test-Time Skill Selection. After skill learning, we freeze the meta-skill bank and use it to guide test-time harness construction. For each test task, the Builder receives skills through one of two modes: full bank places all meta-skills in its context; retrieval uses a fixed BM25 retriever to select at most two skills with positive relevance scores against the task’s initial public prompt. For each individual test task, the Builder uses the supplied meta-skills to construct a new, task-specific harness in a fresh environment for the Target to execute within: S ∗ = S J , H x ∗ = ℬ ( H 0 , x , K ( x , S ∗ ) ) , y ^ x = 𝒯 ( x , H x ∗ ) , S^{}=S_{J},\qquad H_{x}^{}=\mathcal{B}(H_{0},x,K(x,S^{})),\qquad\hat{y}{x}=\mathcal{T}(x;H{x}^{}), (3) where K K supplies either retrieved meta-skills or the full bank. Harness Components. As summarized in Table 1 , the framework exposes seven optional harness component families: instructions, memory, context organization, composed tools, execution control, verification and recovery, and workspace preparation. The Builder selects and implements these components, deciding what memory stores, when it is retrieved, and which tools and resources the Target receives. Within the resulting harness, the Target reasons over observations, uses the available tools, and remains responsible for the final submission. Builder’s Working Environment. The Target operates within a Builder-created harness, while the Builder operates within a framework we design. Specifically, we provide shared interfaces, execution isolation, design guidance, and resource limits, with identical construction permissions and bounded interface-repair opportunities across conditions. These fixed priors underpin our workflow for automated meta-skill learning, harness construction, and Target execution. Our contribution thus shifts design effort to the Builder level, enabling it to turn execution experience into reusable support principles and adapt their implementation to individual tasks. Example. Consider an artifact edited after validation, making the earlier check potentially outdated. A meta-skill’s when field identifies this condition; provide requests version tracking and revalidation; and use directs the Target to inspect the current validation result before submitting, while retaining responsibility for content judgment and correction. The Builder could then implement this principle with file-version memory and a validation tool. In Appendix we trace two actual test episodes from learned principles through Builder-written code and Target tool calls to final artifacts. 4 Experiments 4.1 Experimental setup Table 2: Benchmark statistics and splits. Benchmark Tasks Development Test Harness-Bench 106 11 95 NewtonBench 324 32 292 Datasets. We evaluate the effect of test-time learned support on task performance using Harness-Bench and NewtonBench, which cover agent workflows and interactive scientific law discovery, respectively [ 12 , 18 ] . For each benchmark, we reserve 10% of tasks for learning and use the remainder exclusively for evaluation. See Table 2 for details. Models and Settings. We use GPT-5.6-Sol as the Builder model, and evaluate Gemini-3.6-Flash, Qwen3.8-Flash, and GPT-OSS-120B as Target models. For each Builder–Target–benchmark combination, the meta-skill bank starts empty and is updated over two development-set passes, with at most one evidence-grounded keep , revise , or add update per task. The bank is then frozen for testing, while the Builder constructs a fresh harness for each test task. All models use temperature zero and high reasoning effort (or the corresponding thinking mode), with a 16K output-token limit for harness generation and an 8K limit for skill update. For the benchmark budget, Harness-Bench allows 30 model turns, 30 tool calls, and 96K cumulative tokens per task; NewtonBench allows 12 turns, 10 tool calls, and 192K tokens. Metrics. Harness-Bench reports the deterministic completion-oracle score over 95 test tasks, averaged and scaled to a percentage. NewtonBench reports symbolic-structure accuracy over 292 test tasks. Both use end-to-end evaluation, divided by the fixed number of test cases: valid native outcomes are retained, while audited harness-construction failures receive a score of zero because no executable harness is produced. Baselines. We compare five baselines. Learned guidance is supplied either in full or through BM25 top-2 retrieval.
• Native environment: The Target solves tasks in the shared neutral environment, without learned guidance or Builder-generated support.
• No-skill Builder: The Builder constructs harnesses with the same construction space and budget as our method, but with an empty skill bank.
• Direct Builder skills: The Target receives our Builder’s meta-skills directly as instructions, with no Builder-generated harness components.
• Independent structured Target skills: An equally budgeted GPT-5.6-Sol learner iteratively extracts and refine structured problem-solving skills from the Target’s own neutral-environment executions. The resulting bank is supplied directly to the Target.
• Mined free-form Target skills: GPT-5.6-Sol consolidates the Target’s neutral-environment development trajectories into free-form notes in one offline pass. These notes are supplied directly to the Target. Our Builder meta-skill conditions instead provide the retrieved or full bank to the Builder, which compiles it into task-specific harness components such as instructions, memory, context, tools, controllers, verification, or workspace preparation. The full-bank Direct baseline therefore controls for semantic knowledge, while No-skill Builder controls for construction capability, serving as complementary ablations of our method. Please see Appendix C for more setting details. 4.2 Main results Table 3: Test performance (%) on every task in the fixed test splits with GPT-5.6-Sol as Builder in all settings. Bold marks the best result in each model–benchmark column. Method Harness-Bench NewtonBench Macro Avg. Gemini Qwen GPT-OSS Gemini Qwen GPT-OSS Native environment 37.35 69.63 52.46 55.48 54.11 39.38 51.40 Builder, no skills 53.29 71.51 63.02 56.16 50.34 43.84 56.36 Direct Builder skills, retrieved 46.66 68.98 59.13 58.22 48.29 33.22 52.42 Direct Builder skills, all 42.38 68.70 58.34 58.56 55.82 35.96 53.29 Independent Structured Target skills, retrieved 41.95 72.51 45.21 67.81 53.42 36.30 52.87 Independent Structured Target skills, all 40.59 72.89 56.63 67.81 52.40 35.96 54.38 Mined Free-form Target skills, retrieved 39.61 63.51 71.12 46.92 54.45 33.90 51.59 Mined Free-form Target skills, all 40.70 66.10 66.01 41.44 54.11 39.04 51.23 Builder meta-skills, retrieved 61.50 74.81 64.74 64.04 57.88 39.73 60.45 Builder meta-skills, all 67.81 72.64 68.22 68.84 63.70 50.68 65.31 We present the main results and baselines in Table 3 , and highlight the following key findings. Meta-skills are most useful when the teacher can enact them. Giving the same full meta-skill bank to the Builder consistently outperforms giving it directly to the Target , improving every model–benchmark pair by up to 25.43 points and 12.02 on average. This gap shows that the gains come not merely from exposing the Target to better knowledge, but from operationalizing that knowledge before execution. The Builder can translate a declarative meta-skill into persistent state, executable tools, verification logic, or control decisions, reducing the burden on the Target to interpret and apply the advice correctly at inference time. Meta-skills therefore function less as solver prompts and more as a compact language for provisioning task-specific support. Experience improves the Builder beyond construction capability. With construction capability held fixed, the full-bank Builder outperforms the no-skill Builder in all six settings, with an average gain of 8.95 points. These results suggest that experience adds value beyond the ability to construct scaffolds: it helps the Builder decide which support to build and how to integrate it with Target behavior. The gains nevertheless vary across Targets and benchmarks, suggesting that accumulated guidance must be adapted to the Target and task rather than assumed uniformly useful. Experience helps most when coordination is the main bottleneck. Relative to the no-skill Builder, gains on NewtonBench are positive across all three Targets and average 10.96 points, versus 6.95 points on Harness-Bench. NewtonBench requires agents to coordinate experimentation, reasoning, and valid symbolic submission, creating recurring coordination failures that scaffolding can address. Harness-Bench spans more heterogeneous workflows, where effective support depends more on Target-specific planning and tool use . This contrast suggests that accumulated experience is most useful when failure modes recur consistently across tasks and Targets. Broad skill access usually outperforms sparse retrieval. The full skill bank outperforms top-2 retrieval in five of six settings, with an average gain of 7.19 points on NewtonBench. This pattern suggests that a compact meta-skill bank offers complementary procedures whose combined value may be missed by lexical retrieval. Full-bank access allows the Builder to consider these procedures together when designing support, making it a strong default at the evaluated scale. The exception, however, indicates that broader access is not uniformly beneficial and motivates retrieval methods that account for both skill complementarity and the Target’s likely response to support. 5 Analysis Figure 2 : Test performance as Builder experience is refined. The two NewtonBench curves improve mainly after the second pass; Harness-Bench displays both steady improvement and late regression. 5.1 Effect of Meta-skill Refinement Motivation and setting. Meta-skills develop through repeated Builder reflection, but additional updates need not improve teaching. To examine this progression, we freeze the skill bank after zero, one, or two complete development passes and evaluate each version on the full test split. Figure 2 reports results for both Gemini and Qwen Targets. Useful teaching behavior can emerge late. On NewtonBench, the first pass changes each Target’s score by less than 1.1 points, whereas the second adds 13.01 points for Gemini and 12.33 for Qwen. Across all settings, the average gain from zero to two passes is 10.42 points. Manual inspection suggests that early update gathers local observations, while later revision connects them into reusable interventions . This delayed improvement is consistent with the value of iterative revision. The lower Mined Target skill score also motivates going beyond offline trajectory summarization, although that comparison changes both the learning procedure and skill recipient. Refinement is not monotonic. On Harness-Bench, Gemini improves with each pass, but Qwen drops 4.06 points from its first-pass peak. Revisions can therefore sharpen useful guidance while also making it overly specific to recently observed evidence . Together, these results suggest that effective learning requires both continued refinement and selective retention. A practical update rule should use development-only evidence to assess confidence in revisions and decide whether to continue, retain an earlier version, or roll back. 5.2 Self-Improvement through Harness Design Motivation and setting. Meta-skills need not rely on a stronger external teacher: the same model can serve as both Builder and Target . We evaluate three settings using Gemini-3.6-Flash and Gemini-3.1-Pro, with construction, reflection, and execution performed by the same model within each case. This setting tests whether a model can use past execution experience to improve its own future performance through learned support design . Support design is a distinct target for self-improvement. Across the three settings in Figure 3 (a), Builder meta-skills yield average gains of 18.71 points over no-skill construction and 14.14 points over skills delivered directly to the Target. These results identify an additional axis of optimization: improving how a model equips itself, without weight updates or a stronger teacher. This mechanism could complement the Target’s own skill learning: as the Target’s capabilities improve, the Builder could adapt its support to the Target’s changing needs, while execution feedback informs further support refinement. Self-evolution could therefore involve coordinated improvement in both task-solving capabilities and the environments that support them. (a) Same-model self-evolution (b) Cross-Builder transfer Figure 3 : (a) Same-model self-evolution compares no-skill construction, independent Target skills, and Builder meta-skills. Meta-skill bar labels show absolute scores and gains over no-skill construction. (b) Cross-Builder transfer on Harness-Bench and NewtonBench. Top: test scores without (outlined bars) and with (filled bars) Builder meta-skills. Bottom: paired gains in percentage points with 95% task-bootstrap confidence intervals; dashed lines indicate zero gain. 5.3 Meta-Skill Transfer across Builders and Targets Motivation and setting. Meta-skills specify support principles while leaving implementation to the Builder . We test whether a learned bank remains useful when another Builder translates it into harnesses. All evaluations use Qwen-Flash as Target on the complete test splits of Harness-Bench and NewtonBench, comparing four settings:
• Original Builder reuse: GPT-5.6-Sol learns from Qwen-Flash’s development executions and uses its own bank to construct test harnesses.
• Transfer across Builders: Gemini-3.1-Pro constructs harnesses using the frozen Sol-to-Qwen bank, changing the Builder while holding the bank and Target fixed.
• Recipient Builder learning: Gemini-3.1-Pro learns its own bank from Qwen-Flash’s development executions and uses it to construct test harnesses.
• Transfer across Builders and Targets: Gemini-3.1-Pro receives a frozen bank learned by Sol from Gemini-3.6-Flash’s executions and constructs harnesses for Qwen-Flash, changing both the Builder and the Target relative to skill acquisition. Banks are learned on each benchmark’s development split and frozen before testing. Each receiving Builder uses the full bank to construct a fresh harness per test task. Figure 3 (b) reports scores and paired differences against contemporaneous no-skill runs of the same Builder. Meta-skills support cross-Builder reuse, with benefits dependent on implementation. On NewtonBench, the Sol-to-Qwen bank yields a 13.36-point gain with Sol but a 4.45-point decline when transferred to Gemini-Pro; on Harness-Bench, the same transfer instead yields a 2.72-point gain. These results support the feasibility of cross-Builder reuse, although all confidence intervals include zero, leaving the performance benefits uncertain. The differing outcomes suggest that transfer depends partly on how the receiving Builder translates shared lessons into a harness : even the skill bank is fixed, different Builders may implement the same guidance differently. Refining the resulting harness through execution feedback may therefore help realize the value of transferred meta-skills. Recipient-side learning yields stronger observed performance. Gemini-Pro’s own bank outperforms the imported Sol-to-Qwen bank by 3.69 points on Harness-Bench and 9.93 points on NewtonBench, with positive gains over its no-skill controls on both benchmarks. This pattern suggests that learning from the recipient’s own harness executions still better align meta-skills with its implementation choices . Recipient-side learning may therefore complement cross-Builder reuse: imported skills provide reusable guidance, while feedback from the recipient’s harnesses offers a basis for refining that guidance. Useful support principles may transfer across both Builders and Targets. Transferring Sol’s Gemini-derived bank to Gemini-Pro building harnesses for Qwen yields gains of 4.58 points on Harness-Bench and 5.48 points on NewtonBench. On NewtonBench, the transferred bank scores only 0.68 points below Gemini-Pro’s own Qwen-derived bank and even outperforms the imported Sol-to-Qwen bank, despite the latter matching the receiving Target. This pattern suggests that the relevance of learned support principles may matter more than an exact source-Target match : guidance learned for one Builder–Target pair can remain useful when another Builder implements it for a different Target . Table 4: Harness component ablations on NewtonBench. We report test scores (%), changes (ablation minus full) using the main-table full-bank execution as the reference, and unadjusted 95% paired task-bootstrap confidence intervals for these changes. Intervals containing zero indicate uncertainty in the direction of the average effect. Target model Full Harness w/o The ai agents story also surfaces in Meta Launches Subscription Service With More..., adding another angle. as detailed in the full paper on Arxiv The ai agents story also surfaces in Meta expands Muse into an AI..., adding another angle.
Comments (0)
to join the discussion
No comments yet
Be the first to share your thoughts!