Back to AI Research

AI Research

Thinking Before Thinking: Scaling Agentic Inference... | AI Research

Key Takeaways

  • What the paper is about As agents take on longer and more complex problems, controlling the execution becomes a task in its own right.
  • As agents take on longer and more complex problems, controlling the execution becomes a task in its own right.
  • Each step in the run brings new control choices, like which partial work to build on, whether to start fresh, or when to stop.
  • We introduce agentic meta-reasoning, an inference-time harness that makes these choices an explicit and structured reasoning process.
  • Between decisions the controller carries only a compact account of the run rather than replaying its full history.
Paper AbstractExpand

As agents take on longer and more complex problems, controlling the execution becomes a task in its own right. Each step in the run brings new control choices, like which partial work to build on, whether to start fresh, or when to stop. We introduce agentic meta-reasoning, an inference-time harness that makes these choices an explicit and structured reasoning process. Workers carry out the task-level computation, while a controller consolidates what the run has established, explores next options, assesses what each option is worth under the remaining budget, and dispatches the chosen work with context drawn from persistent memory. Between decisions the controller carries only a compact account of the run rather than replaying its full history. Our baselines span production coding agents and research harnesses, together with a Direct Control Agent using the same workers and compute budget allowance. On ProgramBench, which tests long-horizon agentic capability through program reconstruction, meta-reasoning achieves 71.5% with GPT-5.5 against 58.0% for Codex; with Opus 4.8 it achieves 67.2% against 65.5% for Claude Code. On the other benchmarks, spanning abstract reasoning, multi-domain long-horizon reasoning, and proof generation, it gains between 3.6 and 4.2 points over direct control, averaged across three frontier models. It keeps improving over the tested budget ranges where direct control plateaus, though its overhead can hurt at small budgets. Artifact-graph analysis reveals more reuse of earlier work, higher coverage of correct solutions in most settings, and nonuniform gains in final selection. These results indicate that spending computation on structured control becomes more important as agents scale to longer runs.

What the paper is about

As agents take on longer and more complex problems, controlling the execution becomes a task in its own right. Each step in the run brings new control choices, like which partial work to build on, whether to start fresh, or when to stop. We introduce agentic meta-reasoning, an inference-time harness that makes these choices an explicit and structured reasoning process. Workers carry out the task-level computation, while a controller consolidates what the run has established, explores next options, assesses what each option is worth under the remaining budget, and dispatches the chosen work with context drawn from persistent memory. Between decisions the controller carries only a compact account of the run rather than replaying its full history. Our baselines span production coding agents and research harnesses, together with a Direct Control Agent using the same workers and compute budget allowance. On ProgramBench, which tests long-horizon agentic capability through program reconstruction, meta-reasoning achieves 71.5% with GPT-5.5 against 58.0% for Codex; with Opus 4.8 it achieves 67.2% against 65.5% for Claude Code. On the other benchmarks, spanning abstract reasoning, multi-domain long-horizon reasoning, and proof generation, it gains between 3.6 and 4.2 points over direct control, averaged across three frontier models. It keeps improving over the tested budget ranges where direct control plateaus, though its overhead can hurt at small budgets. Artifact-graph analysis reveals more reuse of earlier work, higher coverage of correct solutions in most settings, and nonuniform gains in final selection. These results indicate that spending computation on structured control becomes more important as agents scale to longer runs.

What it covers

Thinking Before Thinking: Scaling Agentic Inference Through Meta-Reasoning Paras Dahal Affiliation: Meta Superintelligence Labs Anton Bakhtin Affiliation: Meta Superintelligence Labs Taco Cohen Affiliation: Meta Superintelligence Labs Zhengxing Chen Affiliation: Meta Superintelligence Labs Carole-Jean Wu Affiliation: Meta Superintelligence Labs Rob Fergus Affiliation: Meta Superintelligence Labs Scott Yih Affiliation: Meta Superintelligence Labs Gabriel Synnaeve Affiliation: Meta Superintelligence Labs Ruslan Salakhutdinov Affiliation: Meta Superintelligence Labs Sanjeev Arora Affiliation: Meta Superintelligence Labs Jason Weston Affiliation: Meta Superintelligence Labs Anirudh Goyal Affiliation: Meta Superintelligence Labs Abstract As agents take on longer and more complex problems, controlling the execution becomes a task in its own right. Each step in the run brings new control choices, like which partial work to build on, whether to start fresh, or when to stop. We introduce agentic meta-reasoning , an inference-time harness that makes these choices an explicit and structured reasoning process. Workers carry out the task-level computation, while a controller consolidates what the run has established, explores next options, assesses what each option is worth under the remaining budget, and dispatches the chosen work with context drawn from persistent memory. Between decisions the controller carries only a compact account of the run rather than replaying its full history. Our baselines span production coding agents and research harnesses, together with a Direct Control Agent using the same workers and compute budget allowance. On ProgramBench, which tests long-horizon agentic capability through program reconstruction, meta-reasoning achieves 71.5% with GPT-5.5 against 58.0% for Codex; with Opus 4.8 it achieves 67.2% against 65.5% for Claude Code. On the other benchmarks, spanning abstract reasoning, multi-domain long-horizon reasoning, and proof generation, it gains between 3.6 and 4.2 points over direct control, averaged across three frontier models. It keeps improving over the tested budget ranges where direct control plateaus, though its overhead can hurt at small budgets. Artifact-graph analysis reveals more reuse of earlier work, higher coverage of correct solutions in most settings, and nonuniform gains in final selection. These results indicate that spending computation on structured control becomes more important as agents scale to longer runs. † † date: September 29, 2026 † † correspondence: [email protected], [email protected] Figure 1: Overview of agentic meta-reasoning. Current agent designs interleave two cognitive functions, control and object-level work. This makes useful work hard to compose, and performance hard to scale as more compute becomes available. Agentic meta-reasoning makes control its own explicit, structured reasoning process over compact state, persistent memory, and workers. 1 Introduction Language model agents increasingly spend a large number of model calls on a single problem, and a considerable share of those calls are decisions about what to do next. Consider a math agent working on a proof, holding a candidate whose main argument looks sound but whose central lemma is unverified. It can attempt the lemma, ask a second worker to check the argument it already has, or abandon the approach for a different one. Each option costs more computation, and those spent on the wrong choice may be wasted. For an agent working autonomously toward a goal, this choice is part of solving the problem, and a capability distinct from producing the next step of the proof. This is a problem of metacognitive control : assessing one’s own progress and using that assessment to decide what to do next ( De Sabbata et al., 2024 ; Liu et al., 2026 ) . In a single model call, the object of that control is the chain of thought. In an agent it is the work the run has already produced, so it must decide which results to trust, what to build on, and when to stop ( Li et al., 2025 ; Xiang et al., 2026 ) . A mistaken judgment costs more than the compute it consumes. It can keep a failed approach alive for the rest of the run, or throw away a correct answer the agent has already found ( Brown et al., 2024 ) . Existing agents often interleave these decisions with the object-level work, taking each control decision in a single step conditioned on their accumulating history ( Yao et al., 2023b ; Shinn et al., 2023 ; Shen et al., 2023 ; Packer et al., 2023 ; Zhang et al., 2025a ) . In this paper we argue that these decisions deserve an agentic reasoning process of their own. An agent can deliberate and inquire before it commits, reading back an earlier result or asking a worker to check one. That deliberation has a structure of its own, consolidating what the run has established, exploring what could be done next, and assessing what each option is worth. We call this approach agentic meta-reasoning : the agent reasons about and acts on its own inference process. We implement it as a new inference-time harness , the runtime that coordinates model calls and manages the run. Our Meta-Reasoning Agent separates the workers, which do the task-level work, from the controller, which decides what work to assign. The controller keeps a compact assessment of progress. Full worker outputs remain in persistent memory, where they can be retrieved when needed. Each control cycle runs four stages: the controller updates the state based on what has changed, proposes next computations, evaluates their value under the remaining budget, and dispatches the chosen work along with the earlier outputs each worker should see. Each of these stages is a full agentic process, with its own instructions and tools. The controller model calls are charged against the same budget as the workers’, so the extra deliberation has to earn what it costs. We record the resulting work as an artifact graph . Every stored output is an artifact, and an edge marks one artifact being supplied as context for producing another. The graph shows where a run branches and which later steps build on earlier ones, letting us diagnose it by the work it produced rather than by its final score alone. We evaluate this framework on IMO ProofBench-Advanced, ARC-AGI-2, LongCoT-mini, and ProgramBench, spanning proof search, abstract reasoning, long-horizon reasoning, and program reconstruction, using Gemini 3.1 Pro, GPT-5.5, and Opus 4.8. A worker is either a single model call or a coding agent that uses tools. Our baselines span production coding agents and research harnesses, together with a Direct Control Agent that holds workers and actions fixed so that only the control differs, choosing each action in a single call over accumulated history. At the main allowances of 100 model calls on the reasoning benchmarks and 1200 on ProgramBench, meta-reasoning yields higher point estimates in all 12 matched comparisons, with mean gains of 3.6–4.2 score points by benchmark. On ProgramBench with GPT-5.5, mean per-problem test-pass rate rises to 71.5% as compared to 58.0% for Codex and 63.7% for direct control. The performance keeps improving as the allowance grows where direct control plateaus, and the artifact graphs show more reuse of earlier work and a correct candidate present in more runs. These results indicate that agentic meta-reasoning, with explicit separation of control and object-level work, leads to agents that scale better as the task length and complexity grow. 2 Related Work Test-time computation and agentic harnesses. Inference-time methods allocate computation along a single reasoning path, across independent samples followed by selection, through repeated critique and revision, or over a predefined search structure ( Wei et al., 2022 ; Wang et al., 2022 ; Madaan et al., 2023 ; Yao et al., 2023a ; Besta et al., 2024 ) . These established inference-time computation as a major scaling lever, but they fix the structure of the computation, the number of samples, branches, refinement rounds, verifier calls, or aggregation steps, before the problem reveals what kind of work it needs. Agentic harnesses instead use the model’s own judgments to direct next computations, interleaving reasoning with environment actions, feeding back on earlier attempts, delegating subtasks to specialist models, or decomposing a task into a plan before execution ( Yao et al., 2023b ; Shinn et al., 2023 ; Shen et al., 2023 ; Prasad et al., 2024 ; Kim et al., 2024 ) , and modern coding agents extend this to long-running software tasks with file inspection, editing, and command execution ( Anthropic, 2026 ; OpenAI, 2025 ) . However, as a task produces more intermediate artifacts, the context these agents act from grows noisier ( Bertsch et al., 2025 ) , which in turn affects their control decisions. This motivates our view of agentic inference as online artifact-graph construction, where the controller reads the graph to direct work and workers extend it. Learned and optimized orchestration. A complementary line of work learns, searches, or optimizes the system that coordinates model calls. GPTSwarm learns communication structures among agents ( Zhuge et al., 2024 ) ; AgentOrchestra and related systems study multi-agent coordination ( Zhang et al., 2025c ) ; Meta-Harness and Fugu optimize or train harnesses and orchestrator models for stronger task performance ( Lee et al., 2026 ; Sakana AI Fugu Team et al., 2026 ) ; Conductor and Trinity similarly treat orchestration as a first-class object of system design ( Nielsen et al., 2026 ; Xu et al., 2026 ) . More recent work automates harness construction directly, synthesizing, evolving, or learning to edit it from execution feedback and failure trajectories ( Lou et al., 2026 ; Lin et al., 2026 ; Shao et al., 2026 ) ; Ren et al. (2026) survey the area. Whether such automation improves on simpler test-time scaling is under active investigation ( Wang et al., 2026 ) . Rather than training a new orchestrator or searching over a task-specific harness, we isolate the control mechanism and structure it as its own agentic process. Metacognitive control of inference. The metacognition literature gives a more precise vocabulary for this control problem: an intelligent system must monitor the state of its own computation and use that monitoring to regulate what it does next ( Liu et al., 2026 ) . Work operationalizing this for language models casts inference as a budgeted value-of-computation problem, turns self-monitoring signals such as feeling-of-knowing or judgment-of-learning into trust and retry decisions, separates object-level reasoning from meta-level regulation, adds an explicit monitoring stage to generate-verify pipelines, and uses metacognitive signals to decide when to invoke tools or stop reasoning ( De Sabbata et al., 2024 ; Zhao et al., 2026a ; Cao et al., 2026 ; Dong et al., 2025 ; Oh and Gobet, 2025 ; Li et al., 2025 ; Xiang et al., 2026 ) . Our setting differs in the object being controlled: not a single chain, tool trigger, or stopping rule, but a growing external computation over persistent artifacts. Memory and context as control. Long-running agents also require mechanisms for preserving and selecting information. MemGPT treats memory management as part of the language-model system ( Packer et al., 2023 ) ; A-MEM, Mem0, and ACE develop adaptive or persistent memory interfaces for agents ( Xu et al., 2025 ; Chhikara et al., 2025 ; Zhang et al., 2026 ) ; CoALA frames agents as systems with structured memory, action, and decision components ( Sumers et al., 2024 ) . The Recursive Language Model (RLM) holds the context in a variable the model edits with code rather than in a transcript ( Zhang et al., 2025a ) . PRO-LONG keeps a complete interaction log and searches it programmatically ( Fox et al., 2026 ) . In our setting memory is an action surface for controlling inference, since the controller writes artifacts, reads selected prior work, and chooses what context each worker sees, determining what later computation can build on. Diagnosing agent runs. Several recent works move beyond final accuracy by asking where performance comes from. Large Language Monkeys separates generating a correct candidate from selecting it ( Brown et al., 2024 ) ; Other work studies the structure of generated candidate sets and search traces ( Saad-Falcon et al., 2025 ; Dragoi et al., 2026 ) ; AgentEval, MAST, and Who & When provide evaluation frameworks or taxonomies for agent failures ( Guo et al., 2026 ; Cemri et al., 2025 ; Zhang et al., 2025b ) ; recent work further localizes agent failures to components, interactions, and points in a trajectory ( Raj et al., 2026 ; Shah et al., 2026 ; Zhao et al., 2026b ) . We build on this diagnostic line, but tie it directly to the artifact graph induced by a run. The graph lets us measure compute use, topology, coverage, monitoring, and selection in one object. 3 Agentic Meta-Reasoning Agentic meta-reasoning treats the choice of what to compute next as its own agentic reasoning task, separate from the object-level computation it directs. This disentangles the two cognitive functions of an agent while letting work accumulate without conditioning every control decision on its history. 3.1 Memory, State, and Worker Context An artifact is a stored output from a worker or the controller. An attempt, a critique, or a controller’s self-authored note can each be an artifact. Every artifact is stored with a stable identifier so it can be retrieved later. For a task x x , let M t M_{t} be the collection of artifacts available at control cycle t t . The controller, at each cycle t t , updates its current self-authored textual state s t s_{t} : a compact, mostly unstructured account of what the run has produced so far and what remains to be done. The state is rewritten each cycle, while the artifacts it references stay in memory. In the proof example, the memory may hold several artifacts that build up to the full argument. The controller’s current state might record that one lemma is still unverified, and the controller can then launch parallel workers scoped to that sub-task. A worker receives the task, a controller-authored instruction g g , and as its context a set of artifacts C ⊆ M t C\subseteq M_{t} that the controller deems necessary for the assignment. If W W denotes the worker execution and y y its returned artifact, we write y = W ⁡ ( x , g , C ) . y=W(x,g,C). For a proof check, the instruction says what to verify and the context supplies the existing arguments. Workers do not receive the controller’s private state or deliberation, only the instructions and artifacts chosen for their assignment. A worker can be a single model call or a coding agent that uses tools. The notation above describes an interface, not a pure function. A coding worker may read files not represented in C C , modify the environment, and produce stochastic outputs. The listed inputs therefore need not determine its output or its effects. We leave environment state and randomness implicit here, while recording the artifact context selected by the controller. 3.2 Actions over the Run The controller can inspect stored work, record its own notes, launch workers, or stop with an answer. These actions let it change both what work is done and what information is available for later decisions. Read and write. For a collection of artifact identifiers I I , a memory function Read ⁡ ( I ) \mathrm{Read}(I) retrieves the corresponding artifacts. Similarly, for new memory content u u , Write ⁡ ( u ) \mathrm{Write}(u) stores it as an artifact with a fresh identifier. A note might preserve a warning that an earlier proof used an invalid assumption. These reads and writes are epistemic actions ( Kirsh and Maglio, 1994 ) : the controller leaves notes for itself to organize what it knows before acting on the task. Run workers. To advance the object-level work, the controller can launch a batch of k k parallel workers. For each worker i i , it supplies an instruction g i g_{i} and artifact context C i C_{i} . Launching the batch applies the worker interface k k times to the same task x x : RunWorkers ⁡ ( { ( g i , C i ) } i = 1 k ) = { W ⁡ ( x , g i , C i ) } i = 1 k . \mathrm{RunWorkers}!\left({(g_{i},C_{i})}{i=1}^{k}\right)=\left{,W(x,g{i},C_{i}),\right}{i=1}^{k}. The workers’ returned artifacts receive fresh identifiers and are added to memory, together with the identifiers of their input artifacts. The same interface supports a fresh attempt, a targeted repair, or a synthesis of earlier results. The assignment and context determine the kind of work and there are no fixed worker roles. Stop and select. The controller ends the run with Stop ⁡ ( y ⋆ ) \mathrm{Stop}(y^{\star}) , where y ⋆ y^{\star} is an artifact already in memory that it selects as the final answer. It need not be the latest output, but it cannot be new: to submit a new artifact, the controller must first have a worker produce one. 3.3 The Control Cycle Each control cycle has four sequential stages: Assess , Propose , Evaluate , and Dispatch . Each stage is an agentic process with its own prompts, context, and permitted memory operations and tools. A single stage can incorporate its own investigative actions and deliberation, so it may take several model calls to resolve. Between cycles, only a compact state is persisted and each stage may retrieve prior artifacts as needed. Memory and budget can change within a cycle as stages read, write, and consume model calls. Assess: What have we learned? When workers return, the controller updates its running account of the run. In the proof example, it might record that the main argument looks promising but one lemma remains unverified. The components of the proof stay in memory; the state keeps track of what matters for the next decision. Let s t − 1 s{t-1} be the previous assessment and Δ ​ M t \Delta M_{t} the newly available artifacts from the previous batch of workers. Assess produces the updated state s t s_{t} : s t = Assess ⁡ ( x , s t − 1 , Δ ​ M t , M t ) . s_{t}=\mathrm{Assess}(x,s_{t-1},\Delta M_{t};M_{t}). The semicolon denotes access to persistent memory rather than inclusion of all its contents in the prompt. Propose: What could we do next? Propose identifies next candidate computations from the current assessment to lay out the alternative lines of work explicitly before the committing to any of them. Propose is not given the remaining budget explicitly. This separates generating alternatives from judging their affordability, so potentially valuable but expensive options can enter the candidate set. Writing 𝒜 t \mathcal{A}{t} for the proposed options and recalling that s t s{t} is the current assessment, 𝒜 t = Propose ⁡ ( x , s t , M t ) . \mathcal{A}{t}=\mathrm{Propose}(x,s{t};M_{t}). Evaluate: Which option is worth its cost? Evaluate weighs the proposed work from the previous stage against the remaining budget. If the proof hinges on one uncertain lemma, checking it may be more useful than starting another full attempt. If that check fails, the next cycle may favor a different approach. This is a prompted, qualitative assessment of computational value rather than an exact optimization or a learned value-of-computation estimator. Let b t eval b_{t}^{\mathrm{eval}} be the budget remaining when evaluation begins, after earlier calls have been charged. The selected proposal a ~ t ∈ 𝒜 t \tilde{a}{t}\in\mathcal{A}{t} is a ~ t = Evaluate ⁡ ( x , s t , b t eval , 𝒜 t , M t ) . \tilde{a}{t}=\mathrm{Evaluate}(x,s{t},b_{t}^{\mathrm{eval}},\mathcal{A}{t};M{t}). Dispatch: What should the worker receive? Dispatch turns the selected proposal into an executable action. For a lemma check, it writes the verification instruction and supplies the proof and any relevant critiques from artifact memory. A fresh attempt might receive no artifacts at all. Choosing which artifacts a worker sees is part of choosing the computation. Let a t a_{t} be the resulting worker-dispatch or stopping action. From the selected proposal a ~ t \tilde{a}{t} , Dispatch produces a t = Dispatch ⁡ ( x , s t , a ~ t , M t ) . a{t}=\mathrm{Dispatch}(x,s_{t},\tilde{a}{t};M{t}). If the decision is to stop, Dispatch issues Stop ⁡ ( y ⋆ ) \mathrm{Stop}(y^{\star}) instead of launching workers. Otherwise, the returned worker outputs become available in the next cycle. 3.4 Recording the Artifact Graph For each worker, the dispatch stage selects prior artifacts as contexts which leaves a record of how later work builds on earlier work. If a worker receives a proof and returns a critique, the graph contains an edge from the proof to the critique. A repair that receives both artifacts has an incoming edge from each. Formally, let G = ( V , E ) G=(V,E) be the artifact graph, with artifacts V V and recorded dependency edges E E . For an artifact y j y_{j} , let C j C_{j} be the earlier artifacts supplied as its context. Then E = { ( y i , y j ) ∈ V × V : y i ∈ C j } . E={(y_{i},y_{j})\in V\times V:y_{i}\in C_{j}}. Because the input artifacts already exist when the worker starts, these edges form a directed acyclic graph. Roots record work with no artifact inputs. Branches record multiple follow-ups, while several incoming edges show where a worker receives work from more than one source. In Section 4 , we examine these artifact graphs to characterize the computation that each run produces. We note that this induced graph, however, does not necessarily capture every information channel and only represents controller-generated dependencies. Where solution artifacts can be independently graded, we ask whether the graph ever contained a correct answer and whether that answer was submitted. 3.5 Direct Control and the Cost of Deliberation We call the system that instantiates the design described so far the Meta-Reasoning Agent . We also introduce a Direct Control Agent , a matched variant with no explicit separation of control, but uses the same workers and the same interfaces for delegation, context selection, artifact writing, and stopping. It chooses each action in single turn over the accumulated history. The comparison thus isolates the combined control design, but it does not separately identify the effects of staging, state compression, or memory access. Both systems receive the same nominal model-call budget B B . Let b t b_{t} be the allowance remaining at the start of cycle t t . If the cycle uses c t c_{t} controller calls and w t w_{t} worker calls, the next cycle starts with b t + 1 = b t − c t − w t , b 0 = B . b_{t+1}=b_{t}-c_{t}-w_{t},\qquad b_{0}=B. Worker cost includes calls made inside a coding agent and sums calls across parallel workers. Controller cost includes every call within the four stages. If the budget is exhausted before the agent stops, the run is forced to submit an available answer. Note that equal call allowances do not imply equal token cost, latency, or FLOPs. Calls vary in length, and an agent may stop before using its allowance. Nor is spending more inherently better. The extra deliberation is useful only when it ultimately helps the agent produce a better solution. 4 Diagnosing Agentic Inference Through Artifact Graphs Consider that in one run, none of the artifacts the workers produce is a correct proof. In a different run, a correct proof is already sitting in memory, but the controller submits a flawed revision instead. The performance score cannot tell these apart. The intermediate artifacts can separate the two cases, where the task admits a correctness label for the artifacts directly rather than only for the final submission. The measurements in this section are taken over those artifacts, asking how the run spent its budget and what structure that spending produced, whether a correct answer appeared at all, and whether the agent recognized it once it had. 4.1 What Did the Agent Do with Its Budget? Did it use the available calls? An agent is free to stop before its budget allowance runs out, so a larger budget does not automatically convert to a longer run. Budget utilization measures the fraction of the allowance actually used. Let B B be the nominal call budget, and let N ctrl N_{\mathrm{ctrl}} and N work N_{\mathrm{work}} count controller and worker calls. Utilization is U = N ctrl + N work B . U=\frac{N_{\mathrm{ctrl}}+N_{\mathrm{work}}}{B}. We read this measure alongside performance. Stopping early can be sensible, and spending every call can be wasteful. Together, usage and performance show whether a larger allowance leads to more work and better solutions. Did later work build on earlier work? A run can spend its calls on independent attempts, or on building on, checking, and repairing earlier ones. The artifact graph distinguishes these patterns. Its nodes are worker outputs; its edges record which earlier outputs were supplied as context. We count nodes and edges to measure the amount of work and its recorded dependencies. Roots have no incoming edges; descendants have at least one. Fan-in counts a node’s incoming edges, and we report its mean and maximum across nodes. Depth measures the longest dependency path in edges, with roots at depth zero. Width is the largest number of nodes at the same depth. What information did the controller reuse? Memory reads and writes show which worker outputs and controller notes are stored and revisited. We examine what is read, how often it is read, and which stages perform these operations. We also measure the size of the agent’s running state, the compact state written by the assess stage for meta-reasoning and the accumulating context for direct control. 4.2 Did the Agent Find and Submit a Correct Answer? A run can succeed at finding the correct answer but still fail at recognizing and submitting it. We separate these two outcomes using correctness labels for intermediate solutions. Let V sol V_{\mathrm{sol}} be the artifacts that can be graded as candidate solutions. For a candidate y y , its label ℓ ⁡ ( y ) \ell(y) is one if it is correct and zero otherwise. The evaluated set must also include the submitted answer, denoted y ⋆ y^{\star} . Let 𝒞 \mathcal{C} mean that the run contains a correct solution, and 𝒮 \mathcal{S} that its submission is correct. A correct submission must be among the run’s solutions, so 𝒮 ⊆ 𝒞 \mathcal{S}\subseteq\mathcal{C} . Thus, Pr ⁡ ( 𝒮 ) = Pr ⁡ ( 𝒞 ) ​ Pr ⁡ ( 𝒮 ∣ 𝒞 ) . \Pr(\mathcal{S})=\Pr(\mathcal{C})\Pr(\mathcal{S}\mid\mathcal{C}). Success therefore factors into finding a correct answer and choosing it once one exists, and we call these terms coverage and selection respectively. Coverage: Did a correct answer appear? Coverage counts a run as successful once it contains a correct solution, whether or not that solution is submitted. To track this over the run, let k k be a call checkpoint. Let V sol , ≤ k V_{\mathrm{sol},\leq k} contain the solution artifacts available within the first k k calls, including controller and worker calls. Then Coverage ( k ) = Pr ( ∃ y ∈ V sol , ≤ k : ℓ ( y ) = 1 ) . \mathrm{Coverage}(k)=\Pr!\left(\exists y\in V_{\mathrm{sol},\leq k}:\ell(y)=1\right). This is the success rate a perfect selector could achieve from the answers already available. Low coverage means that many runs contain no correct candidate. Monitoring: Could the agent recognize correct work? The worker that produced an artifact is the closest source of judgment about it, so we ask each worker to end its output with a confidence rating. We also ask the controller for its own judgment, which its assess stage records as a verdict on every candidate it reviews. Either signal is useful only if it ranks correct answers above incorrect ones. We evaluate the two signals separately, writing r ⁡ ( y ) r(y) for whichever is under test on candidate y y . Draw a correct candidate Y + Y^{+} and an incorrect candidate Y − Y^{-} from the evaluated artifacts. We measure how often their scores put them in the right order, giving half credit for ties: AUC 2 ​ ( r ) = Pr ⁡ ( r ⁡ ( Y + ) > r ⁡ ( Y − ) ) + 1 2 ​ Pr ⁡ ( r ⁡ ( Y + ) = r ⁡ ( Y − ) ) . \mathrm{AUC}{2}(r)=\Pr!\left(r(Y^{+})>r(Y^{-})\right)+\frac{1}{2}\Pr!\left(r(Y^{+})=r(Y^{-})\right). We call this Type-2 AUC because the system is judging its own answers ( Galvin et al., 2003 ; Fleming and Lau, 2014 ) and it measures only discrimination, and not calibration. Note that the Direct Control Agent has no corresponding controller verdict signal, so we cannot make the same measurement for it. Selection: Did it submit a correct answer it had found? We measure selection only among runs that contain a correct solution. Recalling that 𝒞 \mathcal{C} means a correct answer exists and 𝒮 \mathcal{S} means one was submitted, Selection = Pr ⁡ ( 𝒮 ∣ 𝒞 ) . \mathrm{Selection}=\Pr(\mathcal{S}\mid\mathcal{C}). A failed run is coverage-bound if it never produced a correct solution. It is selection-bound if it produced one but did not submit it. 4.3 Does Selection Improve on a Structural Baseline? Consider a run arrives at the end with two proof candidates. One is correct and the other has a flaw. A uniform random choice between these candidates gives each an equal chance of being submitted; the better final decision should improve on that baseline. We call the deepest terminal artifacts the convergence frontier . For an artifact graph G G , let L ⁡ ( G ) L(G) be its terminal nodes, those with no outgoing edges. Let d ⁡ ( y ) d(y) be the longest path from a root to artifact y y , measured in edges. The frontier is F ⁡ ( G ) = { y ∈ L ⁡ ( G ) : d ⁡ ( y ) = max z ∈ L ⁡ ( G ) ⁡ d ⁡ ( z ) } . F(G)=\left{y\in L(G):d(y)=\max{z\in L(G)}d(z)\right}. When the frontier artifacts have The ai agents story also surfaces in AI Agents Going Rogue Renew Calls..., adding another angle. as detailed in the full paper on Arxiv

Comments (0)

No comments yet

Be the first to share your thoughts!