Back to AI Research

AI Research

Do LLM Agents Execute the Plans They Declare? From... | AI Research

Key Takeaways

  • What the paper is about Large language models (LLMs) enable agents to solve long-horizon tasks by generating a plan and then executing it in an environment.
  • Large language models (LLMs) enable agents to solve long-horizon tasks by generating a plan and then executing it in an environment.
  • However, successful planning requires two distinct capabilities: selecting an appropriate plan for the task and executing it faithfully.
  • Existing planner--executor systems can fail at either stage, while final task success alone cannot distinguish selection from execution failures.
  • Across four benchmarks and three LLMs, we find three consistent patterns.
Paper AbstractExpand

Large language models (LLMs) enable agents to solve long-horizon tasks by generating a plan and then executing it in an environment. However, successful planning requires two distinct capabilities: selecting an appropriate plan for the task and executing it faithfully. Existing planner--executor systems can fail at either stage, while final task success alone cannot distinguish selection from execution failures. We therefore study the Plan Declaration--Execution Gap and introduce Planning-as-Routing, where an LLM declares one of four planning modes: Predefined, Sequential, Hierarchical, or Search, and a deterministic router dispatches the task to the corresponding pattern-specific executor. Across four benchmarks and three LLMs, we find three consistent patterns. First, generic Plan+ReAct often fails to preserve declared planning structure, especially for longer plans: across three benchmarks, only (22)--(45%) of trajectories preserve it, whereas pattern-specific executors enforce the intended structure. Second, planning-mode effectiveness varies across environments and models: Search performs best on ALFWorld, Hierarchical on SWE-bench, and the strongest pattern can vary across models within the same benchmark. Third, the largest gains come from execution: pattern-specific executors improve task success from (0.48) to (0.92) on ALFWorld and from (0.36) to (0.44) on SWE-bench Verified over Plan+ReAct. Current LLMs, however, do not reliably select the strongest mode for each task, although few-shot examples improve selection in some benchmark--model combinations. Overall, reliable agent planning requires both effective mode selection and faithful execution: routing substantially closes the execution gap, while task-specific mode selection remains open.

What the paper is about

Large language models (LLMs) enable agents to solve long-horizon tasks by generating a plan and then executing it in an environment. However, successful planning requires two distinct capabilities: selecting an appropriate plan for the task and executing it faithfully. Existing planner--executor systems can fail at either stage, while final task success alone cannot distinguish selection from execution failures. We therefore study the Plan Declaration--Execution Gap and introduce Planning-as-Routing, where an LLM declares one of four planning modes: Predefined, Sequential, Hierarchical, or Search, and a deterministic router dispatches the task to the corresponding pattern-specific executor. Across four benchmarks and three LLMs, we find three consistent patterns. First, generic Plan+ReAct often fails to preserve declared planning structure, especially for longer plans: across three benchmarks, only (22)--(45%) of trajectories preserve it, whereas pattern-specific executors enforce the intended structure. Second, planning-mode effectiveness varies across environments and models: Search performs best on ALFWorld, Hierarchical on SWE-bench, and the strongest pattern can vary across models within the same benchmark. Third, the largest gains come from execution: pattern-specific executors improve task success from (0.48) to (0.92) on ALFWorld and from (0.36) to (0.44) on SWE-bench Verified over Plan+ReAct. Current LLMs, however, do not reliably select the strongest mode for each task, although few-shot examples improve selection in some benchmark--model combinations. Overall, reliable agent planning requires both effective mode selection and faithful execution: routing substantially closes the execution gap, while task-specific mode selection remains open. The ai agents story also surfaces in Google’s Gemini AI Accessed Three Outside..., adding another angle.

What it covers

Do LLM Agents Execute the Plans They Declare? From Planning-Mode Declaration to Pattern-Specific Execution Subba Reddy Oota Affiliation: ADIA Lab, Abu Dhabi, United Arab Emirates Email: [email protected] Francisco Herrera Affiliation: ADIA Lab, Abu Dhabi, United Arab Emirates Affiliation: University of Granada, Granada, Spain Email: [email protected] Jordi Cabot Sagrera Affiliation: ADIA Lab, Abu Dhabi, United Arab Emirates Affiliation: Luxembourg Institute of Science and Technology, Luxembourg Email: [email protected] Marcos López de Prado Affiliation: ADIA Lab, Abu Dhabi, United Arab Emirates Affiliation: Cornell University, Ithaca, USA Affiliation: Lawrence Berkeley National Laboratory, Berkeley, CA Email: [email protected] Shadab Khan Affiliation: ADIA Lab, Abu Dhabi, United Arab Emirates Email: [email protected] Abstract Large language models (LLMs) enable agents to solve long-horizon tasks by generating a plan and then executing it in an environment. However, successful planning requires two distinct capabilities: selecting an appropriate plan for the task and executing it faithfully. Existing planner–executor systems can fail at either stage: the agent may deviate from a structured plan during execution, or it may faithfully execute a plan that is poorly matched to the task or environment. Final task success alone cannot distinguish between these two sources of failure. We therefore ask: can LLM agents be trusted to execute the plans they commit to, and do different tasks and environments benefit from different planning modes? To study this, we introduce a diagnostic framework for the Plan Declaration–Execution Gap and propose Planning-as-Routing , in which the LLM declares one of four planning modes: Predefined, Sequential, Hierarchical, or Search . A deterministic router sends the task to the corresponding pattern-specific executor. Across four benchmarks and three LLMs, we find three consistent patterns. First, generic Plan+ReAct often fails to preserve the declared planning structure, especially for longer plans: across three benchmarks, only 22 22 – 45 % 45% of trajectories preserve it, whereas pattern-specific executors enforce the intended structure. Second, planning-mode effectiveness varies across environments and models: Search performs best on ALFWorld, Hierarchical on SWE-bench, and the strongest pattern can vary across models within the same benchmark. Third, the largest gains come from plan execution: pattern-specific executors raise task success from 0.48 0.48 to 0.92 0.92 on ALFWorld and from 0.36 0.36 to 0.44 0.44 on SWE-bench Verified over Plan+ReAct. However, current LLMs do not reliably select the strongest planning mode for a task, while few-shot examples can improve planning-mode selection for some benchmark--model combinations. Overall, reliable agent planning requires both selecting an effective planning mode and executing it with a matching executor: routing substantially closes the execution gap, while selecting the right mode for each task remains open. 1 1 1 Project page: https://subbareddy248.github.io/projects/declaration-gap-site/ 1 Introduction The rapid advancement of large language models (LLMs) has accelerated the development of interactive agents Liu et al. (2025) ; Wang et al. (2024) , which are designed to solve complex real-world tasks through multi-turn interactions with external environments such as web browsing Zhou et al. (2024b) ; Deng et al. (2023) , computer use Xie et al. (2024) ; Merrill et al. (2026) , and embodied task execution Shridhar et al. (2020) . To solve such tasks, an agent often needs to decompose a high-level goal into structured steps and perform the necessary actions; planning provides this structure by turning an objective into a course of action that guides the agent toward the goal Yao et al. (2023b) . Consequently, planning has become a central component of both agentic frameworks Shen et al. (2023) ; Webb et al. (2025) and world-model-based approaches that simulate or reason about possible future states before acting Qiao et al. (2024) ; Maes et al. (2026) ; Wang et al. (2026) . This structure is what sustains goal-directed behavior in complex, multi-step tasks, where later decisions depend on the outcomes of earlier actions. However, generating a coherent plan is not sufficient for reliable agent behavior. A plan can fail at two levels. At the execution level, the agent may declare one plan but deviate from it during execution. At the selection level, the declared plan itself may be poorly matched to the task: even faithful execution can fail if the chosen plan does not fit the environment. This raises two fundamental questions: when an agent proposes a plan for solving a task, does it execute the plan it commits to, and is the proposed plan appropriate for the task? Existing agent evaluation harnesses offer limited insight into how agents plan and execute their tasks, because they focus on final task success Liu et al. (2024) ; Zhou et al. (2024b) ; Mialon et al. (2024) ; Xie et al. (2024) ; Pan et al. (2024) . Final task success does not reveal whether an agent followed a deliberate plan or reached the goal through an inefficient trajectory, memorization, or chance Liu et al. (2026b) ; Sun et al. (2026) . Recent work has begun to evaluate plan quality and whether agents follow instructed plans Sun et al. (2026) ; Liu et al. (2026b) ; Jia et al. (2025) . However, these studies largely focus on diagnosing planning behavior rather than jointly examining whether the selected planning mode is appropriate for the task and preserved during execution. They also treat a plan as a sequence of steps to be generated or followed, rather than as a choice among different planning modes. Planning requirements differ across tasks and environments: some tasks can be solved with a fixed plan, whereas others require adaptation, decomposition, or exploration. In this work, we study whether LLM agents select a planning mode for a task, whether that mode is preserved during execution, and whether matching modes to corresponding executors improves task success. (a) Flat ReAct (b) Planning in the Prompt + ReAct (c) Planning-as-Routing Figure 1: Overview of the three execution conditions. (a) Flat ReAct : a single ReAct-style executor selects actions from environment feedback without an explicit plan. (b) Planner–Executor (Plan+ReAct) : a planner first selects a planning mode and produces a plan, which is then passed to a generic ReAct executor; however, the declared plan is not structurally enforced during execution. (c) Planning-as-Routing : the LLM first declares a planning mode, and a deterministic router dispatches the task to a matching pattern-specific executor: predefined, sequential, hierarchical, or search. A verifier checks whether execution preserves the declared plan structure. A useful perspective on this problem comes from human planning: humans do not rely on a single fixed strategy for all tasks. A familiar task calls for a routine plan, a structured but unfamiliar one for step-by-step planning, a multi-stage one for decomposition into subgoals, and an uncertain one for weighing alternative routes before acting Daw et al. (2005) ; Daw et al. (2011) ; Botvinick (2008) ; Mattar and Daw (2018) ; Mattar and Lengyel (2022) . This flexibility suggests that planning in LLM agents should not be treated as a uniform behavior. We call a mismatch between the declared plan and its execution the Plan Declaration–Execution Gap . This distinction helps distinguish failures of planning-mode selection from failures of plan execution. This motivates the following research questions:

• RQ1: As agents take on increasingly long-horizon and multi-step tasks, can we rely on them to execute the planning approach they commit to, or do they fall back to reacting one step at a time?

• RQ2: Do different environments and models benefit from different planning approaches, or is a single planning approach sufficient across web, coding, and embodied tasks?

• RQ3: When an agent declares a planning approach, does executing it through the corresponding planning pattern improve task success compared with passing the same plan to a generic ReAct executor?

• RQ4: When multiple planning approaches are available, can an LLM choose the one that is most appropriate for the task it is trying to solve? Together, these questions clarify whether agent failures arise from choosing the wrong planning mode or from failing to execute the chosen mode faithfully. To address these questions, we systematically study how LLM agents select and execute different planning approaches, and whether matching a declared approach to its corresponding execution pattern improves task performance. For the purpose of this work, we evaluate agents built on multiple LLM families (Qwen3.6 Qwen Team (2026) , DeepSeek-V4 DeepSeek-AI (2026) , Gemma-4-26B Abd et al. (2026) ) across four benchmarks spanning diverse and complex environments: web browsing (WebArena Zhou et al. (2024b) , Mind2Web Deng et al. (2023) ), software engineering (SWE-Bench Verified Jimenez et al. (2024) ), and embodied navigation (ALFWorld Shridhar et al. (2020) ). We compare three conditions that differ by one factor at a time: (i) a no-planning Flat ReAct baseline, (ii) a plan-prompted Flat ReAct baseline (Plan+ReAct), and (iii) our Planning-as-Routing framework with pattern-specific executors. Figure 1 illustrates these three conditions. Our contributions are threefold: (1) We introduce a diagnostic framework for measuring the Plan Declaration–Execution Gap in LLM agents, and show that a ReAct agent told to plan a certain way often does not follow the declared plan, instead reacting locally to intermediate observations, a divergence invisible to final-success metrics. Our analysis goes beyond task success by measuring plan quality, plan adherence, and structural faithfulness, while separately examining whether the declared planning approach is effective for the task. (2) We introduce Planning-as-Routing , an architecture in which the LLM declares a planning mode and a deterministic router sends the task to the corresponding pattern-specific executor: {predefined, sequential, hierarchical, or search}. This design makes the declared planning structure explicit in execution and allows its behavior to be verified from the resulting trajectory. (3) Across web navigation, software engineering, and embodied tasks, and across multiple LLM families, we show that planning pattern effectiveness varies across environments and models, while planning-mode selection remains a key bottleneck: current LLM declarations trail the best fixed planning pattern on every benchmark–model pair. Few-shot demonstrations improve task success through better planning-mode selection, with gains of + 0.004 +0.004 to + 0.16 +0.16 across benchmarks. We will release the code, baseline and planning trajectories across seeds upon publication of this paper. 2 Related Work LLMs as Planning Agents. Prior work has proposed many planning mechanisms for LLM agents, but these methods often instantiate planning behavior through a specific agent architecture. Earlier approaches used ReAct-style reasoning and acting Yao et al. (2023b) , search over reasoning paths Yao et al. (2023a) , symbolic planning Liu et al. (2023) , planner–executor architectures Erdogan et al. (2025) , multi-agent workflows Shen et al. (2023) ; Hong et al. (2024) , and adaptive replanning Liu et al. (2026a) ; Dong et al. (2026) ; Wu et al. (2026) . These approaches can generate and revise task-specific plans, but the underlying planning mechanism is typically fixed by the agent framework. Our work instead treats the planning mode as a per-task choice and asks whether the declared mode is preserved during execution. Agent Evaluation and Trajectory Diagnosis. Agent benchmarks span web navigation Deng et al. (2023) ; Zhou et al. (2024b) ; Pan et al. (2024) , computer use Xie et al. (2024) ; Merrill et al. (2026) , software engineering Jimenez et al. (2024) , embodied tasks Shridhar et al. (2020) , and general assistants Mialon et al. (2024) , but primarily evaluate final task outcomes. Recent work moves toward process-level evaluation through progress metrics, plan-compliance analysis, trajectory diagnosis, and failure taxonomies Ma et al. (2024) ; Liu et al. (2026b) ; Ou et al. (2025) ; Kong et al. (2025) ; Cemri et al. (2025) . Our work complements these approaches by explicitly modeling the planner–executor handoff, allowing us to distinguish failures of planning-mode selection from failures to preserve the declared mode during execution. Discussion of planning architectures and process-level agent evaluation is provided in Appendix A . 3 Methodology 3.1 Planning Patterns. We consider four planning patterns inspired by common forms of human planning ( Mattar and Daw, 2018 ; Mattar and Lengyel, 2022 ) : predefined , sequential , hierarchical , and search . These patterns capture different ways an agent can organize task execution. A predefined plan follows a fixed plan generated before execution, without replanning. A sequential plan executes one step at a time and updates the remaining plan based on intermediate observations. A hierarchical plan decomposes the task into subgoals and coordinates their execution through an orchestrator–worker structure. A search plan generates multiple candidate plans, executes each candidate independently, and selects the most promising resulting trajectory using a rubric-based judge . The workflow of each planning pattern is illustrated in Appendix B , Figs. 4 and 5 . Representative plan structures and execution traces for all four planning modes are provided in Appendix E.2 . 3.2 Problem Formulation. We study planning under three conditions that share the same backbone LLM and action space, differing only in how planning is produced and executed. Condition 1: Flat ReAct (no explicit plan declaration). The LLM receives only the task description and goal, and acts through a standard Flat ReAct loop, one step at a time. No planning strategy is explicitly declared before execution. This condition serves as our implicit-planning baseline, where any planning behavior must emerge locally through the interaction loop. Condition 2: Plan+ReAct (declared but unenforced). The LLM first selects a planning approach and produces a task-specific plan. The plan is then provided to the same generic ReAct executor used in Condition 1. Because the executor retains a flat step-by-step control loop, the requested planning structure is available as context but is not structurally enforced. This condition tests whether prompting alone is sufficient for the intended planning approach to appear during execution. Condition 3: Planning-as-Routing. The LLM first declares one planning mode P ∈ { predefined , sequential , hierarchical , search } . P\in{\textit{predefined},\textit{sequential},\textit{hierarchical},\textit{search}}. A deterministic router maps the declaration to the corresponding pattern-specific executor. Unlike Condition 2, the selected planning structure therefore determines the execution control flow. Planning-as-Routing components. (1) Plan declaration. Given a task, the backbone LLM selects one planning mode (P). We evaluate this declaration step under both zero-shot and few-shot settings. In the zero-shot setting , the model selects a mode from the task description alone. In the few-shot setting , the declaration prompt additionally includes example tasks paired with planning modes, allowing us to test whether demonstrations improve task-conditioned mode selection. (2) Router. A deterministic mapping dispatches P P to its corresponding executor, with no additional model inference. (3) Pattern-specific executors. Each executor imposes a distinct planning and control-flow pattern, following established agent-planning architectures. The predefined executor generates a complete plan before execution and follows it without replanning, similar to plan-then-solve approaches Wang et al. (2023) . The sequential executor follows a planner–executor–replanner loop, where the agent executes the current step, observes the outcome, and revises the remaining plan when necessary Sun et al. (2023) . The hierarchical executor follows an orchestrator–worker structure, where an orchestrator decomposes the task into sub-goals, delegates them to specialized workers, and aggregates their outputs Zhang et al. (2025) ; Choi et al. (2025) . Finally, the search executor generates multiple candidate plans, executes each independently, and selects the most promising trajectory using a rubric-based judge, following prior search-based agent planning approaches ( Zhou et al., 2023 ) . (4) Execution-structure verifier. A rule-based verifier checks whether a Plan+ReAct trajectory executes the declared plan steps in order. We first clean the declared steps by removing non-actionable text, then match each remaining step to the corresponding trajectory actions using benchmark-specific rules. Structure is maintained only when all scorable steps are matched in the declared order, while allowing extra actions between them. Routed runs are not scored this way because their executor dispatch records establish order fidelity by construction. Implementation details and representative matching examples are provided in Appendix E.1 . 3.3 Planning Process Metrics Following prior work Jia et al. (2025) , we evaluate plan adherence and plan quality , together with plan-order faithfulness . Plan adherence measures whether the executed actions complete the declared plan steps. Plan quality measures whether the generated plan is appropriate for the task goal and environment. Plan-order faithfulness measures whether the plan steps or subgoals are executed in their declared order. Plan quality and plan adherence are empirical metrics, whereas plan-order faithfulness is a structural verification for predefined, sequential, and hierarchical; it is not applicable to search, where candidates represent competing alternatives rather than an ordered sequence. 3.4 Pattern-Ceiling Analysis Task success alone cannot distinguish whether an executor is incapable of solving a task or whether the declaration module selected an unsuitable planning pattern. To separate execution limitations from planning-mode selection, we use forced dispatch : for each benchmark–model pair, every planning mode is executed on every task, bypassing the declaration module. This yields a task–mode success matrix m i , p m_{i,p} , where m i , p = 1 m_{i,p}=1 if planning mode (p) solves task (i), and (0) otherwise. From this matrix, we compute each fixed mode’s success S ⁡ ( p ) S(p) , the best fixed-mode performance S ⋆ = max p ⁡ S ⁡ ( p ) S^{\star}=\max_{p}S(p) , and a per-task oracle ceiling DSR max = 1 N ​ ∑ i max p ⁡ m i , p \mathrm{DSR}{\max}=\frac{1}{N}\sum{i}\max_{p}m_{i,p} . Their difference H = DSR max − S ⋆ H=\mathrm{DSR}{\max}-S^{\star} measures the available routing headroom. For a declaration policy π \pi , task success is computed as TSR ⁡ ( π ) = 1 N ​ ∑ i m i , π ⁡ ( i ) \mathrm{TSR}(\pi)=\frac{1}{N}\sum{i}m_{i,\pi(i)} , using the same forced-dispatch matrix without re-running executors. Together with the trajectory verifier, this separates execution failure from selection failure : whether the declared mode is preserved during execution versus whether the selected mode is effective for the task. Additional controls and estimation details are provided in Appendix I . 4 Experimental Setup Environments. We evaluate our agent planning strategies across four agent benchmarks spanning web navigation (Mind2Web Deng et al. (2023) , WebArena Zhou et al. (2024b) ), software engineering (SWE-bench Verified Jimenez et al. (2024) ), and embodied tasks (ALFWorld Shridhar et al. (2020) ). We additionally group tasks using each benchmark’s available task categories to examine whether planning-pattern preferences vary with task type; category definitions and counts are provided in Appendix C Table 9 . Language Models. We evaluate three backbone LLMs: Qwen3.6-35B-A3B Qwen Team (2026) , DeepSeek-V4-Flash DeepSeek-AI (2026) , and Gemma-4-26B Abd et al. (2026) . Model details are provided in Appendix C Table 8 . Each model is evaluated under the same task inputs, tool interfaces, action spaces, and execution budgets across all conditions. The exact environment-step and planning-structure budgets are reported in Appendix D , Table 13 . The same backbone model is used for plan declaration and execution unless otherwise specified. Repeated runs. We evaluate all baselines, declaration settings, and planning patterns with three random seeds (7, 13, and 42), reporting mean and standard deviation across seeds using the same tasks and inference configurations. Evaluation Metrics. We report both task- and process-level metrics. Task success follows each benchmark’s standard protocol: task success rate (TSR) for ALFWorld and WebArena, TSR and step success rate (SSR) for Mind2Web, and patch success rate (PSR) for SWE-bench Verified. Process metrics include plan quality, plan adherence, and structural faithfulness (Section 3.3 ). We also measure execution cost through environment interactions, LLM calls, generated tokens (thinking vs. content), and completed trajectories to control for differences in inference and interaction budget. Inference Settings. For each model, decoding parameters, thinking settings, tool-calling policies, and generation budgets are fixed across conditions for each model. Full inference and sampling configurations are provided in Appendix D . Development and tuning protocol. The four planning-pattern executors, their prompts, and the declaration prompt were fixed before the final evaluation runs, and the same implementations and routing rules are applied to every task within a benchmark. The benchmark tasks were used for inference only; no model parameters were trained or fitted on them. For few-shot declaration, demonstration tasks are excluded from scoring. The null baselines and permutation tests were specified after the main results and are reported as post-hoc analyses. Plan verification and quality judging. Structural fidelity is measured with a rule-based verifier that matches declared plan steps to executed actions and requires all scorable steps to be preserved in order; we validate it against two independent human annotators, with full matching and agreement results in Appendices E and E.3 . Plan quality and adherence are scored separately with an LLM-as-judge pipeline using a 0–3 GPA-style rubric ( Jia et al., 2025 ) ; full prompts and judge configurations are provided in Appendix F . Table 1: Plan-structure maintenance under generic Plan+ReAct, overall and by declared plan length. Overall is the percentage of trajectories preserving the declared structure; Avg. steps is the mean number of declared plan units. Pattern-specific executors are omitted because they enforce their execution structure by design. Values are mean percentages ± \pm SD across seeds 7, 13, and 42. ‡ Sparse bins should be interpreted cautiously. Benchmark Model Overall Avg. steps Structure maintained by declared plan length ≤ 𝟑 \mathbf{\leq 3} 4–5 6–7 8–10 > 𝟏𝟎 \mathbf{>10} ALFWorld DeepSeek-V4 27.4 ± 5.1 27.4{\pm}5.1 6.2 ± 0.3 6.2{\pm}0.3 53.3 ± 37.7 53.3{\pm}37.7 36.7 ± 7.0 36.7{\pm}7.0 21.4 ± 3.0 21.4{\pm}3.0 0.0 ± 0.0 ‡ 0.0{\pm}0.0^{\ddagger} 0.0 ± 0.0 ‡ 0.0{\pm}0.0^{\ddagger} Qwen3.6-35B 23.7 ± 2.7 23.7{\pm}2.7 6.1 ± 0.3 6.1{\pm}0.3 66.7 ± 27.2 66.7{\pm}27.2 36.5 ± 1.3 36.5{\pm}1.3 9.5 ± 3.9 9.5{\pm}3.9 1.2 ± 1.7 1.2{\pm}1.7 0.0 ± 0.0 ‡ 0.0{\pm}0.0^{\ddagger} Gemma-4-26B 21.5 ± 1.9 21.5{\pm}1.9 7.6 ± 0.2 7.6{\pm}0.2 91.7 ± 8.3 91.7{\pm}8.3 49.6 ± 2.2 49.6{\pm}2.2 4.8 ± 4.8 4.8{\pm}4.8 0.0 ± 0.0 0.0{\pm}0.0 0.0 ± 0.0 0.0{\pm}0.0 Mind2Web DeepSeek-V4 44.6 ± 0.9 44.6{\pm}0.9 4.8 ± 0.0 4.8{\pm}0.0 73.3 ± 0.5 73.3{\pm}0.5 39.2 ± 1.4 39.2{\pm}1.4 31.2 ± 1.7 31.2{\pm}1.7 10.0 ± 3.4 10.0{\pm}3.4 5.9 ± 3.9 5.9{\pm}3.9 Qwen3.6-35B 36.1 ± 0.9 36.1{\pm}0.9 4.8 ± 0.0 4.8{\pm}0.0 62.4 ± 1.8 62.4{\pm}1.8 31.5 ± 0.3 31.5{\pm}0.3 17.9 ± 1.1 17.9{\pm}1.1 8.6 ± 1.4 8.6{\pm}1.4 4.2 ± 3.1 4.2{\pm}3.1 Gemma-4-26B 28.2 ± 0.2 28.2{\pm}0.2 5.7 ± 0.0 5.7{\pm}0.0 70.1 ± 2.9 70.1{\pm}2.9 24.8 ± 0.8 24.8{\pm}0.8 13.4 ± 0.5 13.4{\pm}0.5 12.7 ± 1.8 12.7{\pm}1.8 11.3 ± 4.9 11.3{\pm}4.9 SWE-bench DeepSeek-V4 25.4 ± 2.1 25.4{\pm}2.1 4.6 ± 0.1 4.6{\pm}0.1 33.4 ± 1.3 33.4{\pm}1.3 25.6 ± 1.2 25.6{\pm}1.2 15.5 ± 2.5 15.5{\pm}2.5 21.2 ± 12.1 21.2{\pm}12.1 0.0 ± 0.0 ‡ 0.0{\pm}0.0^{\ddagger} Qwen3.6-35B 23.5 ± 1.0 23.5{\pm}1.0 4.8 ± 0.1 4.8{\pm}0.1 47.9 ± 3.2 47.9{\pm}3.2 22.3 ± 0.5 22.3{\pm}0.5 13.6 ± 0.1 13.6{\pm}0.1 5.9 ± 0.3 5.9{\pm}0.3 3.6 ± 3.6 3.6{\pm}3.6 Gemma-4-26B 12.1 ± 0.9 12.1{\pm}0.9 5.6 ± 0.3 5.6{\pm}0.3 33.3 ± 1.2 33.3{\pm}1.2 13.3 ± 0.6 13.3{\pm}0.6 4.1 ± 0.2 4.1{\pm}0.2 5.6 ± 0.7 5.6{\pm}0.7 0.0 ± 0.0 0.0{\pm}0.0 WebArena DeepSeek-V4 67.9 ± 4.4 67.9{\pm}4.4 5.1 ± 0.1 5.1{\pm}0.1 84.5 ± 1.2 84.5{\pm}1.2 68.7 ± 7.4 68.7{\pm}7.4 64.1 ± 1.9 64.1{\pm}1.9 45.0 ± 5.0 45.0{\pm}5.0 24.7 ± 17.0 24.7{\pm}17.0 Qwen3.6-35B 42.2 ± 1.2 42.2{\pm}1.2 6.7 ± 0.2 6.7{\pm}0.2 65.9 ± 6.8 65.9{\pm}6.8 59.1 ± 3.6 59.1{\pm}3.6 39.7 ± 0.3 39.7{\pm}0.3 18.4 ± 4.1 18.4{\pm}4.1 5.4 ± 1.5 5.4{\pm}1.5 Gemma-4-26B 61.4 ± 2.5 61.4{\pm}2.5 5.6 ± 0.1 5.6{\pm}0.1 76.1 ± 1.5 76.1{\pm}1.5 65.9 ± 2.6 65.9{\pm}2.6 53.9 ± 18.5 53.9{\pm}18.5 52.3 ± 2.3 52.3{\pm}2.3 32.5 ± 9.8 32.5{\pm}9.8 5 Results [RQ1]: Pattern-specific execution preserves declared planning structure, while generic Plan+ReAct increasingly deviates on longer plans Generic Plan+ReAct does not reliably preserve the committed planning structure. To examine whether agents preserve their declared planning structure, we first measure structure maintenance under Plan+ReAct. Table 1 reports overall plan-structure maintenance under Plan+ReAct and its variation with declared plan length. We make the following observations: (i) Plan+ReAct reveals a substantial Plan Declaration–Execution Gap across benchmarks. Across ALFWorld, Mind2Web, and SWE-bench, only about 22 22 – 45 % 45% of Plan+ReAct trajectories preserve the declared structure. (ii) Plan structure fidelity also decreases with plan length: on ALFWorld, maintenance falls from (36.5)–(49.6%) for 4–5-step plans to (4.8)–(21.4%) for 6–7 steps and approaches zero for longer plans; Mind2Web shows a similar decline, while SWE-bench shows the same overall pattern from short to medium-length plans. WebArena is a short-plan boundary case, with higher maintenance ((65.6)–(69.1%)) and average plans of only (3.3)–(3.8) steps. Hierarchical/Search plans are particularly difficult for generic ReAct execution to preserve. We next analyze structural maintenance by planning mode. The mode-level breakdown in Appendix G , Table 14 , shows that the declaration–execution gap is most pronounced for richer planning structures. Hierarchical plans are difficult for generic Plan+ReAct to preserve, with structure maintenance typically below 25 % 25% across ALFWorld, Mind2Web, and SWE-bench. Search shows a similar pattern where enough declarations are available. Together with the plan-length results, this indicates that generic ReAct is unreliable for multi-level or multi-candidate planning structures. Verifier agreement with human annotations. We further validate the rule-based verifier against two independent human annotators on 100 sampled Plan+ReAct ALFWorld trajectories. Verifier–human agreement ( κ = 0.31 / 0.34 \kappa=0.31/0.34 ) is comparable to human–human agreement ( κ = 0.35 \kappa=0.35 ); full validation results are reported in Appendix E.3 . Qualitative trajectories illustrate the declaration–execution gap. Figure 2 provides representative trajectories that complement the aggregate results. In the maintained Plan+ReAct example, all declared hierarchical units are reached in their original structural order, even though additional environment actions occur between them. In contrast, the non-maintained example skips multiple units of the declared predefined plan while continuing with later parts of the trajectory. This illustrates that a generic ReAct executor can depart from the committed plan structure while continuing to act in the environment. The pattern-specific executor shows a different behavior: the hierarchical plan is explicitly traversed through its declared tree, with all 13 units dispatched according to the prescribed structure. These examples illustrate the distinction captured quantitatively in Tables 1 and 14 (Appendix G ). This suggests that providing a plan as context does not enforce its structure, whereas a pattern-specific executor makes that structure part of the execution control flow. Figure 2: Qualitative examples of plan-structure preservation. (a) A Plan+ReAct trajectory that maintains the declared hierarchical structure: declared units are reached in order despite intervening actions. (b) A Plan+ReAct trajectory that does not maintain the declared predefined structure, skipping two plan units before continuing with later ones. (c) A routed pattern-specific hierarchical executor, where the declared tree directly determines the dispatch sequence. Green denotes maintained/dispatched plan units and red denotes declared units that are not reached. Table 2: Plan adherence on ALFWorld under the fixed execution budget, evaluated on 134 tasks with three random seeds per model. Values are mean ± \pm standard deviation across seeds. Adherence is the fraction of declared executable plan units completed within the allocated budget; Full P The ai agents story also surfaces in Google launches Gemini 3.8 Flash and..., adding another angle. as detailed in the full paper on Arxiv The ai agents story also surfaces in Andrew Ng Launches OpenWorker to Deliver..., adding another angle.

Comments (0)

No comments yet

Be the first to share your thoughts!