Back to AI Research

AI Research

Qwen-Planner-Agent: A Closed-Loop AI-for-AI Framewo... | AI Research

Key Takeaways

  • What the paper is about The rapid progression of large language models is extending AI from passive content generation into the active workflows of engineeri...
  • The rapid progression of large language models is extending AI from passive content generation into the active workflows of engineering and scientific discovery.
  • This shift raises a compelling question: can AI be both the object of development and an active participant in building next-generation AI systems?
  • We explore this question by building Qwen-Planner-Agent within a closed-loop AI-for-AI framework for scalable development and iterative improvement.
  • Mobile planning offers a demanding test of this approach: complex, long-horizon tasks challenge agent reliability, while costly real-device interaction limits development scalability.
Paper AbstractExpand

The rapid progression of large language models is extending AI from passive content generation into the active workflows of engineering and scientific discovery. This shift raises a compelling question: can AI be both the object of development and an active participant in building next-generation AI systems? We explore this question by building Qwen-Planner-Agent within a closed-loop AI-for-AI framework for scalable development and iterative improvement. Mobile planning offers a demanding test of this approach: complex, long-horizon tasks challenge agent reliability, while costly real-device interaction limits development scalability. The framework connects data production, model training, and deployment through a shared action-feedback-verification contract. (i) AI for Data builds a human-gated agentic data flywheel in which specialized agents construct tasks, collect interaction trajectories, curate and balance training data, and use training feedback to guide subsequent data generation. (ii) AI for Training combines a supervised planning cold start with hybrid-environment online agentic reinforcement learning, where we introduce Competence-Aware Reward-and-Advantage Engineering (CARE) to reduce reasoning and tool-use costs while preserving task performance. (iii) AI drives model--harness co-evolution through an execution-evidence-driven loop that orchestrates memory, skills, and tools at runtime and feeds structured action feedback and preserved failure traces back into coordinated model and harness adaptation. Qwen-Planner-Agent achieves the best overall performance among all evaluated models and systems on MobilePA-Bench, improving over its base model across tool use, memory, skills, and sub-agent coordination. Further evaluations of our model show improvements across non-mobile agentic benchmarks while largely preserving general capabilities.

What the paper is about

The rapid progression of large language models is extending AI from passive content generation into the active workflows of engineering and scientific discovery. This shift raises a compelling question: can AI be both the object of development and an active participant in building next-generation AI systems? We explore this question by building Qwen-Planner-Agent within a closed-loop AI-for-AI framework for scalable development and iterative improvement. Mobile planning offers a demanding test of this approach: complex, long-horizon tasks challenge agent reliability, while costly real-device interaction limits development scalability. The framework connects data production, model training, and deployment through a shared action-feedback-verification contract. (i) AI for Data builds a human-gated agentic data flywheel in which specialized agents construct tasks, collect interaction trajectories, curate and balance training data, and use training feedback to guide subsequent data generation. (ii) AI for Training combines a supervised planning cold start with hybrid-environment online agentic reinforcement learning, where we introduce Competence-Aware Reward-and-Advantage Engineering (CARE) to reduce reasoning and tool-use costs while preserving task performance. (iii) AI drives model--harness co-evolution through an execution-evidence-driven loop that orchestrates memory, skills, and tools at runtime and feeds structured action feedback and preserved failure traces back into coordinated model and harness adaptation. Qwen-Planner-Agent achieves the best overall performance among all evaluated models and systems on MobilePA-Bench, improving over its base model across tool use, memory, skills, and sub-agent coordination. Further evaluations of our model show improvements across non-mobile agentic benchmarks while largely preserving general capabilities.

What it covers

Qwen-Planner-Agent: A Closed-Loop AI-for-AI Framework for Real-World Mobile Planner Agents MAI Team, Alibaba Token Hub, Alibaba Group https://tongyi-mai.github.io/Qwen-Planner-Agent/ Abstract The rapid progression of large language models is extending AI from passive content generation into the active workflows of engineering and scientific discovery. This shift raises a compelling question: can AI be both the object of development and an active participant in building next-generation AI systems? We explore this question by building Qwen-Planner-Agent within a closed-loop AI-for-AI framework for scalable development and iterative improvement. Mobile planning offers a demanding test of this approach: complex, long-horizon tasks challenge agent reliability, while costly real-device interaction limits development scalability. The framework connects data production, model training, and deployment through a shared action-feedback-verification contract. (i) AI for Data builds a human-gated agentic data flywheel in which specialized agents construct tasks, collect interaction trajectories, curate and balance training data, and use training feedback to guide subsequent data generation. (ii) AI for Training combines a supervised planning cold start with hybrid-environment online agentic reinforcement learning, where we introduce Competence-Aware Reward-and-Advantage Engineering (CARE) to reduce reasoning and tool-use costs while preserving task performance. (iii) AI drives model–harness co-evolution through an execution-evidence-driven loop that orchestrates memory, skills, and tools at runtime and feeds structured action feedback and preserved failure traces back into coordinated model and harness adaptation. Qwen-Planner-Agent achieves the best overall performance among all evaluated models and systems on MobilePA-Bench, improving over its base model across tool use, memory, skills, and sub-agent coordination. Further evaluations of our model show improvements across non-mobile agentic benchmarks while largely preserving general capabilities. Figure 1 : Qwen-Planner-Agent achieves frontier-level Overall performance on MobilePA-Bench (left), with strong tool-use, memory, skill-use, and sub-agent capabilities (center) at a low estimated output cost including thinking tokens (right). Contents 1 Introduction 2 Qwen-Planner-Agent 2.1 Overview 2.1.1 AI-for-AI Framework 2.1.2 Task Formulation 2.2 Hybrid Mobile Environment Infrastructure 2.2.1 Complementary Environment Backends 2.2.2 Hybrid Environment Strategy 2.2.3 Training Infrastructure 2.3 AI for Data: An Agent-Driven Data Flywheel 2.3.1 Task Construction 2.3.2 Interaction Trajectory Collection 2.3.3 Training Dataset Composition 2.3.4 Feedback-Driven Data Refinement 2.4 AI for Training: Competence-Adaptive Agentic Optimization 2.4.1 Planning-Oriented Cold Start 2.4.2 Hybrid-Environment Online Agentic Reinforcement Learning 2.4.3 CARE: Competence-Aware Reward-and-Advantage Engineering 2.5 AI for Harness: Toward Model–Harness Co-Evolution 2.5.1 Unified Model–Harness Runtime 2.5.2 Context Management with Skills and Memory 2.5.3 Privacy-Preserving Memory Safety 2.5.4 Model and Harness Refinement from Execution Feedback 3 Experiments 3.1 Mobile Planning Performance and Efficiency 3.2 Ablation Studies 3.2.1 CARE Training Dynamics and Advantage Calibration 3.2.2 Harness Contribution under Fixed Checkpoints 3.2.3 Toward Model-Harness Co-Evolving 3.3 Long-History Memory Evaluation 3.3.1 Evaluation Protocol 3.3.2 Results Across History Lengths 3.4 General Agentic Capability Evaluation 3.5 Qualitative Analysis 4 Related Work 4.1 AI-Assisted AI Development 4.2 Mobile Agents and Interactive Environments 4.3 Agentic Reinforcement Learning 4.4 Agent Harnesses and Model–Harness Co-evolution 5 Conclusion 6 Contributors References A Harness Configuration and Memory Details A.1 Scenario-Adaptation Configuration A.2 Persistent Memory Lifecycle A.3 Memory Evaluation and Reproducibility B Infrastructure Details C Additional Mobile-Planning Case Studies C.1 Procedural Dependencies C.2 Scoped Device and Memory Revisions C.3 Constraint Use Across Dialogue Turns 1 Introduction “Since the design of machines is one of these intellectual activities, an ultraintelligent machine could design even better machines; …” — I. J. Good, Speculations Concerning the First Ultraintelligent Machine , 1965 ( Good, 1965 ) . Recent advances in language-model agents are extending AI beyond generating individual outputs toward executing multi-step workflows guided by interaction and feedback ( Wang et al., 2023a ) . These advances make it increasingly practical to pursue a long-standing ambition: using AI to help build and improve AI. As early as 1950, Turing proposed developing machine intelligence by educating a “child machine” and iteratively refining its design through experimentation ( Turing, 1950 ) . Today, language-model agents invite a further step: involving AI not only as the system being developed, but also as a participant in the development process. This raises a central question: how can these capabilities be organized into a scalable, feedback-driven development lifecycle? We study AI for AI from this perspective, using execution experience to guide coordinated improvements in data production, model training, and deployment. We investigate this question by developing a mobile planner agent for complex, real-world tasks. Completing a high-level user goal requires coordinating actions across applications, maintaining context as states change, recovering from failures, and verifying that the intended outcome has been reached. Developing these capabilities calls for diverse interaction data, effective policy learning, and runtime support suited to changing tasks and resources. Yet real-device interaction is costly and difficult to parallelize, constraining the scale of development and evaluation ( Bai et al., 2024 ; Tang et al., 2026 ) . Task-specific verification provides a complementary opportunity: execution traces and observable outcomes can supply evidence for assessing progress and guiding targeted improvements. Mobile planning therefore provides a concrete setting for studying scalable, AI-assisted agent development. In this report, we present an AI-for-AI framework for developing Qwen-Planner-Agent , a complete agent system that couples a trained Planner Model with a unified Harness. Within this framework, AI interprets execution feedback to identify capability gaps and guide coordinated updates to training data, learning strategies, and runtime support. The AI for Data stage combines AI-assisted task construction and failure diagnosis with automated trajectory collection and curation. Training and development-set feedback guides task generation and sampling adjustments. The AI for Training stage combines a planning-oriented cold start with hybrid-environment online agentic RL. Competence-Aware Reward-and-Advantage Engineering (CARE) adapts rewards and calibrates advantages, with a bounded LLM-based controller configuring predefined reward schedules from training and validation feedback. The AI for Harness stage integrates the Planner Model with a Harness that supplies tool-conditioned Skills, persistent Memory, and execution feedback. AI assists memory consolidation, while failure diagnosis guides data updates and LLM-based Harness revisions. The development process adapts as the agent’s capabilities evolve. Diagnosed failures guide new tasks and data sampling, while updated model behavior informs Harness revisions. In turn, revised Harness instructions shape the context and interaction experience used in subsequent model learning. This reciprocal adaptation links model improvement with runtime refinement across development rounds. Model parameters remain fixed during serving; offline updates are validated and versioned, with human review retained for ambiguous and release-critical decisions. On MobilePA-Bench, Qwen-Planner-Agent 27B achieves the highest Overall score among the evaluated models and agent systems. Its estimated per-task output cost, including thinking tokens, is lower than that of the commercial LLMs in our cost comparison. Beyond mobile planning, our Planner Model, Qwen-Planner-Model, demonstrates broad planning and tool-use capabilities across general agentic benchmarks. Ablation studies further support the effectiveness of model–Harness co-evolution in both mobile planning and general agentic settings. In summary, our contributions are threefold:

• A closed-loop AI-for-AI framework. We present a practical exploration of AI-for-AI through the development of mobile planner agents. AI turns execution feedback into targeted improvements in training data, learning strategies, and runtime support. By adapting these development decisions to evolving agent capabilities and leveraging scalable hybrid environments, our framework provides a closed-loop approach to building and iteratively improving real-world agents.

• Qwen-Planner-Agent. We develop a unified Model–Harness agent system that couples generalizable planning with adaptive runtime support. The Planner Model learns task decomposition, grounded tool use, and failure recovery, while the Harness assembles tool-conditioned Skills, persistent Memory, and execution feedback to accommodate changing resources and user context. AI consolidates execution experience and diagnoses failures to guide model training and Harness refinement. These reviewed updates provide a pathway toward model–Harness co-evolution.

• Performance, generalization, and efficiency. Qwen-Planner-Agent achieves the highest Overall score among the evaluated frontier models and agent systems on MobilePA-Bench, demonstrating strong mobile planning and task-completion capabilities at a lower estimated per-task output cost than the commercial LLMs included in our cost comparison. The Planner Model also demonstrates broad competence across general agentic benchmarks, showing that its planning and tool-use capabilities extend beyond mobile environments. 2 Qwen-Planner-Agent 2.1 Overview 2.1.1 AI-for-AI Framework Figure 2 summarizes the AI-for-AI lifecycle for developing Qwen-Planner-Agent, which comprises a Planner and a Harness. AI for Data (Section 2.3 ) combines AI-assisted task construction with automated interaction collection and curation to supply training tasks and trajectories. AI for Training (Section 2.4 ) learns the planner from these assets through a planning-oriented cold start and competence-adaptive online agentic RL. AI for Harness (Section 2.5 ) equips the planner with a Harness for Skills, Memory, and execution feedback, with AI assisting memory consolidation and Harness refinement. Within the AI-for-AI lifecycle, intermediate evaluation uses a held-out development set rather than the final benchmark test set. Development tasks and their trajectories are excluded from direct training, while their evaluation results and diagnosed failure patterns guide subsequent data generation, training adjustments, and Harness refinement. Training and development-set results, together with deployment traces, feed AI-assisted diagnosis that guides targeted tasks, data reweighting, and Harness revisions. Revised Harness instructions shape the context and trajectories used for subsequent model training, while updated model behavior informs further Harness refinement. This feedback connects the three stages across development rounds. Model and Harness updates are reviewed and versioned offline; model parameters remain fixed during serving. Figure 2 : AI-for-AI lifecycle of Qwen-Planner-Agent, comprising three interconnected phases. (i) AI for Data combines AI-assisted task generation with automated trajectory collection and curation, using model feedback to refine subsequent tasks and training-data composition. (ii) AI for Training combines planning-oriented cold-start training with hybrid-environment online reinforcement learning, while CARE adapts reward scheduling and advantage signals to the model’s evolving competence. (iii) AI for Harness equips the planner with a unified Harness that manages Skills, persistent Memory, and execution feedback, while AI-assisted diagnosis guides subsequent model and Harness refinement and feeds new requirements back into data production and training. 2.1.2 Task Formulation Given a user request x x , Qwen-Planner-Agent interacts with a partially observed environment initialized at state s 0 s_{0} . The underlying state s t s_{t} includes the environment and execution-relevant runtime state and is not directly exposed to the policy. Instead, the policy receives observations o t o_{t} through the environment interface. At step t t , the Harness combines the action–observation history with retrieved memory m t m_{t} and loaded skills 𝒦 t \mathcal{K}{t} to construct the model context. The policy selects an action from the currently available structured action set 𝒜 t \mathcal{A}{t} : c t = ℋ η ( x , m t , 𝒦 t , o ≤ t , a < t ) , a t ∼ π θ ( ⋅ ∣ c t , 𝒜 t ) . c_{t}=\mathcal{H}{\eta}(x,m{t},\mathcal{K}{t},o{\leq t},a_{<t}),\quad a_{t}\sim\pi_{\theta}(\cdot\mid c_{t},\mathcal{A}{t}). (1) Here ℋ η ​ ( ⋅ ) \mathcal{H}{\eta}(\cdot) denotes the Harness context-construction function under instruction configuration η \eta , c t c_{t} the resulting model context, and π θ \pi_{\theta} the policy parameterized by θ \theta . The environment executes a t a_{t} , transitions to the next underlying state s t + 1 s_{t+1} , and returns the next observation o t + 1 o_{t+1} : s t + 1 ∼ P trans ( ⋅ ∣ s t , a t ) , o t + 1 ∼ P obs ( ⋅ ∣ s t + 1 , a t ) . s_{t+1}\sim P_{\mathrm{trans}}(\cdot\mid s_{t},a_{t}),\quad o_{t+1}\sim P_{\mathrm{obs}}(\cdot\mid s_{t+1},a_{t}). (2) P trans ​ ( ⋅ ) P_{\mathrm{trans}}(\cdot) models environment dynamics, while P obs ​ ( ⋅ ) P_{\mathrm{obs}}(\cdot) models the structured feedback exposed to the policy, such as tool results, observable state changes, or execution errors. This formulation accommodates both deterministic and stochastic backends. A task-specific verifier evaluates completion using the interaction history and available execution evidence: v t + 1 = 𝒱 ⁡ ( x , ξ t + 1 , τ ≤ t + 1 ) , where ​ τ ≤ t + 1 = ( o ≤ t + 1 , a ≤ t ) . v_{t+1}=\mathcal{V}(x,\xi_{t+1},\tau_{\leq t+1}),\quad\text{where};\tau_{\leq t+1}=(o_{\leq t+1},a_{\leq t}). (3) The verifier 𝒱 ⁡ ( ⋅ ) \mathcal{V}(\cdot) produces the verification outcome v t + 1 v_{t+1} for request x x from the interaction history τ ≤ t + 1 \tau_{\leq t+1} and available execution evidence ξ t + 1 \xi_{t+1} . The evidence ξ t + 1 \xi_{t+1} includes relevant initial conditions and backend state records available to the verifier; this evidence need not be exposed to the policy. If execution continues, the returned observations enter the next model context. Interaction terminates when the verifier confirms task completion, the agent requests clarification or refuses an unsafe request, or the execution budget is exhausted. The action space covers typed tool calls, memory access and updates, skill selection and loading, clarification or refusal, and task-completion declarations. The agent primarily acts through structured tools rather than pixel-coordinate GUI actions, although tools may return visual observations when needed. Backends may differ in their internal state representations and execution mechanisms while exposing compatible task, action, observation, and verification records. 2.2 Hybrid Mobile Environment Infrastructure Mobile-agent development requires scalable interaction as well as feedback that reflects real execution. We combine programmatic sandboxes, LLM-simulated environments, and selected real-device sessions to support the task interactions formulated in Section 2.1.2 . The infrastructure organizes these backends into shared data-collection, evaluation, and online-training workflows. 2.2.1 Complementary Environment Backends The three backends trade off scalability, scenario coverage, and execution fidelity, making each suitable for different task requirements. Programmatic sandbox environments execute typed tool calls through predefined program logic over structured application databases. Deterministic transitions, task-specific resets, and state-based verification support reproducible, high-throughput interaction. Their coverage is limited to implemented tools and state transitions, so new applications or exceptional behavior require additional engineering. They are therefore most suitable for repeatable tasks with explicit state and completion conditions. LLM-simulated environments use language models to generate environment responses for long-tail interactions that are difficult to implement with fixed logic. They broaden scenario coverage without requiring a dedicated programmatic implementation for every case. However, responses may be inconsistent with prior actions or environment state, so trajectories require task-specific validation before use. Real-device environments execute actions in live device sessions, capturing the effects of OS permissions, authenticated services, cross-application dependencies, and runtime changes. They provide direct evidence of behavior that simulation may miss, but incur higher interaction costs, limited parallelism, and difficult resets. They are therefore valuable for tasks whose completion depends on actual device or service behavior. 2.2.2 Hybrid Environment Strategy We match tasks to backends according to their execution requirements. Tasks with well-defined, reproducible state changes primarily use programmatic sandboxes, while LLM-simulated environments supplement long-tail interactions without suitable fixed implementations. Real-device sessions are used selectively for tasks dependent on live device or service behavior. This allocation combines scalable simulated interaction with prioritized device access where simulation lacks sufficient execution fidelity. A common agent-facing interface makes these experiences usable within the same data and training workflows. Tasks expose typed actions, structured responses, and task-specific completion criteria; their interaction records retain execution errors, observable state changes, and verification outcomes. Backend implementations and internal states remain distinct. Compatible records enter shared curation and rollout-processing pipelines, without assuming that simulated and real-device feedback have identical reliability. 2.2.3 Training Infrastructure Our infrastructure separates model-side training from environment-side execution. The training layer coordinates distributed rollout generation, model updates, and training-resource scheduling. A client–server environment-management layer creates, schedules, and cleans up environment instances and real-device sessions, while the corresponding backends implement resets and state handling. This separation allows model computation and environment capacity to be managed independently. During an online rollout, a policy worker interacts with a backend through the environment-management layer and receives execution feedback after each action. Task verifiers assess completion, while environment validation and trajectory checks determine which records are admitted to training. The training layer uses the admitted trajectories for policy updates. Task failure is distinct from an invalid execution record: unsuccessful interactions can still provide learning and diagnostic evidence. Detailed data-curation rules are described in Section 2.3 . The environment-management layer also supports data collection and evaluation. It remains distinct from the deployment-time Harness, which assembles the Planner Model’s context from Skills, Memory, tools, and execution feedback. Further training and environment-management details are provided in Appendix B . 2.3 AI for Data: An Agent-Driven Data Flywheel The data flywheel in Figure 3 turns capability requirements into executable tasks, collects interaction trajectories, and constructs training datasets that evolve with planner performance. AI assists task construction, failure diagnosis, and targeted data refinement, while automated workflows handle rollout collection and data processing. Training and development-set feedback informs which tasks to generate and how to adjust sampling in the next iteration. The resulting assets support both planning-oriented cold-start training and online reinforcement learning. 2.3.1 Task Construction Task-construction agents translate target capabilities into executable task specifications. Initial objectives come from product requirements, available tool inventories, and representative user scenarios; later iterations also incorporate diagnosed capability gaps, as described in Section 2.3.4 . Each specification contains a user goal, available resources, relevant initial conditions, target capabilities, and completion criteria. It defines what must be accomplished without prescribing a single reference trajectory, allowing different valid plans to satisfy the same objective. Construction jointly considers scenario coverage and capability coverage . Scenario coverage spans application domains, tools, user intents, and interaction patterns. Capability coverage targets information acquisition, tool routing, argument grounding, multi-step dependency handling, state tracking, recovery, and verified task completion. This distinction helps introduce new behavioral requirements rather than merely adding more instances of familiar scenarios. Figure 3 : AI-for-Data workflow. AI-assisted task construction supplies executable tasks for automated interaction collection and curation. Verified trajectories and resettable tasks support planning-oriented cold-start training and online RL, respectively. AI-assisted analysis of rollout and development-set feedback guides targeted task generation and sampling adjustments for the next iteration, with human review before each data release. Solid blue arrows denote automated within-iteration flow; dashed orange arrows denote the human-gated transition between iterations. 2.3.2 Interaction Trajectory Collection Tasks are executed in backends matched to their interaction requirements. Mobile tasks follow the hybrid strategy in Section 2.2.2 ; general-agent tasks use executable service environments, while coding tasks use repository and code-execution environments. An automated rollout workflow runs agent policies through multi-step environment interaction and records their trajectories. Each record preserves selected actions, environment feedback, intermediate outcomes, execution errors, retries, recovery attempts, completion evidence, and the final outcome. Successful trajectories provide candidate supervision for planning, tool use, and task closure. Failed and incomplete trajectories are also retained so that diagnosis can examine where execution diverged, whether recovery was attempted, and why the task remained unfinished. Collection therefore preserves the process leading to an outcome, not only its final success label. 2.3.3 Training Dataset Composition Task-completion verification and data-quality assessment are separate decisions. Environment verifiers determine whether completion criteria have been met, while the curation workflow checks whether the corresponding record is suitable for learning. Automated processing normalizes heterogeneous trajectories, removes malformed or duplicate examples, and checks schema consistency and support for reported outcomes in the execution evidence. Retained records are tagged by capability and linked to failure attributions from rollout analysis. Successful completion alone does not establish training suitability, and unsuccessful interactions may still provide useful diagnostic evidence. Low-confidence, conflicting, safety-sensitive, or insufficiently supported cases are routed to human reviewers, who may approve, correct, reject, or quarantine them. Retained records preserve their task source, execution backend, generating policy, capability annotations, outcome, and failure attribution, making later sampling and repair decisions traceable. The curated data are organized into mobile and non-mobile groups. Mobile data form the core of training and cover cross-application planning, tool execution, state tracking, Memory, Skills, sub-agent coordination, recovery, and task completion. Non-mobile data include general-agent trajectories, coding tasks, and reasoning and instruction-following examples. General-agent trajectories exercise structured tool use and multi-turn planning over services such as Model Context Protocol (MCP) servers ( Anthropic, 2024 ) ; coding tasks add repository understanding, iterative editing, and testing. These examples provide complementary supervision for multi-step problem solving while helping preserve general reasoning and instruction following. The pipeline maintains two training assets: verified interaction trajectories for planning-oriented cold-start training and resettable task instances with reliable completion criteria for online reinforcement learning. These support learning from recorded demonstrations and newly generated interactions, respectively. We configure sampling weights across mobile and non-mobile data, capabilities, difficulty levels, and sources rather than sampling in proportion to raw corpus size. These weights are revised using the feedback described below, keeping mobile planning as the primary objective while retaining complementary non-mobile supervision. 2.3.4 Feedback-Driven Data Refinement As the planner improves, the value of individual tasks changes: mastered tasks may become redundant, whereas unstable behaviors and uncovered capabilities require additional training. AI-assisted analysis of task-level rollouts and capability-level development-set results informs three types of updates: reducing redundant examples, increasing coverage of unstable behaviors, and constructing tasks for missing capabilities. At the task level, rollout-analysis agents examine completion outcomes and failure patterns. Reliably mastered tasks are down-sampled while retaining a small preservation set, and inconsistently completed tasks receive greater weight as stabilization data. Failures involving tool routing, argument grounding, state tracking, recovery, or premature termination are grouped into targeted repair sets. Ambiguous or weakly supported cases return to curation rather than entering the training pool as ordinary examples. At the capability level, analysis agents combine held-out development-set results for Tool Use, Memory, Skills, and Sub-agent with trace-derived failure distributions to distinguish recurring gaps from isolated errors. If the existing pool contains suitable examples, their sampling proportions are adjusted. If coverage is insufficient, task-construction agents generate new tasks with targeted capability requirements, difficulty levels, tool combinations, or interaction patterns. Feedback thus determines both which data to select and which tasks to construct next. Development-set failure patterns guide the construction of new training tasks; the development tasks and their trajectories are not added to the training pool. Proposed task additions and removals, sampling adjustments, repair and quarantine decisions, and quality indicators are recorded together for review. Human reviewers may approve, revise, or reject these changes; only an approved dataset and sampling configuration are frozen for the next training round. Each release preserves a versioned history linking diagnosed capability gaps to the data updates made in response. The feedback loop thus changes both the content and composition of subsequent training data, rather than simply expanding the corpus. 2.4 AI for Training: Competence-Adaptive Agentic Optimization The training pipeline begins with a planning-oriented cold start that converts curated data from the AI-for-Data pipeline into an initial policy following the shared action–feedback–verification contract. Starting from this policy, we perform hybrid-environment online agentic reinforcement learning across programmatic sandboxes, LLM-simulated environments, and selected real-device sessions. Within the RL stage, we propose Competence-Aware Reward-and-Advantage Engineering (CARE), which adapts reward composition and advantage scaling to measured group competence to improve reasoning and tool-use efficiency while preserving task quality. Figure 4 summarizes this process. Figure 4 : Training framework for the Planner Model in Qwen-Planner-Agent. Supervised fine-tuning with turn-level error detection and loss masking establishes a planning-oriented cold-start policy, followed by online agentic reinforcement learning in hybrid environments combining LLM simulation, programmatic sandboxes, and selected real-device interaction. Within the RL stage, CARE (Competence-Aware Reward-and-Advantage Engineering) selects process guidance, accuracy reinforcement, or efficiency optimization according to each trajectory group’s competence. Quality-preserving advantage calibration limits the amplification of small efficiency differences, promoting efficient reasoning and tool use while retaining verified task success as the primary objective. 2.4.1 Planning-Oriented Cold Start We initialize the policy through supervised fine-tuning on curated data from the AI-for-Data pipeline, learning task decomposition, action sequencing, tool use, state-conditioned replanning, and recovery under the shared action–feedback–verification contract. Training prioritizes verified mobile trajectories and includes non-mobile data for broader agentic capabilities, general reasoning, and instruction following, To see alibaba in practice, No GPU? Generate & Train AI... walks through a concrete example. as detailed in the full paper on Arxiv The alibaba story also surfaces in Qwen Developers Open-Source Local-First Search Layer..., adding another angle.

Comments (0)

No comments yet

Be the first to share your thoughts!