Back to AI Research

AI Research

Learning What to Activate: Combinatorial Capability... | AI Research

Key Takeaways

  • Learning What to Activate: Combinatorial Capability Allocation for Long-Horizon Multimodal Agents Modern AI agents often rely on a library of specialized too...
  • Long-horizon multimodal agents rely on specialized capabilities for perception, retrieval, reasoning, verification, and execution.
  • Existing designs typically activate a fixed capability set or invoke a predefined workflow, incurring substantial computational overhead while failing to accommodate stage-dependent capability demands.
  • We introduce \textsc{CoCA}, an on-policy learning framework that recovers a deployable capability-subset policy from sparse conditional comparisons.
  • On states visited by the student policy, the stronger teacher compares the marginal net values of candidate capabilities, conditioned on the currently selected subset.
Paper AbstractExpand

Long-horizon multimodal agents rely on specialized capabilities for perception, retrieval, reasoning, verification, and execution. Existing designs typically activate a fixed capability set or invoke a predefined workflow, incurring substantial computational overhead while failing to accommodate stage-dependent capability demands. In this paper, we study the \textit{combinatorial capability allocation} problem for long-horizon multimodal agent systems, where the system selects a cost-sensitive subset of specialized capabilities at each interaction stage, which is nontrivial since capability values depend on the selected subset, while previous allocations alter the states encountered by subsequent decisions. We introduce \textsc{CoCA}, an on-policy learning framework that recovers a deployable capability-subset policy from sparse conditional comparisons. On states visited by the student policy, the stronger teacher compares the marginal net values of candidate capabilities, conditioned on the currently selected subset. Then, we adopt a conditional utility model to transform such comparisons into an autoregressive capability-subset policy, avoiding explicit enumeration. We further introduce dual-level on-policy distillation to address distribution mismatch both across environment states and within the partial subsets encountered during set construction. Finally, trajectory-level reinforcement learning refines the distilled policy toward task success, activation cost, and allocation stability. At inference time, allocation is performed solely by the lightweight student policy without teacher queries or online updates. Experiments on long-horizon multimodal environments and controlled capability-demand shifts demonstrate the superiority of our method over the state-of-the-art baseline methods.

Learning What to Activate: Combinatorial Capability Allocation for Long-Horizon Multimodal Agents
Modern AI agents often rely on a library of specialized tools—such as perception, reasoning, or retrieval modules—to complete complex, multi-step tasks. Current systems typically use a fixed set of these tools or follow a rigid, pre-programmed workflow. However, these approaches are often inefficient, as they fail to adapt to the specific needs of each stage of a task. This paper introduces CoCA, a framework that enables agents to dynamically select the most effective subset of tools at every step, balancing task performance against the computational cost of activating those tools. The ai agents story also surfaces in Google AI Introduces EnvHarness for Adaptive..., adding another angle.

The Challenge of Dynamic Tool Selection

In long-horizon tasks, choosing which tools to use is a "combinatorial" problem. The value of a specific tool depends on what other tools are being used at the same time and how previous choices have changed the current state of the task. Because there are so many possible combinations, it is too expensive to manually label the "perfect" set of tools for every possible situation. Furthermore, because the agent’s choices influence its future observations, a poor decision early on can lead to a cascade of errors or unstable, oscillating tool usage throughout the task.

How CoCA Works

CoCA addresses these issues through an on-policy learning framework that avoids the need for exhaustive manual labeling. It uses a three-part approach:

  • Conditional Utility Modeling: Instead of trying to pick the best set directly, a "teacher" model evaluates the marginal value of each tool based on the tools already selected. This allows the system to learn preferences through simple pairwise comparisons.

  • Dual-Level Distillation: The system uses a "student" policy that learns to mimic the teacher. To ensure this student performs well in real-world scenarios, it is trained using a mixture of teacher-guided and student-generated sequences. This helps the student handle the distribution shifts that occur when it encounters new states or different tool combinations.

  • Trajectory-Level Optimization: Finally, the system uses reinforcement learning to refine the policy. This ensures that the agent doesn't just pick tools for immediate gain, but also considers the long-term impact on task success, total activation costs, and the stability of its choices over time. The ai agents story also surfaces in Arm unveils AI-native mobile platform for..., adding another angle.

Efficiency at Inference

A key advantage of CoCA is its performance during deployment. While the training process involves a teacher model and complex optimization, the final "student" policy is lightweight. Once trained, the agent can make its own allocation decisions in real-time without needing to query the teacher or perform additional online updates. This makes it highly suitable for practical, long-horizon multimodal applications where speed and efficiency are critical.

Performance and Results

Experiments conducted on long-horizon multimodal environments demonstrate that CoCA outperforms existing baseline methods. By effectively managing the trade-offs between task success, computational overhead, and allocation stability, the framework proves that dynamic, cost-sensitive tool selection is a superior strategy compared to static configurations or predefined workflows. The ai agents story also surfaces in NVIDIA Agent Toolkit Adds Omniverse Libraries..., adding another angle. as detailed in the full paper on Arxiv

Comments (0)

No comments yet

Be the first to share your thoughts!