Learning What to Activate: Combinatorial Capability Allocation for Long-Horizon Multimodal Agents
Modern AI agents often rely on a library of specialized tools—such as perception, reasoning, or retrieval modules—to complete complex, multi-step tasks. Current systems typically use a fixed set of these tools or follow a rigid, pre-programmed workflow. However, these approaches are often inefficient, as they fail to adapt to the specific needs of each stage of a task. This paper introduces CoCA, a framework that enables agents to dynamically select the most effective subset of tools at every step, balancing task performance against the computational cost of activating those tools. The ai agents story also surfaces in Google AI Introduces EnvHarness for Adaptive..., adding another angle.
The Challenge of Dynamic Tool Selection
In long-horizon tasks, choosing which tools to use is a "combinatorial" problem. The value of a specific tool depends on what other tools are being used at the same time and how previous choices have changed the current state of the task. Because there are so many possible combinations, it is too expensive to manually label the "perfect" set of tools for every possible situation. Furthermore, because the agent’s choices influence its future observations, a poor decision early on can lead to a cascade of errors or unstable, oscillating tool usage throughout the task.
How CoCA Works
CoCA addresses these issues through an on-policy learning framework that avoids the need for exhaustive manual labeling. It uses a three-part approach:
Conditional Utility Modeling: Instead of trying to pick the best set directly, a "teacher" model evaluates the marginal value of each tool based on the tools already selected. This allows the system to learn preferences through simple pairwise comparisons.
Dual-Level Distillation: The system uses a "student" policy that learns to mimic the teacher. To ensure this student performs well in real-world scenarios, it is trained using a mixture of teacher-guided and student-generated sequences. This helps the student handle the distribution shifts that occur when it encounters new states or different tool combinations.
Trajectory-Level Optimization: Finally, the system uses reinforcement learning to refine the policy. This ensures that the agent doesn't just pick tools for immediate gain, but also considers the long-term impact on task success, total activation costs, and the stability of its choices over time. The ai agents story also surfaces in Arm unveils AI-native mobile platform for..., adding another angle.
Efficiency at Inference
A key advantage of CoCA is its performance during deployment. While the training process involves a teacher model and complex optimization, the final "student" policy is lightweight. Once trained, the agent can make its own allocation decisions in real-time without needing to query the teacher or perform additional online updates. This makes it highly suitable for practical, long-horizon multimodal applications where speed and efficiency are critical.
Performance and Results
Experiments conducted on long-horizon multimodal environments demonstrate that CoCA outperforms existing baseline methods. By effectively managing the trade-offs between task success, computational overhead, and allocation stability, the framework proves that dynamic, cost-sensitive tool selection is a superior strategy compared to static configurations or predefined workflows. The ai agents story also surfaces in NVIDIA Agent Toolkit Adds Omniverse Libraries..., adding another angle. as detailed in the full paper on Arxiv
Comments (0)
to join the discussion
No comments yet
Be the first to share your thoughts!