Back to AI Research

AI Research

AD-WM: Action-Discriminative World Models for Count... | AI Research

Key Takeaways

  • AD-WM: Action-Discriminative World Models for Counterfactual Model Predictive Control Latent world models are commonly used in robotics to predict how a scen...
  • Latent world models are typically trained to predict factual transitions, whereas model predictive control (MPC) must compare alternative actions from the same state.
  • A model can therefore achieve low factual prediction error yet poorly distinguish candidate actions.
  • We introduce AD-WM, an action-discriminative joint-embedding world model for counterfactual MPC.
  • AD-WM combines residual latent dynamics with predictor-level action-recovery regularization, using inverse dynamics and a normalized recovery objective motivated by conditional mutual information.
Paper AbstractExpand

Latent world models are typically trained to predict factual transitions, whereas model predictive control (MPC) must compare alternative actions from the same state. A model can therefore achieve low factual prediction error yet poorly distinguish candidate actions. We introduce AD-WM, an action-discriminative joint-embedding world model for counterfactual MPC. AD-WM combines residual latent dynamics with predictor-level action-recovery regularization, using inverse dynamics and a normalized recovery objective motivated by conditional mutual information. Both objectives encourage planning transitions to preserve action information; their auxiliary heads are discarded at test time, leaving MPC unchanged. On OGBench-Cube, AD-WM improves hard-start success from 3.7% to 52.0% over a matched LeWM baseline and improves mean success over the reproduced baseline in four of five simulation environments. Planning diagnostics show that factual prediction error and whole-bank action ranking do not follow the closed-loop success ordering, whereas CEM-aligned elite regret tracks success more closely. With a frozen V-JEPA 2 encoder and matched DROID post-training, AD-WM also improves zero-shot transfer to our Franka setup, increasing basic pick-and-place success from 42.2% to 71.1% without lab-specific adaptation. These results suggest that world models for planning should preserve action-dependent differences needed for counterfactual selection, rather than optimize factual prediction accuracy alone. More videos and code are available at this https URL .

AD-WM: Action-Discriminative World Models for Counterfactual Model Predictive Control
Latent world models are commonly used in robotics to predict how a scene will change when an agent takes an action. However, these models are typically trained to predict what actually happened in a recorded video, rather than what would happen if the agent chose a different action. This creates a problem for planning: a model might be very good at predicting the future based on the action taken, but fail to distinguish between different candidate actions during the planning process. AD-WM (Action-Discriminative World Model) addresses this by explicitly training the model to preserve action-dependent information, ensuring that the planner can effectively compare different choices to reach a goal.

The Problem with Factual Prediction

Standard world models often prioritize minimizing "factual prediction error"—essentially trying to match the predicted next state as closely as possible to the observed one. In many scenes, the background remains mostly the same, so a model can achieve low error by simply predicting that nothing will change. While this is accurate for the specific action taken, it often masks the subtle, action-dependent changes (like moving toward or away from an object) that are critical for making smart decisions. Because the model doesn't learn to distinguish these small differences, it struggles to rank candidate actions correctly during planning. The robotics story also surfaces in NVIDIA Launches Cosmos 3 Edge for..., adding another angle.

How AD-WM Works

AD-WM improves planning by adding "action-recovery" objectives to the training process. Instead of just predicting the next state, the model is also tasked with recovering the action taken from the current and predicted states. It uses two main techniques:

  • Residual Latent Prediction: The model focuses on predicting only the change (the increment) between the current state and the next, rather than the entire state.

  • Action-Recovery Regularization: The model uses inverse dynamics and a normalized recovery objective to ensure that the predicted transition contains enough information to identify which action was performed.
    These auxiliary tasks act as a guide during training, forcing the model to capture the specific features that distinguish one action from another. Importantly, these extra "heads" are only used during training; at test time, they are discarded, meaning the model can be used with standard planning algorithms like Model Predictive Control (MPC) without any changes to the deployment pipeline. The robotics story also surfaces in MIT Researchers Develop Method to Make..., adding another angle.

Key Results and Performance

In testing, AD-WM demonstrated significant improvements over existing baselines. On the OGBench-Cube benchmark, it increased "hard-start" success—a measure of how well the agent performs when starting from difficult, unfamiliar positions—from 3.7% to 52.0% compared to a matched baseline. The researchers also found that traditional metrics, like factual prediction error, were poor predictors of actual success. Instead, they introduced "elite regret" diagnostics, which measure how well the model ranks the best candidate actions; this metric proved to be a much more reliable indicator of whether the robot would successfully complete its task.

Real-World Transfer

Beyond simulation, the researchers tested AD-WM’s ability to transfer to a real-world Franka robot. By using a frozen encoder and post-training the model on robot data, they achieved a significant boost in pick-and-place success, rising from 42.2% to 71.1%. This suggests that by focusing on action-discriminative features rather than just visual accuracy, world models can become much more robust and effective when moving from controlled simulations to physical environments. The robotics story also surfaces in Black Forest Labs Unveils FLUX 3..., adding another angle. as detailed in the full paper on Arxiv

Comments (0)

No comments yet

Be the first to share your thoughts!