Back to AI Research

AI Research

A Unifying Perspective on Causal World Models: From... | AI Research

Key Takeaways

  • A Unifying Perspective on Causal World Models: From Observations to Representations to Structure proposes a formal framework to bridge the gap between raw se...
  • World Models (WM) are increasingly seen as a foundation for intelligent agents that can predict, plan, and act beyond their training distribution.
  • Finally, we relate CWMs to the literature on identifiability, clarifying when the components of a WM can be recovered from data and up to which equivalence.
  • With this, we ground WMs in representations and structures that support causal reasoning and informed decision-making.
  • Authors Avinash Kori and Fabrizio Russo argue that current world models often conflate simple prediction with the deeper ability to understand entity properties and interactions.
Paper AbstractExpand

World Models (WM) are increasingly seen as a foundation for intelligent agents that can predict, plan, and act beyond their training distribution. In this paper, we study WMs from a causal perspective across multiple levels of abstraction, ranging from perceptual observations to building a conceptual representation of the structure governing the environment dynamics. We argue that useful WMs must go beyond generative capabilities alone: they should also capture entity properties, entity-to-entity interactions, and entity-to-environment interactions that determine and explain the dynamics of a system. We provide a formal definition of Causal WMs (CWMs) grounded in the tasks they are intended to support, connecting world modelling with existing work in causal representation learning, object-centric learning, causal discovery, structural causal models, and model-based decision-making. Finally, we relate CWMs to the literature on identifiability, clarifying when the components of a WM can be recovered from data and up to which equivalence. With this, we ground WMs in representations and structures that support causal reasoning and informed decision-making.

A Unifying Perspective on Causal World Models: From Observations to Representations to Structure proposes a formal framework to bridge the gap between raw sensory data and the structured causal reasoning required for intelligent agents. Authors Avinash Kori and Fabrizio Russo argue that current world models often conflate simple prediction with the deeper ability to understand entity properties and interactions. By defining Causal World Models (CWMs) as a sequence of commitments—moving from perception to representation, then to causal structure and intervention—the paper provides a roadmap for building agents capable of planning and acting safely outside their training distributions.

Defining Causal World Models

The authors define a CWM as a system that integrates observations, actions, latent representations, and utility. Rather than treating a world model as a single "black box" predictor, the framework decomposes it into specific components: an inference model that maps raw observations to entity-centric variables, a transition model that predicts how those variables change based on actions, and a prediction model that translates those states back into observations. This structure allows agents to perform goal-oriented reasoning, such as evaluating whether a specific action will lead to a desired outcome in a tabletop environment.

From Observations to Relational States

A core contribution of the paper is the state-abstraction pipeline. The authors describe a process where raw observations are transformed into "relational variables." These variables are organized into a structured state where diagonal blocks represent individual entity attributes—such as position or velocity—and off-diagonal blocks encode interactions between entities, such as one object leaning on another. This approach ensures that the model captures the underlying dynamics of the environment rather than just pixel-level patterns, enabling the agent to reason about causality and intervention.

Identifiability and Equivalence

The paper addresses the challenge of identifiability, or determining when the components of a world model can be recovered from data. The authors note that exact recovery is often impossible or unnecessary; instead, they propose "identifiability up to equivalence." This means that two models are considered equivalent if they differ only by transformations that preserve the semantics of the world, such as permuting entity labels or applying invertible affine transformations to latent states. By formalizing these equivalences, the authors provide a way to verify if a model’s internal representation is sufficient for downstream tasks like causal reasoning and decision-making.

Franklin Analysis

The framework provided by Kori and Russo is a conceptual synthesis rather than an empirical study. The evidence supports the conclusion that world models require a modular, component-wise approach to achieve causal reasoning. By grounding the definition of a CWM in existing literature on causal discovery and representation learning, the authors clarify that the "world model" label is currently used for a wide range of distinct technologies. The primary limitation of this approach is that it relies on the causal sufficiency assumption—the requirement that all relevant variables driving the environment's dynamics are either observed or can be inferred—which may be difficult to satisfy in complex, real-world scenarios with unobserved confounders.

Comments (0)

No comments yet

Be the first to share your thoughts!