Back to AI Research

AI Research

A game theory for foundation models shows new paths... | AI Research

Key Takeaways

  • Foundation model agents can achieve stable cooperation in social dilemmas by using similarity inference, a departure from classical game theory predictions.
  • Modern AI agents, however, jointly predict their own future actions alongside external observations.
  • To understand this phenomenon, we introduce the embedded Bayesian agent,' a theoretical model for foundation model agents.
  • By shifting from decoupled to embedded agency, these agents model themselves as part of the universe they inhabit, maintaining epistemic uncertainty about their own decision-making algorithms.
  • We show that by inferring whether others are behaviorally similar, an embedded agent treats its own deliberation during planning as evidence: a decision to cooperate predicts a similar decision by a similar partner.
Paper AbstractExpand

As autonomous agents powered by foundation models are increasingly integrated into social and economic systems, understanding the principles governing their collective behavior is essential for ensuring safety and cooperation. Classical game theory, the dominant framework for modeling rational interaction, is built upon the assumption of `decoupled agency,' where agents treat their own decision-making as independent of the environment and other actors. Modern AI agents, however, jointly predict their own future actions alongside external observations. Here, we report a striking finding: when interacting in stylized social dilemmas, foundation model agents engaging in optimal planning consistently converge to stable cooperation, directly contradicting classical game-theoretic predictions of mutual defection. To understand this phenomenon, we introduce the `embedded Bayesian agent,' a theoretical model for foundation model agents. By shifting from decoupled to embedded agency, these agents model themselves as part of the universe they inhabit, maintaining epistemic uncertainty about their own decision-making algorithms. We show that by inferring whether others are behaviorally similar, an embedded agent treats its own deliberation during planning as evidence: a decision to cooperate predicts a similar decision by a similar partner. We formalize this mechanism of similarity inference through the `embedded equilibrium,' a novel solution concept replacing the Nash equilibrium to provide a foundational game theory for the social behavior of modern AI agents.

Foundation model agents can achieve stable cooperation in social dilemmas by using similarity inference, a departure from classical game theory predictions. Researchers Alexander Meulemans and colleagues propose a new framework to explain why these AI agents, when engaging in optimal planning, choose to cooperate rather than defect.

Rethinking Rational Interaction

Classical game theory relies on "decoupled agency," where an agent views its decision-making as independent of other actors and the environment. This framework typically predicts that rational agents will choose mutual defection in social dilemmas. However, the authors observe that foundation model agents—which jointly predict their own future actions alongside external observations—consistently converge to stable cooperation.

The Embedded Bayesian Agent

To explain this behavior, the paper introduces the "embedded Bayesian agent." Unlike decoupled agents, these models treat themselves as part of the environment they inhabit. They maintain epistemic uncertainty regarding their own decision-making algorithms. Because they are embedded, these agents use their own deliberation process as evidence: when an agent plans to cooperate, it infers that a behaviorally similar partner is likely to reach the same conclusion.

The Embedded Equilibrium

The authors replace the traditional Nash equilibrium with a new solution concept called the "embedded equilibrium." This concept formalizes the mechanism of similarity inference, providing a theoretical foundation for the social behavior of modern AI agents. By modeling themselves as part of the system, these agents can reach cooperative outcomes that classical theory would label as irrational.

Implications for AI Safety

This research suggests that the collective behavior of autonomous agents in social and economic systems is governed by principles that differ from those applied to classical agents. By understanding how foundation models infer similarity and adjust their strategies accordingly, researchers can better predict and influence the cooperative outcomes of these systems, which is essential for ensuring safety as AI integration increases.

Comments (0)

No comments yet

Be the first to share your thoughts!