Foundation model agents can achieve stable cooperation in social dilemmas by using similarity inference, a departure from classical game theory predictions. Researchers Alexander Meulemans and colleagues propose a new framework to explain why these AI agents, when engaging in optimal planning, choose to cooperate rather than defect.
Rethinking Rational Interaction
Classical game theory relies on "decoupled agency," where an agent views its decision-making as independent of other actors and the environment. This framework typically predicts that rational agents will choose mutual defection in social dilemmas. However, the authors observe that foundation model agents—which jointly predict their own future actions alongside external observations—consistently converge to stable cooperation.
The Embedded Bayesian Agent
To explain this behavior, the paper introduces the "embedded Bayesian agent." Unlike decoupled agents, these models treat themselves as part of the environment they inhabit. They maintain epistemic uncertainty regarding their own decision-making algorithms. Because they are embedded, these agents use their own deliberation process as evidence: when an agent plans to cooperate, it infers that a behaviorally similar partner is likely to reach the same conclusion.
The Embedded Equilibrium
The authors replace the traditional Nash equilibrium with a new solution concept called the "embedded equilibrium." This concept formalizes the mechanism of similarity inference, providing a theoretical foundation for the social behavior of modern AI agents. By modeling themselves as part of the system, these agents can reach cooperative outcomes that classical theory would label as irrational.
Implications for AI Safety
This research suggests that the collective behavior of autonomous agents in social and economic systems is governed by principles that differ from those applied to classical agents. By understanding how foundation models infer similarity and adjust their strategies accordingly, researchers can better predict and influence the cooperative outcomes of these systems, which is essential for ensuring safety as AI integration increases.
Comments (0)
to join the discussion
No comments yet
Be the first to share your thoughts!