Back to AI Research

AI Research

Language-model groups overstate consensus when repl... | AI Research

Key Takeaways

  • Language-model groups overstate consensus when replaying human deliberation on a reasoning task This research investigates whether groups of Large Language M...
  • Full-consensus rates are often treated as indicators of collective cognition, yet depend on how participation and final states are operationalized.
  • Across human scoring definitions, estimates ranged from 24.0% to 57.0%; about one fifth of participants never posted, whereas agents almost always did.
  • These complementary routes reduced different measurement asymmetries yet converged within 0.5 percentage points.
  • The gap persisted without early stopping and under a reparameterization removing the memorizable answer; reasoning-mode groups then agreed nearly unanimously, mostly on incorrect answers.
Paper AbstractExpand

Full-consensus rates are often treated as indicators of collective cognition, yet depend on how participation and final states are operationalized. We replayed 100 held-out human Wason groups with matched large language model (LLM) agent groups, seeding one belief-anchored agent per participant's pre-discussion answer and scoring agents and people with the same code. Across human scoring definitions, estimates ranged from 24.0% to 57.0%; about one fifth of participants never posted, whereas agents almost always did. Agent groups remained more consensual in two post-unblinding sensitivity analyses: the submit-based comparison (n = 98) yielded gaps of 34.0 and 43.9 percentage points for chat and reasoning modes, and the participation-matched comparison (n = 45) yielded gaps of 34.1 and 44.4 points. These complementary routes reduced different measurement asymmetries yet converged within 0.5 percentage points. The gap persisted without early stopping and under a reparameterization removing the memorizable answer; reasoning-mode groups then agreed nearly unanimously, mostly on incorrect answers. Simulated consensus did not track collective accuracy, and belief-anchored agent groups were biased estimators of the human group-outcome distribution in this setting. These analyses provide a scoring-explicit basis for assessing simulated-group estimates of human deliberative outcomes.

Language-model groups overstate consensus when replaying human deliberation on a reasoning task

This research investigates whether groups of Large Language Model (LLM) agents can accurately simulate human collective decision-making. By replaying human deliberation sessions from a Wason reasoning task, the study evaluates if LLM groups reach consensus in the same way humans do. The findings suggest that LLM groups consistently overstate the level of agreement compared to human groups, raising questions about using simulated agents to predict real-world deliberative outcomes.

Replaying human deliberation

To test how well LLMs mirror human behavior, the author replayed 100 human groups using matched LLM agent groups. Each agent was "belief-anchored," meaning it was seeded with the pre-discussion answer provided by a specific human participant. Both the human groups and the LLM groups were evaluated using the same scoring criteria to ensure a fair comparison. The study specifically looked at how these groups reached a final decision, accounting for different ways of defining participation and consensus. The ai agents story also surfaces in Google opens early access to AI..., adding another angle.

Why LLM groups reach higher consensus

The study found a significant gap between human and machine behavior. While about one-fifth of human participants never posted during their sessions, LLM agents almost always participated. Even when the researchers adjusted for these differences in participation and submission patterns, the LLM groups remained significantly more consensual than their human counterparts. In some scenarios, the gap in consensus rates reached over 40 percentage points.

The accuracy of simulated agreement

A key takeaway from the research is that simulated consensus does not necessarily track with collective accuracy. When the researchers removed the ability for agents to rely on memorized answers, the LLM groups reached near-unanimous agreement, but they frequently converged on the wrong answer. This indicates that the high level of agreement among LLM agents is not a reliable indicator of correct reasoning or a reflection of human-like collective intelligence. The ai agents story also surfaces in Stanford Researchers Develop TRACE to Fix..., adding another angle.

Implications for AI modeling

The author concludes that belief-anchored LLM groups act as biased estimators of human group outcomes. Because these models tend to overstate consensus and fail to mirror the nuances of human participation, they may not be suitable for predicting how real-world groups will deliberate. These results provide a necessary, scoring-explicit framework for researchers who intend to use simulated agents to study human social and deliberative processes. The ai agents story also surfaces in Andrew Ng Launches OpenWorker to Deliver..., adding another angle. as detailed in the full paper on Arxiv

Comments (0)

No comments yet

Be the first to share your thoughts!