Language-model groups overstate consensus when replaying human deliberation on a reasoning task
This research investigates whether groups of Large Language Model (LLM) agents can accurately simulate human collective decision-making. By replaying human deliberation sessions from a Wason reasoning task, the study evaluates if LLM groups reach consensus in the same way humans do. The findings suggest that LLM groups consistently overstate the level of agreement compared to human groups, raising questions about using simulated agents to predict real-world deliberative outcomes.
Replaying human deliberation
To test how well LLMs mirror human behavior, the author replayed 100 human groups using matched LLM agent groups. Each agent was "belief-anchored," meaning it was seeded with the pre-discussion answer provided by a specific human participant. Both the human groups and the LLM groups were evaluated using the same scoring criteria to ensure a fair comparison. The study specifically looked at how these groups reached a final decision, accounting for different ways of defining participation and consensus. The ai agents story also surfaces in Google opens early access to AI..., adding another angle.
Why LLM groups reach higher consensus
The study found a significant gap between human and machine behavior. While about one-fifth of human participants never posted during their sessions, LLM agents almost always participated. Even when the researchers adjusted for these differences in participation and submission patterns, the LLM groups remained significantly more consensual than their human counterparts. In some scenarios, the gap in consensus rates reached over 40 percentage points.
The accuracy of simulated agreement
A key takeaway from the research is that simulated consensus does not necessarily track with collective accuracy. When the researchers removed the ability for agents to rely on memorized answers, the LLM groups reached near-unanimous agreement, but they frequently converged on the wrong answer. This indicates that the high level of agreement among LLM agents is not a reliable indicator of correct reasoning or a reflection of human-like collective intelligence. The ai agents story also surfaces in Stanford Researchers Develop TRACE to Fix..., adding another angle.
Implications for AI modeling
The author concludes that belief-anchored LLM groups act as biased estimators of human group outcomes. Because these models tend to overstate consensus and fail to mirror the nuances of human participation, they may not be suitable for predicting how real-world groups will deliberate. These results provide a necessary, scoring-explicit framework for researchers who intend to use simulated agents to study human social and deliberative processes. The ai agents story also surfaces in Andrew Ng Launches OpenWorker to Deliver..., adding another angle. as detailed in the full paper on Arxiv
Comments (0)
to join the discussion
No comments yet
Be the first to share your thoughts!