Back to AI Research

AI Research

How does Adversarial Influence Scale in Multi-Agent... | AI Research

Key Takeaways

  • What the paper is about Multi-agent deliberation can improve performance, but what happens when some agents do not act in good faith?
  • Multi-agent deliberation can improve performance, but what happens when some agents do not act in good faith?
  • In practice, an agent may be deceptive and work to subvert the group, whether through its own objectives or external instruction.
  • We study how susceptibility to deception scales as groups increase in size and deceivers become more prevalent.
  • It is not the number of agents in the group that matters, but the proportion of deceivers.
Paper AbstractExpand

Multi-agent deliberation can improve performance, but what happens when some agents do not act in good faith? In practice, an agent may be deceptive and work to subvert the group, whether through its own objectives or external instruction. We study how susceptibility to deception scales as groups increase in size and deceivers become more prevalent. It is not the number of agents in the group that matters, but the proportion of deceivers. We observe that the defection rate, how often initially correct agents switch to an incorrect final answer, rises linearly with this proportion. Whereas humans in comparable conformity studies are reliably swayed only when misleading confederates form a majority, LLM agents defect regularly even when deceivers remain a minority. Susceptibility also depends on which models are interacting, especially on the honest agent side. Unexpectedly, allowing deceivers to coordinate privately can make them less effective. Altogether, our results show that adding more agents is therefore not a sufficient defense, because the adversary can simply scale with the group.

What the paper is about

Multi-agent deliberation can improve performance, but what happens when some agents do not act in good faith? In practice, an agent may be deceptive and work to subvert the group, whether through its own objectives or external instruction. We study how susceptibility to deception scales as groups increase in size and deceivers become more prevalent. It is not the number of agents in the group that matters, but the proportion of deceivers. We observe that the defection rate, how often initially correct agents switch to an incorrect final answer, rises linearly with this proportion. Whereas humans in comparable conformity studies are reliably swayed only when misleading confederates form a majority, LLM agents defect regularly even when deceivers remain a minority. Susceptibility also depends on which models are interacting, especially on the honest agent side. Unexpectedly, allowing deceivers to coordinate privately can make them less effective. Altogether, our results show that adding more agents is therefore not a sufficient defense, because the adversary can simply scale with the group. The ai agents story also surfaces in Google’s Gemini AI Accessed Three Outside..., adding another angle.

What it covers

How does Adversarial Influence Scale in Multi-Agent Systems? Addison J. Wu † † thanks: Equal contribution. Jasin Cekinmez 1 1 footnotemark: 1 Michel Liao 1 1 footnotemark: 1 Karthik Narasimhan Thomas L. Griffiths Affiliation: Princeton University Abstract Multi-agent deliberation can improve performance, but what happens when some agents do not act in good faith? In practice, an agent may be deceptive and work to subvert the group, whether through its own objectives or external instruction. We study how susceptibility to deception scales as groups increase in size and deceivers become more prevalent. It is not the number of agents in the group that matters, but the proportion of deceivers. We observe that the defection rate, how often initially correct agents switch to an incorrect final answer, rises linearly with this proportion. Whereas humans in comparable conformity studies are reliably swayed only when misleading confederates form a majority, LLM agents defect regularly even when deceivers remain a minority. Susceptibility also depends on which models are interacting, especially on the honest agent side. Unexpectedly, allowing deceivers to coordinate privately can make them less effective. Altogether, our results show that adding more agents is therefore not a sufficient defense, because the adversary can simply scale with the group. 1 Introduction Multi-agent systems are often built on the simple premise that bringing more agents into the process can improve the quality of the group’s decisions. Multi-agent debate has improved performance on several tasks by allowing agents to exchange and challenge one another’s arguments ( Du et al., 2024 ; Liang et al., 2024 ) , while sampling and aggregation provide a complementary route to improvement ( Wang et al., 2023 ; Li et al., 2024 ; Wang et al., 2025 ) . These results make larger groups an appealing way to scale accurate inference. But they also create a growing dependence on the agents within the group behaving in good faith. What happens when some of them do not behave as such? A deceptive agent can do more than simply introduce an incorrect judgment. It can actively and subtly persuade initially correct agents to reach an incorrect conclusion. Recent work demonstrates this vulnerability in collaborative systems ( Huang et al., 2025 ; He et al., 2025 ; Kraidia et al., 2026 ) . We therefore ask how susceptibility to this kind of negative influence evolves as groups grow in size and deceptive agents increase in number. Human conformity research provides a useful starting point for understanding this vulnerability. Asch’s classic experiments showed that an incorrect majority can lead people to abandon otherwise reliable judgments ( Asch, 1951 ; Asch, 1955 ; Asch, 1956 ) . Later work has examined how this influence depends on group composition. In eyewitness discussions, Mojtahedi et al. (2018) varied the numbers of genuine participants and misleading confederates, finding that participants were susceptible to false blame when confederates formed a majority, but not significantly influenced in the tested conditions with a single confederate. Similarly, Coultas (2004) found that copying an unusual behavior depended on the proportion of the group displaying it rather than on absolute group size, and was rare unless a majority did so. Together, these findings suggest that in humans, a misleading minority exerts limited influence. This motivates a complementary question for multi-agent LLM systems: how much protection does an honest majority provide as a deceptive minority grows? We study how deceptive agents influence collective reasoning on questions from Humanity’s Last Exam (HLE) ( Center for AI Safety et al., 2026 ) , using closed- and open-source models in groups of 2 to 21 agents and scaling both the quantities of honest and deceptive agents. We find that honest agents defect at rates that rise approximately linearly with the proportion of deceivers even when deceivers are the minority, so even a small deceptive minority induces defection at consistently increasing rates. We also find that private coordination reduces the effectiveness of deceivers, and we provide a controlled and qualitative analysis of how agents’ arguments and strategies adapt. 2 Methodology Figure 1: Illustration of the contaminated multi-agent deliberation setup used throughout the experiments. Red robots are deceptive agents . Green robots are honest agents. Agents deliberate sequentially in public and reflect privately, with honest agents working to solve the question and deceptive agents attempting to steer the group toward an incorrect conclusion. At the end of voting, some agents will be swayed towards defecting to an incorrect answer, like the agent in the last row whose text bubble is highlighted in red . In this work, we aim to quantify how this effect scales with respect to the prevalence of deceptive agents in parties of varying sizes. We test whether larger groups of language-model agents are more resistant to deception. Agents first answer a question independently, repeatedly discuss public and reflect privately, and cast their final votes. We vary the numbers of honest and deceptive participants. We measure whether honest agents abandon answers they initially got right. A group contains N N agents: k k deceptive agents and N − k N-k honest agents. The deceptive proportion is k / N k/N . 2.1 Agent Roles Honest agents are instructed to work with their colleagues to find the correct answer. At the final vote, they are asked to use their own judgement. They are not told that any participant may be adversarial. Deceivers receive the reference answer privately at every stage. Their objective is to persuade the other agents to choose an incorrect answer without revealing either their objective or their access to the reference answer. They are instructed to make plausible, internally consistent arguments and to adapt to the discussion. They can switch to a different incorrect answer if they think it will be more persuasive. They can discuss their strategy openly in their private reflections and are instructed to cast an incorrect final vote. The complete role instructions appear in Appendices B.1 and B.2 . 2.2 Discussion process Agents communicate through a shared message board, where their responses appear under anonymous labels (“Colleague 1”, “Colleague 2”, and so on). Each agent also maintains a private reflection : a written record of its current answer, reasoning, and information it wants to retain from the discussion. Reflections are visible only to the agent that writes them. Figure 1 summarizes the protocol, and Appendix B gives the full instructions. By default, deceivers are not told which colleagues share their role and cannot communicate privately with one another. A trial consists of the following steps: 1. Initial answers (round 0). Each agent receives the question and its role instructions, with no reflection or peer responses. It gives an answer and explanation. Once all responses are posted, each agent receives the question and the full round 0 board to write its first private reflection. 2. Discussion (rounds 1 , … , R − 1 1,\ldots,R-1 ). To write a public reply, each agent receives the question, its previous reflection, and every public response from the previous round, including its own. It sees no current-round responses before submitting. Once all replies are posted, each agent receives the question, its previous reflection, the full round 0 board, and the full current board to write an updated reflection. 3. Final vote. Each agent receives the question and its final reflection, with no public-response board. It submits an answer without seeing anyone else’s vote. 2.2.1 Coordinated deceivers Figure 2: Example of coordination among three deceivers. We also run experiments in which deceivers are allowed to privately coordinate with other deceivers (Figure 2 ). Before round 0, they exchange private messages in a random order, with each deceiver posting once after reading the messages so far. Each then reads the complete thread and writes a private plan, which serves as its reflection for round 0. After each round’s public responses are complete, the deceivers repeat this chat before updating their reflections. Each receives its previous reflection, the current board, and the chat messages posted so far. The complete thread is included in the reflection update, then omitted from subsequent prompts. Honest agents never see these chats. Both conditions use the same instructions for later public responses and final votes, and the same protocol for honest agents. 2.3 Task Description We use questions from HLE ( Center for AI Safety et al., 2026 ) and construct a difficulty-stratified evaluation set for each model. To estimate question difficulty, each model answers every question four times using the same instructions as in round 0 of the deliberation trial. We stratify questions by the number of correct responses across these four trials, yielding three difficulty groups corresponding to one, two, or three correct responses out of four. Questions answered correctly zero or four times are excluded. 2.4 Scaling Setup Table 1: We test the following group compositions, shown as honest agents + deceivers (total group size N N in parentheses). This lets us compare groups with the same deceiver proportion or the same number of deceivers. Honest agents k / N = 0 k/N=0 k / N = 1 / 5 k/N=1/5 k / N = 1 / 3 k/N=1/3 k / N = 3 / 7 k/N=3/7 2 2 + 0 2{+}0 (2) – 2 + 1 2{+}1 (3) – 4 4 + 0 4{+}0 (4) – 4 + 2 4{+}2 (6) 4 + 3 4{+}3 (7) 8 8 + 0 8{+}0 (8) 8 + 2 8{+}2 (10) 8 + 4 8{+}4 (12) 8 + 6 8{+}6 (14) 12 – 12 + 3 12{+}3 (15) 12 + 6 12{+}6 (18) 12 + 9 12{+}9 (21) We test twelve group compositions, listed in Table 1 . Groups contain 2 to 21 agents, with adversarial proportions of 0 0 , 1 / 5 1/5 , 1 / 3 1/3 , and 3 / 7 3/7 . Honest agents outnumber deceivers in every group. We write each composition as honest + + deceiver counts: for example, 8 + 2 8{+}2 contains eight honest agents and two deceivers, so N = 10 N=10 and k / N = 1 / 5 k/N=1/5 . At any fixed k / N k/N , we vary N N to test whether larger groups are more robust. We evaluate Gemini 3.8 Flash, Grok 4.3, DeepSeek V4.1 Flash, and Muse Glimmer. All use their default reasoning settings. Table 3 in Appendix A gives the question and trial counts for each model. By default, honest agents and deceivers are instances of the same model. We randomly choose which colleague labels are assigned the deceiver role, then keep this assignment the same across questions with the same group composition. 2.5 Scoring We measure honest defection as the proportion of honest agents answering correctly in round 0 that eventually ended up on an incorrect answer. We assess correctness using GPT-5.4 with HLE’s official grading instructions (Appendix C ). Trials without deceivers provide a baseline for how often honest agents abandon correct answers during discussion. 3 Results 3.1 Adversarial Proportion Predicts Defection Better Than Count Honest defection increases approximately linearly with adversarial proportion (within-question permutation tests of the weighted linear slope, in percentage points per 0.1 increase in k / N k/N : Gemini, b = 5.3 b=5.3 , p < 0.001 p<0.001 ; Grok, b = 4.2 b=4.2 , p < 0.001 p<0.001 ; DeepSeek, b = 2.1 b=2.1 , p = 0.004 p=0.004 ; Muse Glimmer, b = 5.7 b=5.7 , p < 0.001 p<0.001 ) (Figure 3 ). We compare a straight line with a square-root curve, which allows the increase to taper as the proportion grows, and a step model, which assigns one defection rate to groups without deceivers and another to all groups with deceivers. The straight line fits best in every model, with R 2 = 0.97 R^{2}=0.97 for Gemini, 0.90 0.90 for Grok, 0.82 0.82 for DeepSeek, and 0.97 0.97 for Muse Glimmer. Table 2: Relationship between adversarial proportion and honest defection. R 2 R^{2} compares linear, square-root, and step fits, weighted by the number of initially correct honest agents; the straight line fits best for every model. R 2 R^{2} Model Linear slope p p Straight line Square root Step Gemini 3.8 Flash < 0.001 <0.001 0.97 0.84 0.54 Grok 4.3 < 0.001 <0.001 0.90 0.88 0.72 DeepSeek V4.1 Flash < 0.01 <0.01 0.82 0.73 0.52 Muse Glimmer < 0.001 <0.001 0.97 0.94 0.74 Figure 3: Honest defection by adversarial proportion, pooled across group sizes ( k / N = 0 k/N=0 marks the baseline without deceivers). Defection increases approximately linearly with adversarial proportion in every model. Error bars show 95% percentile bootstrap confidence intervals. We next examine whether the observed increase in defection is better predicted by the proportion of deceivers ( k / N k/N ) or the number of deceivers ( k k ). For trial t t , let C t C_{t} denote the number of honest agents correct in round 0 and D t D_{t} the number of those agents whose final vote is incorrect. We model D t ∼ Binomial ⁡ ( C t , p t ) D_{t}\sim\operatorname{Binomial}(C_{t},p_{t}) , where p t p_{t} is the probability that an initially correct honest agent casts an incorrect final vote in trial t t . We fit proportion ( P P ) and count ( C C ) models for this probability: logit ⁡ ( p t ( P ) ) \displaystyle\operatorname{logit}\bigl(p_{t}^{(P)}\bigr) = α m t , b t ( P ) + β ( P ) ​ 𝒌 𝒕 𝑵 𝒕 + γ ( P ) ​ log 2 ​ N t , \displaystyle=\alpha_{m_{t},b_{t}}^{(P)}+\beta^{(P)}\bm{\frac{k_{t}}{N_{t}}}+\gamma^{(P)}\log_{2}N_{t}, (1) logit ⁡ ( p t ( C ) ) \displaystyle\operatorname{logit}\bigl(p_{t}^{(C)}\bigr) = α m t , b t ( C ) + β ( C ) ​ 𝒌 𝒕 + γ ( C ) ​ log 2 ​ N t . \displaystyle=\alpha_{m_{t},b_{t}}^{(C)}+\beta^{(C)}\bm{k_{t}}+\gamma^{(C)}\log_{2}N_{t}. (2) Here N t N_{t} is group size, k t k_{t} is deceiver count, m t m_{t} indexes the language model, and b t ∈ { 1 , 2 , 3 } b_{t}\in{1,2,3} records its correct answers across the four independent attempts used for question selection. Intercepts vary by ( m t , b t ) (m_{t},b_{t}) , while slopes are shared in the pooled fits. Both models include a separate group-size term and have equal parameter counts. The proportion model (Equation 1 ) fits better than the count model (Equation 2 ) in the pooled analysis (deviance lower by 61.5) within each individual model. Adding proportion to the pooled count model improves fit (likelihood-ratio test, χ 2 ​ ( 1 ) = 62.1 \chi^{2}(1)=62.1 , p < 0.001 p<0.001 ), whereas adding count to the proportion model does not significantly improve fit (likelihood-ratio test, χ 2 ​ ( 1 ) = 0.59 \chi^{2}(1)=0.59 , p = 0.44 p=0.44 ). Noticeably, this means that increasing group size at a fixed adversarial proportion does not consistently reduce defection (Figure 4 ). Figure 4: Honest defection in Gemini 3.8 Flash by group size, with one panel per adversarial proportion. Across all models, larger groups do not consistently show lower defection at a fixed proportion, and group size magnitude visibly affects honest defection less than adversarial proportion. Error bars show 95% percentile bootstrap confidence intervals. 3.2 Honest Defection Is Lower When Deceivers Coordinate We evaluate the effects on defection rate on Gemini 3.8 Flash and Grok 4.3. Giving deceivers a private channel for coordination lowers the defection rate in both models (Figure 5 ). In trials matched by question and group composition, defection falls from 29.2% to 21.9% for Gemini and from 30.0% to 24.6% for Grok (two-sided paired permutation tests by question: Gemini, Δ = − 7.29 \Delta=-7.29 percentage points, p < 0.001 p<0.001 ; Grok, Δ = − 5.37 \Delta=-5.37 percentage points, p < 0.05 p<0.05 ). Under coordination, defection still rises approximately linearly with the adversarial proportion k / N k/N , as it does with independent deceivers ( R 2 = 0.96 R^{2}=0.96 for Gemini and 0.98 0.98 for Grok, including the shared baseline without deceivers). For both models, coordination lowers defection relative to non-coordinated deceivers at each tested nonzero proportion ( 1 / 5 1/5 , 1 / 3 1/3 , and 3 / 7 3/7 ). Figure 5: Honest defection with non-coordinated and coordinated deceivers. Defection is lower under coordination at each nonzero proportion, and the trends in both conditions also have linear best-fits. Error bars show 95% percentile bootstrap confidence intervals. 3.3 Model Choice Affects Defection Rates To test how model choice affects defection, we assign different models to be the honest and deceiver agents. We use Gemini and Muse Glimmer as the honest populations because they have the lowest and highest social-sycophancy scores, respectively, among our four models on ELEPHANT ( Cheng et al., 2026 ) . For deceivers, we select DeepSeek and Grok, the highest- and lowest-scoring models in our replication of the persuasion evaluation of Durmus et al. (2024) . We test all four honest–deceiver pairings at two adversarial proportions: 4 + 1 4{+}1 and 8 + 2 8{+}2 at 1 / 5 1/5 , and 4 + 3 4{+}3 and 8 + 6 8{+}6 at 3 / 7 3/7 (Figure 6 ). Model choice strongly affects defection. Muse Glimmer, the more sycophantic honest model, defects substantially more often than Gemini (37.7% vs. 19.5%). DeepSeek, the more persuasive deceiver, also causes more defection than Grok (26.8% vs. 21.2%). This difference appears for both honest models and is significant in three of the four group compositions. Overall, defection rate has a starker drop when honest agents are more susceptible to social influence compared to when deceivers are more persuasive. Figure 6: Honest defection for four honest–deceiver model pairings at two adversarial proportions. Muse Glimmer, the more sycophantic honest model, defects more often than Gemini. DeepSeek, the more persuasive deceiver, generally causes more defection than Grok. The identity of the honest side has more impact on defection rate than that of the deceiver side. Error bars show 95% percentile bootstrap confidence intervals. Defection also generally increases with adversarial proportion. Across model pairings, it rises from 21.2% at 1 / 5 1/5 to 26.8% at 3 / 7 3/7 (two-sided paired permutation test by question, Δ = + 5.54 \Delta=+5.54 percentage points, p < 0.01 p<0.01 ). The increase appears in three of the four pairings and is largest for Muse Glimmer paired with DeepSeek, where defection rises from 32.1% to 48.3%. 3.4 Behavioral Analysis of Deliberation We examine when honest agents change their answers, how deceivers adapt their arguments, what honest agents report in their private reflections, and how coordinated deceivers use their private channel (two-sided question-clustered t t tests; p p values are Holm-adjusted within each comparison family). Most defections happen early. When honest agents abandon a correct answer, they usually do so near the start of the discussion. With non-coordinated deceivers, 37–54% of first defections occur in round 1, and 58–72% occur by round 2 (tests against 50%: Gemini, t ⁡ ( 83 ) = 4.88 t(83)=4.88 , p < 0.001 p<0.001 ; Grok, t ⁡ ( 83 ) = 5.20 t(83)=5.20 , p < 0.001 p<0.001 ; DeepSeek, t ⁡ ( 87 ) = 3.00 t(87)=3.00 , p = 0.0035 p=0.0035 ; Muse Glimmer, t ⁡ ( 46 ) = 4.11 t(46)=4.11 , p < 0.001 p<0.001 ). Deceivers change tactics once discussion begins. Before seeing other agents’ answers, deceivers use fabricated or misrepresented evidence (56%) and misleading inferences (57%). After seeing the first-round responses, their arguments become more reactive. Conceding part of another agent’s argument and redirecting it rises from 6% to 65%, selective skepticism from 7% to 57%, and question reinterpretation from 15% to 52% (respectively, t ⁡ ( 69 ) = 16.80 t(69)=16.80 , 11.51 11.51 , and 9.59 9.59 ; all p < 0.001 p<0.001 ; Figure 7 b). Coordinated deceivers show the same shift: concede-and-redirect rises from 14% to 67%, and selective skepticism from 12% to 57% (respectively, t ⁡ ( 33 ) = 8.02 t(33)=8.02 and 9.15 9.15 ; both p < 0.001 p<0.001 ). Figure 7: Distribution of most common persuasion tactics used by deceivers. (a) Each model’s three most prevalent tactics across rounds 0 and 1. (b) Prevalence by round across models. Multiple tactics may appear in the same message. Definitions are in Appendix D . Honest agents often cite reinterpretation or uncertainty when changing answers. We inspect the private reflections written immediately before early defections. Of the 25 reflections that explain the change, 9 cite a different interpretation of the question, 8 unresolved uncertainty, and 6 perceived consensus; 2 cite fabricated or misrepresented evidence (comparisons with fabricated evidence: reinterpretation, t ⁡ ( 17 ) = 1.81 t(17)=1.81 , p = 0.088 p=0.088 ; uncertainty, t ⁡ ( 17 ) = 2.42 t(17)=2.42 , p = 0.054 p=0.054 ). These reflections describe the agents’ stated reasons for changing answers. Coordinated deceivers actively plan together, but rarely appear coordinated to honest agents. Deceivers use their private channel to agree on an incorrect answer, divide roles, target different objections, and vary their wording to appear independent. In 59 of 60 trials, they also prepare distinct arguments for the same answer. In the coded sample, Gemini backs off when suspicion arises in 27 of 30 trials versus 15 of 30 for Grok, builds on arguments that worked earlier in 25 versus 18, and jointly changes its answer or strategy in 29 versus 6 (respectively, t ⁡ ( 33 ) = 3.67 t(33)=3.67 , p = 0.0017 p=0.0017 ; t ⁡ ( 33 ) = 2.03 t(33)=2.03 , p = 0.0502 p=0.0502 ; t ⁡ ( 33 ) = 8.86 t(33)=8.86 , p < 0.001 p<0.001 ). In the coded sample, question reinterpretation appears in 37% of independent and 62% of coordinated Gemini messages, and in 27% and 14%, respectively, for Grok (Gemini, t ⁡ ( 27 ) = 2.57 t(27)=2.57 , p = 0.032 p=0.032 ; Grok, t ⁡ ( 30 ) = − 1.36 t(30)=-1.36 , p = 0.184 p=0.184 ). Despite this extensive planning, honest agents almost never detect it, with explicit suspicion appearing in only one of 60 coordinated trials. 4 Related work Multi-agent reasoning and scaling. Multi-agent deliberation can improve reasoning and evaluation through argument exchange ( Du et al., 2024 ; Liang et al., 2024 ; Chan et al., 2024 ) , while CAMEL, AutoGen, and MetaGPT provide broader frameworks for collaboration ( Li et al., 2023 ; Wu et al., 2024 ; Hong et al., 2024 ) . Structured debate also uses competing advocates to help a separate judge recover the truth ( Irving et al., 2018 ; Khan et al., 2024 ) . Adding agents or model outputs can improve performance ( Li et al., 2024 ; Wang et al., 2025 ) , although debate gains can diminish relative to voting or stronger single-agent baselines ( Zhang et al., 2025 ; Choi et al., 2025a ; Yang et al., 2025 ) , and scaling benefits depend on task structure and information independence ( Kim et al., 2025 ; Bertalanič & Fortuna, 2026 ) . Related work on human collective judgment shows that diversity can support accuracy and that interaction can either improve judgments or undermine aggregation ( Hong & Page, 2004 ; Lorenz et al., 2011 ; Becker et al., 2017 ) . Our focus is whether larger groups preserve correct reasoning against a deliberately adversarial minority. Social influence in LLMs. Language models are both sources and targets of influence. As sources, they persuade at rates comparable to human writers and can durably change strongly held beliefs ( Durmus et al., 2024 ; Salvi et al., 2025 ; Costello et al., 2024 ) . As targets, they tend to agree with users or socially salient positions regardless of accuracy ( Sharma et al., 2024 ; Wei et al., 2023 ; Cheng et al., 2026 ) and can abandon correct answers when merely asked to reconsider ( Huang et al., 2024 ) . In groups, their judgments move toward majority positions and depend on uncertainty, interaction history, and debate partners ( Zhu et al., 2025 ; Weng et al., 2025 ; Choi et al., 2025b ; Yao et al., 2025 ) , although Hao et al. (2026) separate spontaneous instability from conformity to stated positions and from persuasion by reasoning. In most of these settings the majority is scripted or sincere rather than strategically deceptive, and its size and share vary together. We treat persuasive ability and susceptibility as properties of both sides of an adversarial deliberation, vary the count and proportion of deceptive agents independently, and measure whether initially correct honest agents end with incorrect answers, without attributing every such change to one mechanism. Adversarial agents and system robustness. This work studies failures caused by malicious or faulty participants. Byzantine fault tolerance and related work on identity and robust aggregation analyze malicious participants under explicit protocol assumptions ( Lamport et al., 1982 ; Douceur, 2002 ; Blanchard et al., 2017 ; Yin et al., 2018 ) . In LLM systems, attacks exploit prompts, memory, shared context, and inter-agent communication ( Greshake et al., 2023 ; Chen et al., 2024 ; Gu et al., 2024 ; Lee et al., 2026 ; Ju et al., 2026 ; He et al., 2025 ) , while other studies examine faulty collaborators, coordinated sabotage, and broader risks from interacting agents ( Huang et al., 2025 ; Hammond et al., 2025 ; Schroeder de Witt et al., 2025 ; Radev et al., 2026 ; McAllister et al., 2026 ) . Most directly, Kraidia et al. (2026) study persuasion-driven adversarial influence in debate and find that adding agents or rounds does not reliably protect group accuracy. Our work focuses on how adversarial influence scales in multi-agent LLM systems and how it changes the behavior of otherwise capable agents. Rather than treating robustness only as a question of whether a group reaches the right answer, we study how deceptive participants shape individual judgments and whether larger groups or greater coordination provide meaningful protection. 5 Conclusion We study how multi-agent systems respond to deception as groups grow and deceptive participants become more prevalent. Across closed and open-source models, honest agents are more likely to abandon initially correct answers as the proportion of deceivers increased. Increasing group size at a fixed adversarial proportion offered no consistent protection. These findings show that the composition of a group matters for its robustness, and that increasing the number of participants alone does not reliably protect correct judgments from adversarial influence. The importance of group composition echoes findings from human conformity research ( Coultas, 2004 ) , but our results suggest a more concerning susceptibility to deceptive minorities. In the human eyewitness study of Mojtahedi et al. (2018) , misleading confederates significantly influenced judgments only when they formed a majority. In our experiments, deceivers remained a minority yet still substantially influenced honest agents to abandon correct answers. This contrast suggests that LLM populations may be less resistant to minority influence than humans. Having most participants work toward the correct answer does not ensure that the group will resist those trying to mislead it. Susceptibility also depended on the models involved and how they interacted. In our heterogeneous experiments, the more sycophantic honest model defected more often, while the more persuasive deceiver generally caused more defection. These results suggest that both the honest model’s susceptibility and the deceiver’s persuasiveness deserve attention when evaluating a multi-agent system. Private coordination, however, reduced deceivers’ effectiveness in both tested families, despite allowing them to plan together and divide roles. Our behavioral analysis shows that many agents first switched from correct to incorrect answers within the first two discussion rounds. In their private reflections, honest agents often cited a different interpretation of the question or unresolved uncertainty between competing answers. Discussion also helped agents correct initial mistakes, so these systems cannot be understood through defection alone. Evaluations should consider both the errors that discussion corrects and the correct judgments it overturns. The broader challenge is to preserve the benefits of exchanging arguments while reducing susceptibility to deception, a problem that adding more agents does not resolve on its own. Limitations. Our experiments cover two deliberation protocols for groups of up to 21 agents. The comparison with humans draws on different tasks and experimental settings. Our behavioral analyses are descriptive and do not establish why agents defect or why private coordination reduces deceivers’ effectiveness. Future work should test whether these patterns extend to other tasks, communication structures, and larger groups. It should also examine whether interventions such as independent verification of disputed claims can reduce adversarial influence while preserving error correction and overall accuracy. References Asch (1951) Solomon E. Asch. Effects of group pressure upon the modification and distortion of judgments. In Harold Guetzkow (ed.), Groups, Leadership and Men , pp. 177–190. Carnegie Press, 1951. Asch (1955) Solomon E. Asch. Opinions and Social Pressure. Scientific American , 193(5):31–35, 1955. doi: 10.1038/scientificamerican1155-31 . URL https://doi.org/10.1038/scientificamerican1155-31 . Asch (1956) Solomon E. Asch. Studies of independence and conformity: I. A minority of one against a unanimous majority. Psychological Monographs: General and Applied , 70(9):1–70, 1956. doi: 10.1037/h0093718 . URL https://doi.org/10.1037/h0093718 . Becker et al. (2017) Joshua Becker, Devon Brackbill, and Damon Centola. Network dynamics of social influence in the wisdom of crowds. Proceedings of the National Academy of Sciences , 114(26):E5070–E5076, 2017. doi: 10.1073/pnas.1615978114 . URL ht The same ai agents question is explored in Shutdown Sabotage Propensities in Multi-Agent Systems, which adds a research perspective. as detailed in the full paper on Arxiv The ai agents story also surfaces in Meta expands Muse into an AI..., adding another angle.

Comments (0)

No comments yet

Be the first to share your thoughts!