What the paper is about
Multi-agent LLM systems are often expected to improve as team size increases, yet the scaling behavior may depend on task structure. Our central contribution is to introduce Steiner's taxonomy of group tasks as a framework for analyzing multi-agent LLM scaling and focusing the analysis on disjunctive and compensatory tasks. We model independently sampled agents as conditionally independent given the item, which yields their large-team limits: plurality voting converges to the model's modal answer, and averaging converges to the model's item-level bias. Across selected representative benchmarks, 13 open-weight models, and teams of up to 30 agents, we find qualitatively different scaling behavior. On disjunctive tasks, the probability that at least one agent is correct grows by 5-20 points with team size, but plurality voting over agents that answer directly realises almost none of this potential, as the model predicts to within 0.5 points on average. Multi-round revision raises accuracy considerably, yet the gain is nearly the same with one peer as with 29. In contrast, scaling provides little benefit on Fermi estimation, despite its natural suitability for aggregation: item-level biases shared across the samples of a model account for about 87% of the squared error, so averaging reduces error by only about 6%. Combining model families helps on Fermi estimation but does not surpass the strongest member on disjunctive tasks. These results show that task structure, together with the mechanism combining member outputs, is a fundamental determinant of team scaling. The same large language models question is explored in CERA-MoA, which adds a research perspective.
What it covers
Multi-agent Scaling Across Disjunctive and Compensatory Tasks Carolina Fortuna Blaž Bertalanič Abstract Multi-agent LLM systems are often expected to improve as team size increases, yet the scaling behavior may depend on task structure. Our central contribution is to introduce Steiner’s taxonomy of group tasks as a framework for analyzing multi-agent LLM scaling and focusing the analysis on disjunctive and compensatory tasks. We model independently sampled agents as conditionally independent given the item, which yields their large-team limits: plurality voting converges to the model’s modal answer, and averaging converges to the model’s item-level bias. Across selected representative benchmarks, 13 open-weight models, and teams of up to 30 agents, we find qualitatively different scaling behavior. On disjunctive tasks, the probability that at least one agent is correct grows by 5–20 points with team size, but plurality voting over agents that answer directly realises almost none of this potential, as the model predicts to within 0.5 points on average. Multi-round revision raises accuracy considerably, yet the gain is nearly the same with one peer as with 29. In contrast, scaling provides little benefit on Fermi estimation, despite its natural suitability for aggregation: item-level biases shared across the samples of a model account for about 87% of the squared error, so averaging reduces error by only about 6%. Combining model families helps on Fermi estimation but does not surpass the strongest member on disjunctive tasks. These results show that task structure, together with the mechanism combining member outputs, is a fundamental determinant of team scaling. [Open code placeholder] 1 Introduction Motivation. With the rapid emergence of Large Language Model (LLM) multi-agent frameworks ( Wu et al., 2024 ; Li et al., 2023 ; Du et al., 2024 ) , researchers have deployed LLM ensembles and multi-agent debates across a multitude of reasoning and decision-making problems under the implicit assumption that scaling team size N N universally enhances system performance. As LLMs are trained on human generated knowledge largely encoded through language, that inherently also captures cultural and behavioral traits, we turn to observing multi-agent behavior in a similar way human group behavior was observed and studied. The ability of a group to outperform its individual members is among the most studied phenomenon in cognitive science, economics, and social psychology. In his classic 1907 experiment, Francis Galton observed that while individual fairgoers made highly scattered guesses regarding the weight of an ox, the median of their estimates fell within 0.8 % 0.8% of the true weight ( Galton, 1907 ; Surowiecki, 2004 ) . Conversely, the foundational Ringelmann effect ( Ingham et al., 1974 ) reveals the opposing tension in collective dynamics: as group size increases, teams frequently suffer from diminishing per-member returns and performance saturation. In Ivan Steiner’s seminal taxonomy of group processes ( Steiner, 1972 ) , resolving whether a collective achieves Galtonian crowd wisdom or succumbs to Ringelmann-style process loss depends on the underlying task structure. The tasks observed by Galton are the so-called compensatory tasks Steiner (1972) where group performance is determined by statistical aggregation of continuous estimates. In human crowds, because individual participants possess idiosyncratic cognitive biases and independent heuristic noise, their random errors compensate for one another ( Larrick and Soll, 2006 ; Davis-Stober et al., 2014 ) . Ringelmann originally observed this performance drop on additive tasks , where per-member productivity decayed as group size increased due to coordination losses and social loafing ( Latané et al., 1979 ; Ingham et al., 1974 ) . It was later shown by Kerr and Bruun (1983) that group performance losses are fundamentally mediated by Steiner’s task typology and the perceived dispensability of individual effort. While additive and conjunctive tasks frequently trigger free-rider effects or coordination bottlenecks, compensatory tasks can circumvent these losses when members provide independent, unbiased estimates that aggregate symmetrically. This established that collective human team scaling laws are fundamentally task-dependent . The same observation has been recently made by Kim et al. (2026) : LLM agent team performance is task dependent. At the same time, inspired by Ringlemann’s findings, Bertalanič and Fortuna (2026) developed a scaling law for effective team size. Prior art. Without classifying them into the Steiner taxonomy nor recognizing them explicitly as disjunctive or compensatory tasks, the community has already investigated related multi-agent performance. The rich state of the art on disjunctive tasks represented by benchmarks such as MMLU, GSM8K, MATH500, ARC reports gains from sampling and voting ( Wang et al., 2023 ; junyou li et al., 2024 ) , and from debate ( Du et al., 2024 ) although these gains can saturate or reverse as the number of calls grows ( Chen et al., 2024 ; Bertalanič and Fortuna, 2026 ) . Debate induces conformity ( Weng et al., 2025 ) , and correlated errors limit ensembles ( Kim et al., 2025 ) . Improvements have also been reported for compensatory tasks represented by Fermi Kalyan et al. (2021) ; Liu et al. (2025) and guesstimate Chuang et al. (2025) problems. The closest related works are Chen et al. (2024) , who explain non-monotone voting by a mixture of easy and hard queries, the basis of our plurality limit, and Bertalanič and Fortuna (2026) , who fit a law for the effective team size N eff / N N_{\text{eff}}/N across 44 conditions and find that only heterogeneous teams escape hard ceilings; homogeneous teams also saturate in Yang et al. (2026) , and LLM teams fall short of their best member in Pappu et al. (2026) . This work. Our central contribution is to introduce Steiner’s taxonomy of group tasks as a framework for analyzing multi-agent LLM scaling enabling the formalization and characterization of multi-agent performance on formal task types: 1. We connect benchmarks extensively used the evaluation of LLM systems to the Steiner taxonomy and mathematically formalise agent behaviour on these tasks, including the large-team limits of independently sampled agents (Section 2 ). 2. We investigate multi-agent scaling with 13 open-weight models and teams of up to 30 agents, analysing the independence and diversity assumptions (Section 3.2 ). On disjunctive tasks, pass@ N N grows steadily with team size, whereas plurality voting converges to the model’s modal answer, as predicted to within 0.5 points by a conditional-independence model. Revision rounds raise accuracy, but almost equally with one peer or 29 (Section 3.3 ). 3. We show that unbiasedness fails for compensatory Fermi estimate tasks at the level of individual items: item-level bias accounts for about 87% of the squared log-error, so averaging reduces error by only about 6%; part of the apparent cross-model correlation is due to implausible reference labels (Section 3.4 ). 4. We show that heterogeneous teams improve substantially on their average member but, on disjunctive tasks, don’t exceed their strongest member; on the compensatory Fermi estimation, a team of five 7B–8B models outperforms its members (Sections 3.3 and 3.4 ). 2 Taxonomy and Theoretical Scaling Framework 2.1 Steiner’s Task Taxonomy and Modern LLM Benchmarks We map representative benchmarks used to evaluate LLM-based multi-agent teams onto Steiner’s task types ( Steiner, 1972 ) , according to the rule by which a multi-agent system combines member contributions (Table 1 ); the mapping describes a combination rule rather than an intrinsic property of a benchmark. The remainder of the paper focuses on disjunctive tasks, which depend on the skill of individual members, and compensatory tasks, which depend on error cancellation; we leave the other task types for future work. Disjunctive Tasks in which the outcome is established through Y = max i { y i } Y=\max_{i}{y_{i}} or plurality/majority/verifier selection and/or debate. The tasks are unitary and optimizing and group success requires that at least one agent discovers a valid solution, and the group then recognizes it as such. Examples include mathematical reasoning, multiple-choice knowledge and science questions, and code generation (Table 1 ), which MASes address through plurality voting ( Wang et al., 2023 ) , verifier re-ranking, and debate ( Du et al., 2024 ; Wang et al., 2025 ) . Table 1: Categorization of tasks according to the Steiner taxonomy. Disjunctive and compensatory tasks are briefly discussed in this section; with additional details on all tasks in Appendix A . Task Combination Operator Output Domain Top-Conference Benchmarks (2024–2026) LLM Aggregation Mechanism Dis- junc- tive Y = max i { y i } Y=\max_{i}{y_{i}} (or plurality) Discrete ( 𝒴 \mathcal{Y} ) MATH-500, GSM8K, MMLU-Pro ( Wang et al., 2024 ) , GPQA ( Rein et al., 2024 ) , BigCodeBench ( Zhuo et al., 2025 ) , ARC Plurality voting, verifier selection, multi-agent debate Com- pen- satory Y = ( ∏ i = 1 n x i ) 1 n Y=\left(\prod_{i=1}^{n}x_{i}\right)^{\frac{1}{n}} (or median) Continuous ( ℝ + , [ 0 , 1 ] \mathbb{R}^{+},[0,1] ) ForecastBench ( Karger et al., 2025 ) , Guesstimate ( Chuang et al., 2025 ) , RealFP ( Kalyan et al., 2021 ) , SynthFP, SciOly Geometric/arithmetic mean, median, calibrated Brier aggregation Addi- tive Y = ∑ i y i Y=\sum_{i}y_{i} or ⋃ i S i \bigcup_{i}S_{i} Cumulative sets DAT ( Olson et al., 2021 ) , open-ended ideation, parallel instruction synthesis Deduplicated set union, parallel greedy generation Con- junc- tive Y = min i { y i } Y=\min_{i}{y_{i}} Boolean / state chain SWE-bench ( Jimenez et al., 2024 ) , TheAgentCompany ( Xu et al., 2026 ) , OSWorld ( Xie et al., 2024 ) , BrowseComp ( Kim et al., 2026 ) When modeled as strict sequential pipeline, multi-step dependency verification Dis- cre- tionary Y = f group ( { y i } ) Y=f_{\text{group}}({y_{i}}) Dynamic Heteroge- neous AgentBench ( Liu et al., 2024 ) , ChatEval ( Chan et al., 2024 ) , MoA ( Wang et al., 2025 ) Meta-agent router, judge arbitration, dynamic role assignment Compensatory Tasks in which the outcome is determined by Y = Aggr ( x 1 , … , x n ) Y=\text{Aggr}(x_{1},\dots,x_{n}) via statistical aggregation (geometric mean, arithmetic mean, or median) of continuous numerical estimates ( x i ∈ ℝ + x_{i}\in\mathbb{R}^{+} ) or calibrated probabilities ( p i ∈ [ 0 , 1 ] p_{i}\in[0,1] ). xamples include forecasting ( Karger et al., 2025 ) , guesstimation ( Chuang et al., 2025 ) , and Fermi estimation ( Kalyan et al., 2021 ; Liu et al., 2025 ) . 2.2 Modeling Disjunctive Tasks Through Voting 2.2.1 The Pure Disjunctive Case In Steiner’s classical disjunctive framework ( Steiner, 1972 ) , a task is disjunctive if group success requires only that at least one team member generates the correct solution y ∗ ∈ 𝒴 y^{}\in\mathcal{Y} and that the group recognises it. Let p ∈ ( 0 , 1 ) p\in(0,1) denote the probability that an individual agent independently produces the correct answer. For independent agents and an ideal verifier, the probability of collective success is P disj ( N ) = 1 − ( 1 − p ) N P_{\text{disj}}(N)=1-(1-p)^{N} , so the collective failure probability decays exponentially with team size N N : P ( failure ) = ( 1 − p ) N → 0 P(\text{failure})=(1-p)^{N}\to 0 when N → ∞ N\to\infty . We refer to this, estimated empirically, as the proportion of items on which at least one agent is correct, as pass@ N N . 2.2.2 Majority Voting and the Condorcet Jury Theorem In the absence of an oracle verifier, multi-agent frameworks aggregate discrete candidate outputs through majority or plurality voting ( Wang et al., 2023 ; Du et al., 2024 ) . For a binary decision between alternative solutions, let each agent cast an independent vote V i ∼ Bernoulli ( p ) V_{i}\sim\text{Bernoulli}(p) for the correct solution y ∗ y^{} , where p > 1 2 p>\frac{1}{2} . The majority vote decision y ^ maj = ∑ i = 1 n V i > n 2 \hat{y}{\text{maj}}=\sum{i=1}^{n}V_{i}>\frac{n}{2} succeeds with probability ∑ k > N / 2 ( N k ) p k ( 1 − p ) N − k \sum_{k>N/2}\binom{N}{k}p^{k}(1-p)^{N-k} . By Hoeffding’s inequality under Bernoulli random variables where a i = 0 , b i = 1 a_{i}=0,b_{i}=1 and deviation t = p − 1 2 t=p-\frac{1}{2} , the probability that the majority vote errors is strictly upper-bounded by: P ( y ^ maj ≠ y ∗ ) ≤ exp ( − 2 n ( p − 1 / 2 ) 2 ) P(\hat{y}{\text{maj}}\neq y^{*})\leq\exp\left(-2n\left(p-1/2\right)^{2}\right) (1) Under uncorrelated agent assumption, Eq. 1 shows that for agent accuracies strictly greater than random guessing ( p > 0.5 p>0.5 ), voting ensembles are bounded by exponential error suppression as team size N N increases. 2.3 Modeling Compensatory Tasks through Geometric Aggregation Next, we contrast the disjunctive voting paradigm with the classic statistical model of group judgment on compensatory tasks ( Galton, 1907 ; Surowiecki, 2004 ; Steiner, 1972 ) . Because continuous estimation tasks (e.g., Fermi problems) span many orders of magnitude, errors are evaluated in logarithmic space ( Kalyan et al., 2021 ) . Let g ∈ ℝ + g\in\mathbb{R}^{+} be the positive ground truth, and x i ∈ ℝ + x{i}\in\mathbb{R}^{+} be agent i i ’s estimate for i ∈ { 1 , … , N } i\in{1,\dots,N} . The signed log-error is e i = log 10 ( x i / g ) e_{i}=\log_{10}\left(x_{i}/g\right) . The team estimate X ¯ N \bar{X}{N} is formed via geometric mean aggregation, corresponding to arithmetic mean in log-error space: e ¯ N = log 10 ( X ¯ N g ) = 1 N ∑ i = 1 N e i \bar{e}{N}=\log_{10}\left(\frac{\bar{X}{N}}{g}\right)=\frac{1}{N}\sum{i=1}^{N}e_{i} (2) 2.3.1 The Statistical Principle of Team Superiority To show why an averaged team judgment is expected to be more accurate than a typical individual ( Gigone and Hastie, 1997 ; Armstrong, 2001 ) , we analyse the expected value and the variance of the collective error e ¯ N \bar{e}{N} . In modelling human team performance, the theory relies on i) unbiased estimates where on average, individual human errors balance out to zero across the crowd: 𝔼 [ e i ] = 0 for all i ∈ { 1 , … , N } \mathbb{E}[e{i}]=0\quad\text{for all }i\in{1,\dots,N} and ii) independence where team members do not influence one another, meaning individual errors are statistically uncorrelated: Cov ( e i , e j ) = 0 for all i ≠ j \text{Cov}(e_{i},e_{j})=0\quad\text{for all }i\neq j . If each individual’s judgment error has a variance of Var ( e i ) = σ 2 \text{Var}(e_{i})=\sigma^{2} , the variance of the team’s collective estimate is calculated as: Var ( e ¯ N ) = 1 N 2 ∑ i = 1 N Var ( e i ) = σ 2 / N \text{Var}(\bar{e}{N})=\frac{1}{N^{2}}\sum{i=1}^{N}\text{Var}(e_{i})=\sigma^{2}/N , so that the standard error of the team’s judgment shrinks at the rate SE ( e ¯ N ) = σ / N \text{SE}(\bar{e}{N})=\sigma/\sqrt{N} . As team size N N grows, the variance approaches zero ( lim N → ∞ Var ( e ¯ N ) = 0 \lim{N\to\infty}\text{Var}(\bar{e}{N})=0 ). This shows that the averaged team judgment will systematically be more accurate and consistent than any single individual judgment ( Clemen, 1989 ; Page, 2007 ) . 2.4 Effective Team Capacity and Scaling Limits Across Task Types Comparing the mathematical foundations of the ideal disjunctive model (Section 2.2 ) and the ideal compensatory model (Section 2.3 ) reveals the theoretical origin of the scaling dichotomy: the failure probability of independent agents decreases exponentially in disjunctive tasks while it drops polynomially in compensatory tasks. More details of what behavior is expected from correlated and biased agents on disjunctive and compensatory task types are available in Appendix B.1 . For correlated draws (i.e. answers sampled from a language model), the Kish design effect ( Kish, 1965 ) provides the effective team capacity N eff N{\text{eff}} defined as the size of an ideal, independent ensemble ( ρ = 0 \rho=0 ) that achieves the equivalent variance reduction as a correlated ensemble of size N N : N eff ( N ) = N / ( 1 + ( N − 1 ) ρ ) N_{\text{eff}}(N)=N/(1+(N-1)\rho) . For any positive pairwise error correlation ρ > 0 \rho>0 : lim n → ∞ N eff ( N ) = 1 / ρ \lim_{n\to\infty}N_{\text{eff}}(N)=1/\rho . An important consequence arises when agents are independent samples from the same model, as in our experiments. Conditional on the item k k , the outputs of such agents are independent and identically distributed with an item-specific distribution P k P_{k} . The pairwise correlation of their correctness across items then equals ρ = Var k ( p k ) / ( p ¯ ( 1 − p ¯ ) ) \rho=\mathrm{Var}{k}(p{k})/(\bar{p}(1-\bar{p})) , where p k p_{k} is the single-agent success probability on item k k and p ¯ \bar{p} its mean ( Boland, 1989 ; Ladha, 1992 ) ; that is, ρ \rho measures heterogeneity of item difficulty rather than interaction between agents. Such a team can be seen as a “crowd within” a single model, whose repeated judgments are more correlated than those of different individuals ( Vul and Pashler, 2008 ; Herzog and Hertwig, 2009 ) , and correlated judges add little information ( Hogarth, 1978 ) . Proposition 2.1 (Task-type scaling under correlation and bias) . Let agents be conditionally independent given the item, and let N eff ( N ) ≤ 1 / ρ N_{\text{eff}}(N)\leq 1/\rho be defined as above. Disjunctive scaling regime. With an oracle verifier, accuracy is 1 − 𝔼 k [ ( 1 − p k ) N ] 1-\mathbb{E}{k}[(1-p{k})^{N}] , which increases to the fraction of items with p k > 0 p_{k}>0 . Under plurality voting, the failure probability on item k k decays exponentially in N N if the correct answer is the unique mode of P k P_{k} and tends to one if it is not, so that accuracy converges to the fraction π \pi of items on which the correct answer is the modal answer, that is, of the easy items of Chen et al. (2024) . For binary outcomes, the attainable voting gain satisfies | π − p ¯ | ≤ 2 p ¯ ( 1 − p ¯ ) ( 1 − ρ ) . |\pi-\bar{p}|\leq 2,\bar{p}(1-\bar{p})(1-\rho). (3) The gain from voting is therefore small when ρ \rho is large, even though the exponential convergence of Eq. 1 holds on every item with p k > 1 2 p_{k}>\frac{1}{2} . Compensatory scaling regime. For continuous estimates with systematic bias b b , variance σ 2 \sigma^{2} , and equicorrelated errors with correlation ρ \rho , the mean squared error of the geometric mean is linear in 1 / N eff 1/N_{\text{eff}} : MSE comp ( N ) = b 2 + σ 2 N eff ( N ) → N → ∞ b 2 + σ 2 ρ . \text{MSE}{\text{comp}}(N)=b^{2}+\frac{\sigma^{2}}{N{\text{eff}}(N)}\xrightarrow[N\to\infty]{}b^{2}+\sigma^{2}\rho. (4) With item-level biases b k b_{k} , the same decomposition reads MSE ( N ) = 𝔼 k [ b k 2 ] + 𝔼 k [ σ k 2 ] / N \mathrm{MSE}(N)=\mathbb{E}{k}[b{k}^{2}]+\mathbb{E}{k}[\sigma{k}^{2}]/N , so the attainable reduction is the within-item share of error, 1 − β 1-\beta with β = 𝔼 [ b k 2 ] / ( 𝔼 [ b k 2 ] + 𝔼 [ σ k 2 ] ) \beta=\mathbb{E}[b_{k}^{2}]/(\mathbb{E}[b_{k}^{2}]+\mathbb{E}[\sigma_{k}^{2}]) . For example, with ρ ≈ 0.9 \rho\approx 0.9 (effective capacity N eff ≈ 1.1 N_{\text{eff}}\approx 1.1 ), averaging can remove at most σ 2 ( 1 − ρ ) ≈ 0.1 σ 2 \sigma^{2}(1-\rho)\approx 0.1,\sigma^{2} , which is small relative to a dominant bias term b 2 b^{2} . Proofs in Appendix B.3 . 3 Evaluation of Disjunctive vs. Compensatory Scaling 3.1 Experimental Setup To evaluate empirically whether LLM multi-agent team performance follows Proposition 2.1 , we generated a corpus of approximately 6.8 × 10 7 6.8\times 10^{7} per-agent generations for homogeneous and heterogeneous teams, evaluating all models on the same grid of team sizes N ∈ { 1 , 2 , 3 , 5 , 7 , 10 , 15 , 20 , 25 , 30 } N\in{1,2,3,5,7,10,15,20,25,30} and communication rounds (Rounds 1 to 3). Models: We evaluate 13 open-weight models between 3B and 20B parameters, ordered by parameter count: smollm3-3b, qwen2.5-3b, qwen3-4b, mistral-7b, olmo2-7b, qwen2.5-7b, llama3.1-8b, marin-8b, qwen3-8b, phi4-14b, qwen3-14b, r1-distill-qwen-14b, and gpt-oss-20b. Unless stated otherwise, averages are macro-averages over the 13 models. Benchmark Tasks: Disjunctive tasks: GSM8K ( Cobbe et al., 2021 ) (grade-school mathematics), GSM-Hard ( Gao et al., 2023 ) (GSM8K problems with larger numbers), MATH-500 ( Hendrycks et al., 2021b ) (competition mathematics), MMLU-Hard, a difficult subset of MMLU ( Hendrycks et al., 2021a ) (multitask knowledge), and ARC-Challenge ( Clark et al., 2018 ) (science question answering). Compensatory task: RealFP Fermi problems ( Kalyan et al., 2021 ) , which require numerical order-of-magnitude estimation across physics, geography, and demographics. Inference, Deliberation, and Metric Computation Protocol. All models are decoded at temperature T = 0.4 T=0.4 , with three runs per configuration on the full test sets (Appendix C ). The prompts request the final answer before a brief rationale (Appendix C ); this answer-first format isolates the answer distribution of the instruction-tuned model itself rather than that of a test-time reasoning procedure, and makes answers robust to truncation. Two models, gpt-oss-20b and r1-distill-qwen-14b, reason before the formatted answer regardless of the prompt and thus provide a contrast between the two regimes. In Round 1, agents answer independently; in Rounds 2 and 3, each agent sees its own and its teammates’ previous responses and may revise its answer. Disjunctive answers are aggregated by plurality voting, and we additionally report pass@ N N ; Fermi estimates are aggregated by the geometric mean (Eq. 2 ) and scored by the absolute log-error (ALE) and by within-one-order-of-magnitude accuracy. We clipped ALE at ± 3 \pm 3 orders of magnitude so that implausible reference labels (Section 3.2 ) and occasional runaway estimates do not dominate the mean. The conclusions do not depend on this threshold (Appendix Table 7 ). For RealFP, we analyse the runs generated with the decomposition prompt (10 of the 13 models; Appendix C ). Confidence intervals are 95% item-bootstrap intervals. Figure 1: Scaling with team size (average over 13 models; bands are 95% item-bootstrap intervals). The team answer is the plurality vote on the disjunctive tasks and the geometric mean on a representative compensatory RealFP (right), where accuracy means an estimate within one order of magnitude. Exact values are given in Appendix Table 4 . 3.2 Empirical Scaling Dynamics on Disjunctive vs Compensatory Tasks Disjunctive From Fig. 1 and Table 4 we can see that the probability that at least one of N N independent agents is correct (pass@ N N , Section 2.2 ) rises from the single-agent accuracy to pass@30 of 54.4 % 54.4% on GSM8K (single agent: 34.3 % 34.3% ), 47.1 % 47.1% on MATH-500 ( 29.7 % 29.7% ), 67.7 % 67.7% on MMLU-Hard ( 52.4 % 52.4% ), 28.9 % 28.9% on GSM-Hard ( 18.0 % 18.0% ), and 88.4 % 88.4% on ARC-Challenge ( 83.1 % 83.1% ). In Steiner’s terms, the potential productivity of the team increases with its size. Allowing agents to revise their answers over multiple rounds leads to large accuracy gains on arithmetic tasks. GSM8K accuracy nearly doubles from 34.1 % 34.1% (Round 1, N = 5 N=5 ) to 61.9 % 61.9% (Round 3, N = 5 N=5 ); GSM-Hard accuracy rises from 18.7 % 18.7% to 32.9 % 32.9% and MATH-500 accuracy from 31.3 % 31.3% to 39.7 % 39.7% . The gain is, however, nearly independent of team size: on GSM8K it is 26.5 26.5 points with a single peer ( N = 2 N=2 ) and 26.6 26.6 points with 29 peers (paired difference − 0.1 -0.1 , 95% CI [ − 0.6 , 0.3 ] [-0.6,0.3] ). Because the Round-1 format asks for the answer before the rationale, a substantial part of this gain is likely due to the opportunity to reason before committing to an answer; a single agent given a five-fold token budget improves by 12.4 12.4 points on GSM8K (Appendix C.1 shows an example). On compensatory problems, of which Fermi are representative, Fig. 1 and Table 4 show that increasing team size N N provides little improvement in either round. Under Round 1 averaging, accuracy within one order of magnitude rises from 31.6 % 31.6% for a single agent to 33.1 % 33.1% at N = 7 N=7 and 33.4 % 33.4% at N = 30 N=30 , and the clipped log-error decreases from 1.84 1.84 to 1.73 1.73 at N = 5 N=5 and 1.71 1.71 at N = 30 N=30 . Revision does not help: Round-3 accuracy differs from Round-1 accuracy by between − 0.6 -0.6 and + 0.7 +0.7 points across team sizes. The potential of Fermi teams nevertheless grows: the proportion of items on which at least one agent is within one order of magnitude rises from 31.6 % 31.6% for a single agent to 57.3 % 57.3% at N = 30 N=30 (Section 2.3 ; Appendix Table 4 ), but averaging cannot select that estimate. Continuous estimates lack verifiable intermediate steps, so revision rarely corrects erroneous orders of magnitude. RealFP contains implausible reference values, such as 2 × 10 90 2\times 10^{90} for the number of a person’s ancestors; on 144 of the 529 items (27.2%), the median estimate of the ten models deviates from the reference by more than three orders of magnitude. On the remaining 385 items (open markers in Fig. 1 ; fermi-clean in Table 4 ), all levels are higher, but the pattern is unchanged: Round-1 accuracy rises only from 40.9 % 40.9% for a single agent to 43.6 % 43.6% at N = 30 N=30 and the clipped log-error decreases from 1.52 1.52 to 1.39 1.39 at N = 5 N=5 and 1.37 1.37 at N = 30 N=30 , whereas pass@30 reaches 70.8 % 70.8% . The gap between the potential and the realised performance of Fermi teams therefore does not arise from erroneous labels. 3.3 The Anatomy of the Disjunctive Ceiling While Section 3.2 showed that disjunctive tasks benefit substantially from multi-agent deliberation compared to compensatory tasks, a critical question remains: does disjunctive performance scale monotonically as team size grows? Figure 1 shows that it does not: rather than approaching the pass@ N N curve, realised disjunctive accuracy exhibits its own ceiling. In this section, we analyse the empirical evidence and the mechanisms governing this ceiling. The Early Saturation of Independent Ensembles (Round 1) Round-1 plurality voting yields almost no improvement with team size: from N = 2 N=2 to N = 30 N=30 , accuracy changes by + 0.3 +0.3 points on GSM8K (95% CI [ 0.0 , 0.5 ] [0.0,0.5] ), + 1.0 +1.0 on GSM-Hard, + 1.3 +1.3 on MATH-500, + 0.8 +0.8 on MMLU-Hard, and + 0.8 +0.8 on ARC-Challenge. As Proposition 2.1 predicts, plurality voting converges to the model’s modal answer on each item. We estimated the item-level answer distributions from the Round-1 answers of the N = 30 N=30 teams (up to 90 samples per item) and predicted plurality accuracy at every team size by resampling. Across 650 model–task–team-size configurations, predicted and observed accuracies differ by 0.48 points on average (median 0.18; Pearson r = 0.999 r=0.999 ; Appendix Figure 2 ). The intra-class correlation of correctness ranges from 0.53 to 0.97 (mean 0.78), and the predicted large-team limit exceeds single-agent accuracy by only 0.4–1.2 points per task. Consistent with Eq. 3 , the gain remains within the bound 2 p ¯ ( 1 − p ¯ ) ( 1 − ρ ) 2\bar{p}(1-\bar{p})(1-\rho) for all 65 model–task pairs and uses on average only 11–21% of it (Appendix Figure 3 ). The two models that reason before answering have a lower ρ \rho ( 0.66 0.66 against 0.81 0.81 ), a larger predicted voting gain ( + 3.4 +3.4 against + 0.2 +0.2 points), and observed gains of + 7.2 +7.2 and + 2.8 +2.8 points from one to 30 agents, against − 1.0 -1.0 to + 1.0 +1.0 for the others, consistent with Eq. 3 and with the success of self-consistency under chain-of-thought prompting ( Wang et al., 2023 ) ; the model predicts both regimes (mean error 0.19 and 0.43 points, excluding MATH-500). A higher temperature is no substitute: for qwen2.5-7b, raising T T from 0.2 to 1.0 lowers ρ \rho and raises pass@30 but leaves plurality accuracy unchanged, because the extra samples spread over incorrect answers (Appendix Table 5 ). Additional queries cannot produce a correct plurality on items whose modal answer is wrong; the difference between pass@30 and plurality accuracy at N = 30 N=30 (4.8–20.1 points) is the process loss of plurality voting. The Peak-and-Decay Regime in Deliberated Teams (Round 3) Multi-turn communication nearly doubles accuracy on gsm8k ( 34.1 % → 62.0 % 34.1%\to 62.0% at the peak) and gsmhard ( 18.0 % → 33.5 % 18.0%\to 33.5% ). Deliberated teams do not, however, continue to improve with size (Appendix Table 6 ): accuracy peaks at an intermediate team size and declines by 0.4 0.4 – 1.1 1.1 points towards N = 30 N=30 on the mathematical and knowledge tasks, most clearly on gsm8k ( − 1.0 -1.0 points relative to N = 5 N=5 , 95% CI [ − 1.4 , − 0.7 ] [-1.4,-0.7] ), and is flat on arc. The largest declines occur on multi-step symbolic tasks, consistent with an incorrect intermediate result shared by several peers acting as an attractor during revision. Table 2: Per-Model Scaling and Correlation Breakdown on Disjunctive Tasks (Round 3; accuracy in %, macro-averaged over the five benchmarks). Δ 5 = Acc 5 − Solo \Delta_{5}=\text{Acc}{5}-\text{Solo} and Δ decay = Acc 30 − Acc 5 \Delta{\text{decay}}=\text{Acc}{30}-\text{Acc}{5} , with paired 95% item-bootstrap intervals in brackets; ρ \rho is the intra-class correlation of correctness. † Best five-model pool, selected and evaluated on disjoint halves of the items (mean over 10 splits; brackets give the range across splits); for pools, ρ \rho is the mean correlation between members. Model Solo 𝐍 = The same large language models question is explored in Extracting Arguments, Not Just Classifying Them, which adds a research perspective. as detailed in the full paper on Arxiv The same large language models question is explored in When Can Agents Forget Their Reasoning?..., which adds a research perspective.
Comments (0)
to join the discussion
No comments yet
Be the first to share your thoughts!