Back to AI Research

AI Research

Bayesian council method weights AI votes by measured reliability, including negative weights

Key Takeaways

  • BDA aggregates structured proposals, challenges and concessions without extra model calls, while requiring labeled calibration data and stable agent behavior for its defense to.
  • BDA aggregates structured proposals, challenges and concessions without extra model calls, while requiring labeled calibration data and stable agent behavior for its defense to work.
  • Three AI models agreeing on an answer does not make that answer certain.
  • A new paper examines how a multi-model council can return a confidence estimate that reflects correctness, while resisting members that persist in giving wrong answers.
  • In [Counting Moves, Weighing Voices](https://arxiv.org/abs/2610.02005), the authors propose Bayesian Dialectical Argumentation, or BDA.

Three AI models agreeing on an answer does not make that answer certain. A new paper examines how a multi-model council can return a confidence estimate that reflects correctness, while resisting members that persist in giving wrong answers.
In Counting Moves, Weighing Voices, the authors propose Bayesian Dialectical Argumentation, or BDA. It converts structured dialogue moves into evidence weighted by each agent's measured reliability. Their reported gains concern the tested councils and adversarial protocols; the method's calibration and robustness have explicit assumptions.

Counting commitments instead of judging prose

The council protocol produces proposals, challenges and concessions. BDA records which agent endorsed or rejected which candidate answer. A proposal endorses its answer, a concession endorses the proposal it accepts, and a challenge rejects its target's answer.
The aggregator discards free-text arguments and self-reported confidence. That design prevents a more persuasive explanation or inflated confidence number from changing the posterior on its own. Changing the answer endorsed by a move still changes the evidence, so the method does not eliminate every route for an adversary to influence a council.
BDA computes answer probabilities from these counts using an annotator model. In the binary, single-endorsement case, the decision reduces to a weighted majority with reliability log-odds as weights. An agent whose estimated reliability falls below chance can receive a negative weight: its endorsement counts against the candidate it names.
The per-agent version learns its priors from labeled calibration folds. This is supervised aggregation, and the paper gives supervised baselines the same label budget. The aggregation itself needs no additional LLM calls; generating the council's deliberation remains separate work with its own cost.

Calibration has to be earned

The authors report a clean binary expected calibration error of 0.016 for per-agent BDA and an accuracy improvement of 1.7 percentage points over plurality voting. They also acknowledge that a recalibrated re-derivation baseline matches its calibration and beats its accuracy. The results therefore do not support calling BDA the best method on every metric.
For more than two candidate answers, the paper extends the model to per-agent confusion matrices. Challenges can produce correlated pseudo-observations in that setting, so the authors temper the composite likelihood and choose its settings through cross-validation. Calibration there is an empirical correction rather than a universal guarantee from Bayesian notation.
The evaluation uses binary claim-verification datasets, three-choice PubMedQA questions and a four-choice MMLU subset. The primary council combines Gemma-3-27B-IT, Phi-4 and GPT-4.1-nano, with three rounds of deliberation. Those details define the environment in which the authors measured the method.

Persistent adversaries are a specific threat model

Reliability weighting needs an unreliable member to be identifiable during calibration and remain associated with a stable seat. A member that behaves well during calibration but becomes hostile later inherits the system's trust until the profile is refitted. Rotation and changing behavior limit what a stored per-agent reliability can detect.
Under the tested live adversarial coalitions, the confusion-model variant leads binary and PubMedQA accuracy, while logistic stacking leads the positional MMLU setting. This split is useful evidence against an overly broad robustness claim.
BDA gives council designers a way to separate agreement from confidence and measure whose commitments deserve weight. Its strongest practical lesson is to audit calibration and attack assumptions together, instead of treating additional voices as an automatic reliability improvement.

Comments