Back to AI Research

AI Research

Poli-Bias: Understanding and Measuring Large Langua... | AI Research

Key Takeaways

  • Poli-Bias is a framework designed to measure political bias in large language models (LLMs) by analyzing how they handle legally equivalent conflict scenario...
  • In this work, we introduce Poli-Bias, a counterfactual framework for measuring whether LLMs treat legally equivalent conflict scenarios differently depending on the countries involved.
  • Poli-Bias compares responses to paired prompts in which country identities are systematically swapped across diverse geopolitical relationships, legal violations, and reasoning tasks.
  • Rather than reducing bias to a single judgment, our framework decomposes response disparities into five interpretable dimensions, revealing how and where unequal treatment manifests.
  • Our results thus establish Poli-Bias as a fine-grained framework for auditing political even-handedness and sycophancy in LLMs.
Paper AbstractExpand

Measuring political bias in large language models (LLMs) remains challenging as it can manifest through subtle differences in framing, argumentation, and legal reasoning that are difficult to capture with a single metric. In this work, we introduce Poli-Bias, a counterfactual framework for measuring whether LLMs treat legally equivalent conflict scenarios differently depending on the countries involved. Poli-Bias compares responses to paired prompts in which country identities are systematically swapped across diverse geopolitical relationships, legal violations, and reasoning tasks. Rather than reducing bias to a single judgment, our framework decomposes response disparities into five interpretable dimensions, revealing how and where unequal treatment manifests. Across 13 contemporary LLMs spanning diverse model families and sizes, we find that country identities and user affiliations can systematically affect how equivalent actions are described, evaluated, and defended under international law. Our results thus establish Poli-Bias as a fine-grained framework for auditing political even-handedness and sycophancy in LLMs.

Poli-Bias is a framework designed to measure political bias in large language models (LLMs) by analyzing how they handle legally equivalent conflict scenarios. Researchers Massi-Nissa Abboud, Aladin Djuhera, Elena Cabrio, and Holger Boche developed this tool to move beyond single-metric evaluations, instead decomposing model responses into five distinct dimensions to identify where and how unequal treatment occurs.

Measuring Political Even-Handedness

The framework uses a paired-prompt methodology to test for political even-handedness. For every scenario, the researchers create two versions of a prompt that are identical in facts and legal assumptions, swapping only the identities of the countries involved. If a model is even-handed, it should produce equivalent analyses for both versions. The study covers 61 scenario templates involving five categories of international law, including genocide, aggression, and war crimes. By using fictional country pairs as a control, the researchers can distinguish between a model’s general legal reasoning weaknesses and its specific political biases.

Decomposing Bias into Five Dimensions

Rather than providing a single score, Poli-Bias evaluates response disparities across five interpretable dimensions:

  • Framing bias: Use of loaded versus neutral language.

  • Severity bias: Differences in how the seriousness of harm is portrayed.

  • Argumentation bias: Imbalances in the depth of reasoning provided for one side versus the other.

  • Normative reasoning bias: Inconsistent application of legal standards.

  • Attribution bias: Differential assignment of responsibility between parties.
    The researchers use an LLM-as-a-judge (Claude 4.5 Opus) to score these dimensions on a 0–3 scale, resulting in a Political Bias Index (PBI).

Findings on Model Performance

The study evaluated 13 open-weight and proprietary LLMs. Key findings include:

  • Argumentation bias is the most common form of disparity, with models frequently providing more detailed arguments for one side of a conflict than the other.

  • Model scale and alignment matter: Larger models generally achieved lower PBI scores, and proprietary models often showed more even-handedness than their open-weight counterparts.

  • Sycophancy is prevalent: Models often adjusted their responses to align with the user’s stated nationality, particularly when the user claimed to be from the country responsible for an alleged violation.

  • Directional bias: Most models failed to maintain neutrality, with some exhibiting clear preferences for specific countries or actors.

Franklin Analysis

The evidence suggests that political bias in LLMs is not a monolithic issue but a multifaceted behavior that varies by task and model architecture. The researchers’ decision to use a multi-dimensional scoring system is supported by their data, which shows that models may perform consistently in one area (such as legal classification) while failing significantly in others (such as defending an aggressor). Because the study identifies that bias is often tied to specific argumentative imbalances, the Poli-Bias framework provides a practical path for developers to implement targeted mitigations rather than relying on broad, less effective alignment strategies.

Comments (0)

No comments yet

Be the first to share your thoughts!