Back to AI Research

AI Research

EduFair-Bench: Evaluating Pedagogical Fairness of L... | AI Research

Key Takeaways

  • EduFair-Bench: Evaluating Pedagogical Fairness of LLM Tutors Across Student Demographics Large language models are increasingly being used as personal tutors...
  • Large language models (LLMs) are increasingly deployed as tutors, but it is unclear whether they support all students equally well.
  • We introduce \textbf{EduFair-Bench}, a benchmark for auditing the pedagogical fairness of LLM tutors---whether tutoring quality varies systematically with student demographics.
  • Tutoring quality is scored on five turn-level pedagogical metrics and four conversation-level dimensions, using an LLM judge validated against three-annotator consensus on 180 tutor turns.
  • Bias is measured via paired Wilcoxon signed-rank tests and bootstrap effect-size confidence intervals.
Paper AbstractExpand

Large language models (LLMs) are increasingly deployed as tutors, but it is unclear whether they support all students equally well. We introduce \textbf{EduFair-Bench}, a benchmark for auditing the pedagogical fairness of LLM tutors---whether tutoring quality varies systematically with student demographics. EduFair-Bench pairs a multi-domain question bank (mathematics, physics, chemistry) with a controlled simulation in which a fixed LLM student interacts with each tutor across nine demographic levels spanning four dimensions: gender, immigration background, first language, and socioeconomic status (SES). Tutoring quality is scored on five turn-level pedagogical metrics and four conversation-level dimensions, using an LLM judge validated against three-annotator consensus on 180 tutor turns. Bias is measured via paired Wilcoxon signed-rank tests and bootstrap effect-size confidence intervals. Two ablations (demographic cues conveyed through names; conflicting demographic information between tutor and student) disentangle tutor-driven from student-driven bias. Across five tutors, we find that model capability and demographic fairness are largely orthogonal: the smallest model is the most consistent while the four more capable tutors all exhibit wide demographic gaps with no clear capability-to-fairness ordering, pedagogy-specific RL training redistributes rather than removes bias, and language- and immigration-related cues produce larger gaps than gender- and SES-related cues.

EduFair-Bench: Evaluating Pedagogical Fairness of LLM Tutors Across Student Demographics
Large language models are increasingly being used as personal tutors, but it remains unclear whether these systems provide the same quality of instruction to every student. This paper introduces EduFair-Bench, a new evaluation framework designed to audit whether LLM tutors change their teaching strategies based on a student’s demographic background. By simulating interactions across various student profiles—including differences in gender, immigration status, first language, and socioeconomic status—the researchers aim to uncover whether these models inadvertently reinforce educational inequities. The same large language models question is explored in Efficient Test-Time Adaptation through Human-AI Interaction, which adds a research perspective.

How the Benchmark Works

To measure fairness, the researchers created a controlled simulation where five different LLM tutors interact with a fixed "student" model. The benchmark uses a large bank of questions covering mathematics, physics, and chemistry. By keeping the student’s ability and the questions consistent, any variation in how the tutor responds can be attributed to the demographic information provided to it.
The team evaluates these interactions using a combination of five pedagogical metrics (such as how well the tutor recognizes mistakes, provides scaffolding, or offers actionable feedback) and four conversation-level metrics (such as the length of responses and the frequency of questions). To ensure accuracy, they used an LLM judge that was validated against human annotators to score the quality of the tutoring sessions. The same ai evaluation question is explored in Beyond Aggregate Scores, which adds a research perspective.

Disentangling Bias

A key challenge in auditing AI is determining whether bias comes from the tutor’s internal stereotypes or from the way the student interacts with the tutor. To solve this, the researchers used two specific testing methods:

  • Implicit Cues: The tutor is given only a name associated with a specific demographic to see if the model reacts to subtle social signals.

  • Opposite Conditions: The researchers provided conflicting demographic information to the tutor and the student to separate tutor-driven biases from potential student-side behaviors.

Key Findings

The study revealed several important patterns regarding how LLMs perform as tutors:

  • Capability vs. Fairness: There is no clear link between how "smart" a model is and how fair it is. In fact, the smallest model tested was the most consistent, while more powerful models showed significant gaps in how they treated different student groups.

  • Training Limitations: Using reinforcement learning (RL) to tune models for better pedagogy did not eliminate bias; instead, it tended to redistribute it.

  • Domain-Specific Differences: The way bias manifests depends on the subject matter. In mathematics, demographic signals primarily affected how much help or "scaffolding" the tutor provided. In physics and chemistry, the bias was most visible in the tone of the tutor’s feedback.

  • Signal Strength: Cues related to a student’s first language and immigration background resulted in larger disparities in tutoring quality than cues related to gender or socioeconomic status. The same large language models question is explored in Everything in Moderation, which adds a research perspective. as detailed in the full paper on Arxiv

Comments (0)

No comments yet

Be the first to share your thoughts!