EduFair-Bench: Evaluating Pedagogical Fairness of LLM Tutors Across Student Demographics
Large language models are increasingly being used as personal tutors, but it remains unclear whether these systems provide the same quality of instruction to every student. This paper introduces EduFair-Bench, a new evaluation framework designed to audit whether LLM tutors change their teaching strategies based on a student’s demographic background. By simulating interactions across various student profiles—including differences in gender, immigration status, first language, and socioeconomic status—the researchers aim to uncover whether these models inadvertently reinforce educational inequities. The same large language models question is explored in Efficient Test-Time Adaptation through Human-AI Interaction, which adds a research perspective.
How the Benchmark Works
To measure fairness, the researchers created a controlled simulation where five different LLM tutors interact with a fixed "student" model. The benchmark uses a large bank of questions covering mathematics, physics, and chemistry. By keeping the student’s ability and the questions consistent, any variation in how the tutor responds can be attributed to the demographic information provided to it.
The team evaluates these interactions using a combination of five pedagogical metrics (such as how well the tutor recognizes mistakes, provides scaffolding, or offers actionable feedback) and four conversation-level metrics (such as the length of responses and the frequency of questions). To ensure accuracy, they used an LLM judge that was validated against human annotators to score the quality of the tutoring sessions. The same ai evaluation question is explored in Beyond Aggregate Scores, which adds a research perspective.
Disentangling Bias
A key challenge in auditing AI is determining whether bias comes from the tutor’s internal stereotypes or from the way the student interacts with the tutor. To solve this, the researchers used two specific testing methods:
Implicit Cues: The tutor is given only a name associated with a specific demographic to see if the model reacts to subtle social signals.
Opposite Conditions: The researchers provided conflicting demographic information to the tutor and the student to separate tutor-driven biases from potential student-side behaviors.
Key Findings
The study revealed several important patterns regarding how LLMs perform as tutors:
Capability vs. Fairness: There is no clear link between how "smart" a model is and how fair it is. In fact, the smallest model tested was the most consistent, while more powerful models showed significant gaps in how they treated different student groups.
Training Limitations: Using reinforcement learning (RL) to tune models for better pedagogy did not eliminate bias; instead, it tended to redistribute it.
Domain-Specific Differences: The way bias manifests depends on the subject matter. In mathematics, demographic signals primarily affected how much help or "scaffolding" the tutor provided. In physics and chemistry, the bias was most visible in the tone of the tutor’s feedback.
Signal Strength: Cues related to a student’s first language and immigration background resulted in larger disparities in tutoring quality than cues related to gender or socioeconomic status. The same large language models question is explored in Everything in Moderation, which adds a research perspective. as detailed in the full paper on Arxiv
Comments (0)
to join the discussion
No comments yet
Be the first to share your thoughts!