Back to AI Research

AI Research

StudentBench: AI and human tutoring yield equivalen... | AI Research

Key Takeaways

  • StudentBench: AI and human tutoring yield equivalent GRE learning gains This research introduces StudentBench, a platform and evaluation suite designed to de...
  • Artificial intelligence offers an unprecedented opportunity to augment human capabilities, yet progress at the frontier has focused primarily on advancing model capabilities.
  • Using StudentBench, we measured learning gains on Quantitative and Verbal GRE questions across 2,383 human participants receiving AI tutoring, human tutoring, or no tutoring.
  • In a second study, expert human tutors compared LLM-generated lesson plans and practice problems through 2,028 pairwise rubric evaluations.
  • Together, the two studies clearly separate AI tutors across: (1) lesson planning, (2) practice-problem creation, (3) conversational pedagogy, (4) cost, and (5) engagement.
Paper AbstractExpand

Artificial intelligence offers an unprecedented opportunity to augment human capabilities, yet progress at the frontier has focused primarily on advancing model capabilities. We introduce StudentBench, a suite of AI teaching evaluations and a public platform that enables large-scale data collection with over 175,000 student-AI messages to study whether large language models (LLMs) produce learning gains equivalent to human tutoring. Using StudentBench, we measured learning gains on Quantitative and Verbal GRE questions across 2,383 human participants receiving AI tutoring, human tutoring, or no tutoring. We establish that AI tutoring is statistically equivalent to expert human tutoring for GRE learning gains (p = .015), and in five of the seven GRE domains, the best performing AI tutor surpassed the human tutor, on average. In a second study, expert human tutors compared LLM-generated lesson plans and practice problems through 2,028 pairwise rubric evaluations. Together, the two studies clearly separate AI tutors across: (1) lesson planning, (2) practice-problem creation, (3) conversational pedagogy, (4) cost, and (5) engagement. Surprisingly, one AI tutor achieved learning gains equivalent to human tutoring (p = .044) at 918 times lower cost (USD 0.0052 for AI versus USD 4.81 for human, per percentage point gained). For Quantitative GRE sessions, faster AI replies correlated with more student messages, more messages with more correct practice, and more correct practice with larger learning gains (all p < .002). The StudentBench platform is freely available at this https URL .

StudentBench: AI and human tutoring yield equivalent GRE learning gains
This research introduces StudentBench, a platform and evaluation suite designed to determine if large language models (LLMs) can provide tutoring that is as effective as human instruction. By analyzing over 175,000 student-AI interactions, the study compares the learning outcomes of students receiving AI tutoring, human tutoring, or no tutoring at all. The goal is to understand if AI can democratize access to high-quality, one-on-one education by providing effective, low-cost support for students preparing for the GRE. The google story also surfaces in Google’s Gemini AI Accessed Three Outside..., adding another angle.

How the study works

The researchers conducted a large-scale experiment with 2,383 participants across Quantitative and Verbal GRE sections. Students were randomly assigned to work with either an expert human tutor or one of 13 different AI tutors, while a control group watched unrelated educational videos. To ensure the results were accurate, the study used pre-tests and post-tests created by former GRE exam writers to prevent models from having seen the questions during their training. The AI tutors were given minimal software scaffolding, allowing the researchers to measure the raw teaching capabilities of the models rather than the effectiveness of the interface design.

Key findings on learning and cost

The study found that AI tutoring is statistically equivalent to expert human tutoring in terms of student learning gains. In five of the seven GRE domains tested, the best-performing AI tutors actually surpassed the average performance of human tutors. Beyond effectiveness, the research highlighted a massive disparity in cost-efficiency. One AI tutor achieved learning gains equivalent to a human tutor at 918 times lower cost, with an average inference cost of just $0.067 for an entire one-hour session. This suggests that AI has the potential to make individualized instruction significantly more affordable. The google story also surfaces in Google launches Gemini 3.8 Flash and..., adding another angle.

The impact of speed and engagement

The researchers also examined how the mechanics of AI interaction influence learning. In Quantitative GRE sessions, they discovered that faster AI response times were strongly linked to higher student engagement. When AI tutors replied more quickly, students sent more messages, engaged in more practice, and ultimately achieved larger learning gains. Additionally, the study used expert human tutors to evaluate AI-generated lesson plans and practice problems, creating a leaderboard that helps distinguish which models are best suited for specific teaching tasks, such as lesson planning versus conversational pedagogy.

Important considerations

While the results demonstrate that AI is capable of human-level tutoring, the researchers note that these findings may represent a lower bound of AI potential. As models continue to improve and software scaffolding becomes more sophisticated, AI tutoring performance is expected to increase further. The study also emphasizes that pedagogical styles vary by model family, suggesting that the training practices of different AI companies significantly impact how a model teaches. The StudentBench platform and the data collected are available to the public to support ongoing research into how these technologies can best augment human learning. The google story also surfaces in Google opens early access to AI..., adding another angle. as detailed in the full paper on Arxiv

Comments (0)

No comments yet

Be the first to share your thoughts!