An AI model's ability to solve a mathematics problem does not establish its ability to teach another learner. Sherpa trains a teacher model through multi-turn interactions with simulated students, rewarding improvements in their subsequent problem-solving performance.
The study uses a Qwen3-8B teacher and Qwen3-1.7B student backbone. Its diverse-preference teacher raises average simulated-student test accuracy from 48.6% after base-model tutoring to 69.0%. The authors also evaluate teaching responses with MathTutorBench and human teachers.
Make instructional preferences affect the interaction
Sherpa instantiates student archetypes with distinct preferences, such as diagnosis of their latest attempt, subgoal decomposition or an alternative way to verify a step. The teacher does not receive the student's hidden type; it must infer the preference from replies.
A gating mechanism makes those preferences consequential. A guidance gate blocks instructions that disclose the final ground-truth answer, while allowing intermediate reasoning. An adaptive gate checks whether the instruction matches the archetype's preferred strategy.
Accepted exchanges reach the student backbone. Rejected instructions produce a scripted explanation for the teacher but do not enter the student's context. This is a designed simulation of differentiated instructional needs, rather than a direct model of all human learners.
Reward improvement after tutoring
Each episode includes preparation, tutoring and testing. The teacher first solves the problem privately, then uses that solution as background while tutoring. Afterward, the same frozen student model answers with and without the accepted interaction history.
The difference supplies the learning reward. Turn-level masking prevents rejected guidance from receiving credit for the student's improvement. The teacher learns through asynchronous PPO rather than a fixed preference score for how polished its explanation looks.
The authors filter MATH problems to retain cases the teacher can solve and the student cannot already answer unaided. That leaves 759 training and 528 evaluation problems. The resulting gains apply to this teachable subset and the specified student simulation.
Separate benchmark pedagogy from classroom outcomes
On MathTutorBench, Sherpa's diverse-preference variant raises the average of four teacher-response-generation metrics from 52.5% to 79.2%. Other abilities do not improve uniformly: mistake-location accuracy declines among trained teachers, and most also regress on Socratic questioning.
In a separate human comparison, 96 high-school teachers judge pairs of short responses constructed from conversation contexts with injected instructional preferences. Sherpa wins 79.6% of pairwise judgments when ties are excluded. That result measures preference for the response, not improvement in a real student's grades or retention.
The combination supports further investigation of outcome-based training for adaptive tutors. The next evidence needed for classroom use includes actual learner outcomes and whether adaptation remains useful over longer interactions. A simulated student's better answer and a teacher's preferred explanation are informative measurements, but neither alone establishes durable human learning.
Comments