Calling a reasoning prompt unlearnable makes a stronger claim than observing that a model improves slowly on it. A new reanalysis finds that some prompts previously given that label do improve during reinforcement learning, while the set selected for study changes substantially with the measurement procedure.
Chandak Chakma and colleagues examine that distinction in Unlearnable, or Unmeasured? On the Reliability of Difficulty Labels in RLVR. Their findings concern reinforcement learning with verifiable rewards, where automatically checked answers provide the learning signal.
Slow learning survives the reanalysis
The primary experiments train Qwen2.5-0.5B with GRPO on 1,023 MATH prompts, across five independent seeds. The researchers retain the earlier study's easy, learnable and unlearnable cohorts and its 0.1 difficulty threshold.
Prompts labeled unlearnable improve at roughly one third of the learnable group's rate in this configuration. They remain substantially worse, but their learning curve is not flat. The authors consequently describe persistent differences in trainability, rather than proof that those prompts cannot improve.
They also explain why training slopes are more informative here than a final cohort average. Under the sampling procedure, prompts stop contributing when a sampled group's rewards have zero variance. The prompts represented near the end need not be the same ones represented earlier.
A difficulty label depends on sampling
Estimated difficulty comes from a finite number of sampled answers. A prompt near the threshold can fall on either side when the same model is evaluated again. Combining decisions across runs introduces another choice: whether any run or every run must meet an exclusion condition.
Holding the same 292 candidate hard prompts fixed, the published five-seed construction retains 74 prompts. An alternative aggregation retains 228. Individual seeds retain between 145 and 150. The aggregation rule therefore changes the selected set more than switching between those individual training seeds.
The paper models reproducibility as a sampling-budget problem. In its setting, predicted agreement between independently constructed difficulty sets is about 0.63 with eight rollouts per prompt and 0.73 with 16. Reaching 0.90 agreement requires a projected 242 rollouts. That projection exceeds the 128-rollout fitting range and is specific to the model, dataset and threshold, not a universal recommended budget.
Using one pooled 128-response estimate also yields more reproducible membership than requiring agreement across four separately thresholded blocks of 32, at the same total sampling cost.
Measurement does not settle the mechanism
The authors revisit gradient-similarity evidence used to explain the slow-learning group. Hard prompts yield fewer successful rollouts, so their gradients are estimated from fewer samples. Matching the correct-rollout count reduces the reported similarity ratio from 2.327 to 1.539, but leaves a substantial difference. The paper weakens part of that explanation without resolving the underlying cause.
Its reproduction is not identical to the earlier training setup. Primary training and evaluation cap generation at 1,024 tokens, versus the original scripts' 5,120. A longer evaluation-cap control supports the instability finding, but does not recreate training under the longer cap.
The practical conclusion is to measure the reliability of a difficulty-defined set before using it to filter training data or explain model limits. A stable benchmark average does not necessarily mean that the same individual prompts receive the same label on another run.
Comments (0)
to join the discussion
No comments yet
Be the first to share your thoughts!