Back to AI Research

AI Research

PEPO uses nearby token uncertainty to improve reinforcement-learning credit assignment

Key Takeaways

  • Researchers propose local entropy weighting for language-model reinforcement learning.
  • Mathematical reasoning experiments favor PEPO over uniform and batch-level weighting, while i
  • Mathematical reasoning experiments favor PEPO over uniform and batch-level weighting, while its theoretical position-invariance result remains conditional on smooth trends and a small window.
  • Rewarding a language model for a correct answer does not tell the training algorithm which words made the answer correct.
  • The proposal targets reinforcement learning with verifiable rewards, where a checker can score an answer.

Rewarding a language model for a correct answer does not tell the training algorithm which words made the answer correct. Yun Kim and Nojun Kwak address that problem in their paper on Proximal Entropy Policy Optimization, or PEPO, by changing how token-level uncertainty is compared during reinforcement learning.
The proposal targets reinforcement learning with verifiable rewards, where a checker can score an answer. It does not replace the checker or introduce a new reasoning model. It changes the distribution of credit within a generated response.

Why a global uncertainty ranking can mislead

Group Relative Policy Optimization, or GRPO, gives every token in a rollout the same sequence-level advantage. Entropy-based alternatives try to focus updates on uncertain tokens, which may correspond to important choices in a reasoning path.
The paper identifies two confounders: harder prompts raise uncertainty across a response, and early positions tend to be more uncertain than later ones. A token can therefore rank highly across a training batch because of the question it belongs to or where it appears, rather than because it marks a consequential decision.
The researchers illustrate the difficulty effect using Llama-3.2-3B-Instruct rollouts grouped by pass rate. Their reported global entropy differs across hard, moderate, and easy groups, while their local measure stays nearly constant. That observation motivates a relative comparison rather than treating raw entropy as direct causal evidence of importance.

Weight successful responses locally

PEPO compares a token's entropy with a centered window of nearby tokens. A softmax gives greater relative weight to uncertainty that stands out locally. The weights are then normalized so the total advantage magnitude matches standard GRPO.
The implementation applies this weighting to positive-reward rollouts and leaves failed rollouts unweighted. The authors explain that emphasizing negative updates at uncertain positions could discourage exploration. Their experiments use a window of 101 tokens.
The mathematical guarantee needs care: a constant shift in sequence entropy leaves the local measure unchanged. Robustness to positional trends is approximate, under a smooth monotonic trend and a sufficiently small window. It is not a proof that entropy reveals every token's true contribution to an answer.

What the reasoning experiments show

The study evaluates Qwen3-1.7B, Qwen3-4B, and Llama-3.2-3B-Instruct on MATH500, AMC, and AIME 2024 and 2025. Its main table reports means and sample standard deviations over three runs, using avg@1 for MATH500 and avg@16 for the other benchmarks.
PEPO has the highest mean benchmark score among the compared methods for each tested model. Replacing global entropy with proximal entropy inside the separate 80/20 method also improves its mean scores, which helps isolate the local comparison's contribution from PEPO's other choices.
The captured paper additionally reports transfer to single-stream reinforcement learning, where each prompt produces one rollout. These are mathematical reasoning results for the listed models and training setups, not established gains for arbitrary chat, coding, or multimodal tasks. The practical takeaway is a specific training hypothesis worth reproducing: local uncertainty may allocate credit more usefully than a batch-wide ranking.

Comments (0)

No comments yet

Be the first to share your thoughts!