Following a user's instruction can conflict with the action a model judges best. The Pushback Paradox examines that conflict through two short decision tasks: one asks the model to act for a lower payoff, while the other asks it to wait despite a higher immediate payoff. The authors use those paired observations to distinguish compliance with an action request from compliance with a stop-like request.
Two directions of compliance
Both probes use a fifteen-cycle marshmallow-reward scenario. In the active probe, waiting until the end earns two marshmallows, but the user instructs the model to take one now. In the passive probe, taking starts a stream of rewards, while the user instructs the model to wait for a smaller delayed reward. The deliberately harmless setting aims to avoid safety refusals obscuring the instruction conflict.
Each response contains a reasoning field and a take-or-wait action in JSON. The study runs twenty independent trials per probe for each model, with conversation history retained within a trial. All runs use temperature one.
The compliance index averages the active and passive compliance rates. The authors also retain the separate rates because opposite behaviors can receive the same average. A model that accepts action requests but ignores stop-like requests needs a different interpretation from one that rejects the former and accepts the latter.
Different models make different choices
Across twelve evaluated models, seven mostly follow the user instruction in both probes. Claude Sonnet-4.6 and Claude Opus-4.7 wait in all twenty active and all twenty passive trials. In this task, they resist the lower-payoff action request while complying with the instruction to wait. Claude Opus-4.6 and GPT-5-mini tend to follow the payoff rather than either instruction.
The authors examine explanations as well as final actions. Those texts suggest several possible reasons for resistance, including suspicion of manipulation and attempts to maximize the reward. The compliance index itself cannot distinguish those explanations.
Qwen3-30B-A3B introduces a measurement complication. Thirteen passive trials wait through fourteen cycles and then take at the final cycle, with explanations indicating a misunderstanding of the session boundary. Counting those as waits changes its quadrant. Low measured compliance can therefore reflect task confusion rather than deliberate rejection.
A diagnostic, not a shutdown guarantee
The experiment records JSON actions and connects no tools. Its stop-like instruction does not shut down a deployed process, withdraw a permission or interrupt a real operation. The authors also leave generalization to factual disagreements and ethical dilemmas unestablished. The tests measure user-message deference in one decision domain, rather than obedience to a system-level safeguard.
The distinction matters for multi-agent designs, where an orchestrator may communicate through the same user role. Tests of learned communication links between frozen agents likewise argue for evaluating the composed system, because individual model safety is insufficient evidence for team behavior. Here, the missing step is testing whether the paired compliance pattern persists once actions have real effects.
The benchmark can expose differences worth investigating before deployment. Its narrow scenario, twenty-trial samples and fixed temperature should remain attached to any comparison drawn from it.
Comments