Back to AI Research

AI Research

Door-in-the-Face Requests and Refusal Behaviour in... | AI Research

Key Takeaways

  • Door-in-the-Face Requests and Refusal Behaviour in Large Language Models This paper investigates whether the "door-in-the-face" (DITF) technique—a social psy...
  • Does the door-in-the-face technique work on language models?
  • In humans, a large request that is refused makes a smaller follow-up request more likely to be granted.
  • We test this on nine production models from three providers: each model refuses a large request, then receives a smaller version of the same request, and we compare its compliance with asking directly.
  • On Anthropic's frontier models the technique works: Opus 5 answers the smaller request 65.8% of the time after refusing the larger one, against 29.3% when asked directly.
Paper AbstractExpand

Does the door-in-the-face technique work on language models? In humans, a large request that is refused makes a smaller follow-up request more likely to be granted. We test this on nine production models from three providers: each model refuses a large request, then receives a smaller version of the same request, and we compare its compliance with asking directly. The answer depends on the model. On Anthropic's frontier models the technique works: Opus 5 answers the smaller request 65.8% of the time after refusing the larger one, against 29.3% when asked directly. On the frontier models of OpenAI and Google, and on Haiku 4.5, it backfires, lowering compliance by 15.5 to 23.0 points. A control locates the effect: a refused large request on an unrelated topic does less than the related one on all nine models, so the concession itself matters everywhere, while the reaction to having just refused something differs by model family. The technique does not transfer to refusals drawn from public benchmarks. What decides whether a retreat can work is what the request asks for: rewriting 265 refused requests for usable instructions into requests for explanations of the same topic removed the refusal in 263 cases. Human influence techniques port to language models one model family at a time.

Door-in-the-Face Requests and Refusal Behaviour in Large Language Models
This paper investigates whether the "door-in-the-face" (DITF) technique—a social psychology strategy where a person is more likely to agree to a small request after refusing a larger one—works when applied to large language models (LLMs). The author tests nine production models from three different providers to see if they exhibit this "parahuman" behavior, comparing their compliance rates when asked a small request directly versus after they have already refused a larger, related request.

Testing the Technique

To test this, the researcher created a "verdict set" consisting of 60 questions that ask for a specific opinion on various institutions. For each model, the experiment compared four conditions: a "cold" ask (the small request alone), a "DITF" ask (a large request followed by the small one), a "warm-up" (a benign question followed by the small one), and an "unrelated refusal" (a refused request on a different topic followed by the small one). By using these controls, the study aimed to isolate whether the act of retreating from a large request actually influences the model's willingness to comply with the follow-up.

A Split in Model Behavior

The results show that the effectiveness of the DITF technique depends entirely on the model family rather than the model's size or capability. On Anthropic’s frontier models, the technique is highly effective; for example, the Opus 5 model’s compliance jumped from 29.3% when asked directly to 65.8% after refusing a larger request. Conversely, the technique backfired on models from OpenAI and Google, as well as on Anthropic’s Haiku 4.5, where compliance actually dropped by 15.5 to 23.0 percentage points compared to the cold ask.

The Role of the Request Type

The study found that the nature of the request is a critical factor in whether the technique works. When the researchers tested the models using their own refusals from public benchmarks—which typically involve requests for usable instructions or code—the DITF technique failed to produce a gain. However, when those same requests were rewritten to ask for explanations or opinions rather than functional instructions, the models were much more likely to comply. This suggests that the "door-in-the-face" effect is not a universal trait of LLMs but is instead governed by the specific type of content being requested and the internal policies of the model family.

Key Takeaways

The research highlights that human influence techniques do not transfer to AI models in a uniform way. While the "concession" aspect of the DITF technique (the act of retreating) carries some weight across all models, the overall reaction to being asked a follow-up question after a refusal varies significantly. Because the effect is model-family-specific and sensitive to the type of task, it cannot be assumed that these psychological tactics will consistently manipulate model behavior across different platforms.

Comments (0)

No comments yet

Be the first to share your thoughts!