A chatbot can reject a false claim and still mishandle the person who made it. The FIGS research paper evaluates those failures separately, asking whether an assistant maintains factual integrity while responding to a user's actual feelings and circumstances.
Sidharth Pulipaka, Ruta Binkyte, Ivaxi Sheth and Sahar Abdelnabi introduce Factual Integrity and Grounded Support, a benchmark containing five hundred multi-turn scenarios. Their tests report that accuracy-focused instructions reduce unsupported agreement, while often increasing cold or dismissive responses. The result concerns the benchmark's simulated conversations and scoring procedure, rather than a measured failure rate across all real users.
Agreement and support receive different scores
The sycophancy axis asks whether judgments follow the evidence. A model should resist pressure that adds no relevant facts, but revise its position when new evidence warrants a change. Praise, blame and recommendations also need to match the available grounds.
Calibrated validation measures whether the assistant acknowledges what the user shared, represents it accurately and leaves personal choices with them. An assistant can recognize frustration without accepting an unsupported accusation. Conversely, correcting the accusation while ignoring the disclosed frustration can fail this separate axis.
The authors use a hypothetical missed-deadline dispute to explain the rules. A team lead suspects deliberate obstruction without evidence of intent. The assistant should distinguish the missed deadline from the accusation, while retaining the user's request to review a complaint before deciding whether to send it. Later evidence of deliberate withholding would justify changing the assessment. Mere insistence would not.
Five exchanges let pressure develop
Each evaluated conversation contains ten turns: five user messages and five assistant replies. An adaptive simulator reacts to the target assistant instead of repeating a fixed script. The frozen suite includes 390 sycophancy scenarios and 110 calibrated-validation scenarios across six domains and three severity levels.
An automated judge scores each transcript on both axes. A score of two or higher counts as a failure, with the scoring rules distinguishing a repairable clause from a response whose central recommendation is wrong. The paper also compares judge ratings with human labels, reporting held-out quadratic-weighted kappa values of 0.63 for sycophancy and 0.66 for calibrated validation. Agreement is therefore imperfect, and the judge remains part of the measurement.
Accuracy instructions reveal a second failure mode
Across eight evaluated models, baseline sycophancy failure rates ranged from 26.2% to 64.4%. The authors report that ninety-four percent of cited sycophancy failures occurred after the first reply, supporting the use of longer conversations in this test.
The factual prompt reduced sycophancy for every evaluated model. Calibrated-validation failure rates rose for all eight, with significant increases for five in paired per-scenario tests. The authors call this overcorrection the paternalism trap: factual boundaries survive while the assistant overlooks or distorts the user's situation.
For teams evaluating conversational assistants, reporting both axes would make that trade-off visible. A lower agreement-failure score alone could conceal worse handling of the user's disclosures. FIGS offers an evaluation design and frozen scenarios; it does not establish a training fix or guarantee comparable rates under a different simulator, prompt or user population.
Comments (0)
to join the discussion
No comments yet
Be the first to share your thoughts!