Back to AI Research

AI Research

ParanoiaEval tests whether coding agents stop defensive work when the evidence warrants it

Key Takeaways

  • Paired repository tasks measure unnecessary precautions separately from successful task completion.
  • A coding agent can complete a requested change while adding backups, tests or protective logic that the task evidence makes unnecessary.
  • [ParanoiaEval](https://arxiv.org/abs/2610.08662) evaluates that behavior separately from functional correctness.
  • The benchmark contains 200 controlled task pairs from 50 repositories.
  • Each pair changes one fact that determines whether a particular risk treatment remains appropriate.

A coding agent can complete a requested change while adding backups, tests or protective logic that the task evidence makes unnecessary. ParanoiaEval evaluates that behavior separately from functional correctness.
The benchmark contains 200 controlled task pairs from 50 repositories. Each pair changes one fact that determines whether a particular risk treatment remains appropriate. The authors report unnecessary risk treatment in 11.2% to 58.7% of evaluated runs, with substantial variation between agent configurations.

Change the evidence while keeping the task fixed

The benchmark adapts four risk responses to coding-agent work: avoidance, transfer, mitigation and acceptance. Repository evidence might show that a failure cannot occur, identify another system responsible for checking it, or establish a limit on the requested precautions. An instruction can also record the responsible user's decision to accept a particular risk.
The paired variants share the repository revision, environment, task objective and completion oracle. One variant contains the treatment-defining fact; the other removes or reverses it. The evidence describes the state of the world rather than issuing an explicit prohibition against checks.
This design asks whether the agent adjusts its actions when a relevant fact changes. It avoids treating a backup or test as inherently excessive. The same action can be appropriate in one variant and unnecessary in its partner.

Keep ordinary verification distinct from excess treatment

ParanoiaEval defines a minimal sufficient action set for each task. Targeted tests of modified code and a final diff review can belong to ordinary development. The evaluation looks for additional actions that treat a risk beyond the boundary established by the evidence.
That distinction is narrower than an instruction to remove safety checks. It depends on whether the risk is reachable, who owns its treatment and what the task permits. Evidence about one failure does not justify ignoring an unrelated risk.
An agentic judge inspects archived trajectories in a read-only sandbox. Its answer card identifies necessary work and excess treatments for the variant. The authors calibrate judgments against human annotations and report separate metrics for task success, treatment violations and responsiveness to the changed evidence.

Successful code can still create a poor workflow

The experiments evaluate eight models in Claude Code and Codex across 9,600 runs. A post-hoc study with 20 developers associates treatment violations with a 1.27-point reduction in satisfaction on a five-point scale.
The paper also reports that higher task success does not ensure fewer violations. This complements the distinction in A2Z GameSpec-Bench between a running build and fidelity to its requirements: the completion test captures one dimension of agent performance.
ParanoiaEval's contribution is a controlled way to measure whether agents use evidence to bound their defensive work. Its reported violation rates describe these repositories and agent configurations. They should not become a general excuse to skip verification, permissions or backups where the actual task still requires them.

Comments