Back to AI Research

AI Research

Agent Error Dataset links failed runs to diagnoses and tested corrections

Key Takeaways

  • The Agent Error Dataset collects 50,228 error-diagnosis pairs with trace provenance.
  • Matched replay tests support some proposed corrections, while training results distinguish impr
  • Matched replay tests support some proposed corrections, while training results distinguish improved internal label agreement from uneven external transfer.
  • A failed agent run contains more useful evidence than a final pass-or-fail score.
  • The [Agent Error Dataset paper](https://arxiv.org/abs/2609.40111) preserves what an agent observed, which actions it took and how the environment responded, then links that trace to a diagnosis and proposed correction.

A failed agent run contains more useful evidence than a final pass-or-fail score. The Agent Error Dataset paper preserves what an agent observed, which actions it took and how the environment responded, then links that trace to a diagnosis and proposed correction. Its authors use the resulting records to study failure analysis and several forms of post-training, with separate requirements for each use.

Count diagnoses without confusing them with independent failures

The collection contains 50,228 error-diagnosis pairs from 9,961 source tasks, spanning 33 environments, nineteen harness families and 23 policy models. A pair joins a failed execution to one recorded diagnosis. Multiple diagnoses of the same run therefore add pairs without adding independent runs.
Records retain source traces, execution metadata, review decisions and any available replay outcomes. That provenance lets researchers inspect a diagnosis or generate another one without repeating the original rollout. The authors also distinguish agent mistakes from infrastructure or grader faults: an unsuccessful outcome alone does not identify an avoidable decision.
Their Agentic Error-to-Training pipeline collects natural failures, proposes a diagnosis and correction, checks those claims against visible evidence, runs a replay where supported, and creates training views. Collection membership does not mean a record contains a verified repair. Diagnosis training can use a trace-supported label without replay; recovery targets need an executed, passing continuation.

Test a correction against a retry from the same checkpoint

In the replay-supported cohort, the researchers compare a proposed correction with a fresh retry of the original action from the same checkpoint. They match policy, harness, budget and verifier settings. Across 3,062 matched replay pairs, first-proposal corrections increase verifier pass rates from 18.4% to 51.1%, a difference of 32.7 percentage points.
That experiment tests whether a correction helps under its execution conditions. It does not prove that the diagnosis names a unique root cause. Both the correction and original-action outcome remain in the record, including negative or zero contrasts.
The paper separates this replay cohort from its diagnosis-training release. That separation prevents a successful correction test from being mistaken for evidence that a trained model can locate the error on its own.

Separate internal learning from external transfer

Full-diagnosis fine-tuning on 1,656 source tasks increases Qwen3-8B's exact-step agreement with internal teacher labels from 47.2% to 63.6%, averaged over three seeds on a 943-case holdout. The score measures agreement with recorded annotations, including their conventions and defects. It is not a universal ranking of diagnostic ability.
External results are mixed. The paper reports that scores depend on the benchmark, answer format and overlap with training tasks. Mean TrajErrBench accuracy remains below the base model, and some Who&When comparisons lose accuracy under different evaluation protocols.
For actor training, an action-only repair recipe scores 6.67 percentage points above success-only training on WebShop-lite in a single-seed comparison. The authors caution that those recipes differ in task pools, exposure and optimizer updates, so the difference cannot be attributed to repair supervision alone. The dataset supplies inspectable failure experience; deciding which records improve a particular agent still requires objective-specific tests.

Comments (0)

No comments yet

Be the first to share your thoughts!