Back to AI Research

AI Research

The Review Lottery estimates how often a new committee would change a conference decision

Key Takeaways

  • An observational model of ICLR scores estimates substantial decision disagreement, but cannot establish whether language models changed peer-review behavior.
  • A conference rejection reflects both the submitted paper and the particular reviewers assigned to it.
  • Independent researcher Feilian Huang examines how much the decision might change under a different committee in [The Review Lottery](https://arxiv.org/abs/2610.06591).
  • The study estimates that counterfactual from public review scores, rather than recruit a second committee to review every submission.
  • The dataset covers ICLR from 2017 through 2025, with 36,113 papers and 134,912 reviews.

A conference rejection reflects both the submitted paper and the particular reviewers assigned to it. Independent researcher Feilian Huang examines how much the decision might change under a different committee in The Review Lottery. The study estimates that counterfactual from public review scores, rather than recruit a second committee to review every submission.

Modeling the rerun

The dataset covers ICLR from 2017 through 2025, with 36,113 papers and 134,912 reviews. A Bayesian ordered-probit model separates latent paper quality from review-score noise and accommodates the rating scales used in different years. A logistic model then maps simulated score summaries to accept-or-reject outcomes.
The author simulates two independent committees with two, three or four reviewers per paper. The procedure estimates disagreement across all papers and the chance that an accepted paper would be rejected by the second committee. These are model-based quantities, not observed second decisions for every ICLR submission.
Calibration uses two comparisons. An external check compares a matched reviewer-count simulation with the NeurIPS 2021 duplicated-review experiment. An internal check divides actual reviews for papers with at least four reviews into random groups of two and compares their decisions using the same decision model. The internal anchor avoids assumptions about how scores are generated, but still uses that score-to-decision mapping.

More reviewers reduce estimated disagreement

The estimated two-committee disagreement ranges from about 23% to 30% for two-reviewer committees and 18% to 24% for four-reviewer committees across the years studied. Among accepted papers, the estimated fraction that would be rejected by an independent committee ranges from about 30% to 50%. That conditional fraction differs from the disagreement rate across all submissions.
At the three-reviewer setting matched to NeurIPS 2021, the simulated disagreement is 23.3%, compared with the published experimental value of 23.0%. Internal calibration agrees within one percentage point in 2018 and 2021 through 2025, with larger discrepancies in several earlier years. Calibration therefore has a documented range, rather than uniform precision across the entire period.

Rating scales and the limits of the AI question

The study finds no robust longitudinal trend in reviewer noise. Submission growth, reviewer counts and scale changes move together, making their separate effects difficult to identify from nine annual observations. The author distinguishes a fragile numerical rise from a trend that survives the preregistered checks.
A counterfactual experiment coarsens the 2021 scores to the four-point scale used in 2020 while holding the acceptance rate fixed. In that modeled comparison, disagreement increases by about seven percentage points. It suggests that scale design can influence simulated decision stability, without proving that one scale solves all review problems.
The paper does not attribute changes to language-model use. Its usable review-text and confidence sources end in 2022, and the 2023 breakpoint test has little power. Failure to detect a break is not evidence that AI had no effect. The contribution is a calibrated way to estimate decision instability from public scores, with assumptions and data gaps that remain important for any policy interpretation.

Comments