Back to AI Research

AI Research

LLM-SoccerArena: Benchmarking LLMs on Real-World Pr... | AI Research

Key Takeaways

  • LLM-SoccerArena: Benchmarking LLMs on Real-World Predictions in Sports Large language models are increasingly used to help make decisions about the future, b...
  • Large language models (LLMs) increasingly support decisions about uncertain future events, yet evaluating their ability to forecast real-world outcomes remains difficult.
  • In particular, existing benchmarks are typically static and retrospective, and therefore cannot test how information is synthesized by LLMs to predict future events under uncertainty.
  • We introduce LLM-SoccerArena ( this https URL ), a prospective live benchmark that evaluates how well LLMs forecast real-world sports events before the outcomes are known.
  • LLM-SoccerArena provides (1) a prospective live benchmark protocol, (2) a public open-source platform, and (3) a factorial benchmark design together with tournament-related questions (e.g., which team will win).
Paper AbstractExpand

Large language models (LLMs) increasingly support decisions about uncertain future events, yet evaluating their ability to forecast real-world outcomes remains difficult. In particular, existing benchmarks are typically static and retrospective, and therefore cannot test how information is synthesized by LLMs to predict future events under uncertainty. We introduce LLM-SoccerArena ( this https URL ), a prospective live benchmark that evaluates how well LLMs forecast real-world sports events before the outcomes are known. LLM-SoccerArena provides (1) a prospective live benchmark protocol, (2) a public open-source platform, and (3) a factorial benchmark design together with tournament-related questions (e.g., which team will win). LLM-SoccerArena automatically records timestamped, schema-validated forecasts of unresolved events, together with prompts, model versions, tool traces, and costs. The factorial design varies along four dimensions: (1) model version (e.g., GPT-5.5, Claude Opus 4.8); (2) information access; (3) prompting strategy, and (4) forecast horizon. We demonstrate LLM-SoccerArena through a large-scale evaluation of the 2026 FIFA World Cup, in which seven LLMs generated forecasts for all 104 matches and 15 tournament-related questions. We provide a detailed analysis of model performance across information access, prompting strategy, and forecast horizon. As a result, LLM-SoccerArena provides new evidence about the forecasting performance of state-of-the-art LLMs. For example, LLMs with web access outperform those without, but only by a small margin (i.e., a 0.023 improvement in Brier score). Overall, LLM-SoccerArena provides a flexible, open-source platform for prospective benchmarking of unresolved events. LLM-SoccerArena will be continuously updated, and can be directly applied to future national and international tournaments and league competitions.

LLM-SoccerArena: Benchmarking LLMs on Real-World Predictions in Sports
Large language models are increasingly used to help make decisions about the future, but it is difficult to test how well they actually predict real-world events. Most existing benchmarks are static, meaning they test models on questions where the answers are already known, which can lead to models simply "remembering" facts rather than truly forecasting. LLM-SoccerArena addresses this by creating a live, prospective benchmark that forces models to make predictions about sports events before they happen, ensuring that the models are genuinely synthesizing information under uncertainty.

A Live, Prospective Testing Ground

The core of LLM-SoccerArena is its ability to record forecasts for unresolved events. By using a standardized protocol, the platform registers upcoming soccer matches and tournament questions, then triggers models to provide predictions at specific times before kickoff. Because the outcomes are unknown at the time of the forecast, the benchmark prevents "data leakage"—where a model might accidentally rely on information it shouldn't have. The platform records everything from the model’s final prediction and confidence level to its reasoning process and the specific tools it used to gather information.

Factorial Design for Deeper Insights

To understand what makes a model a better forecaster, the researchers used a "factorial design." This means they tested the same models under different conditions to see how specific variables influence accuracy. They varied four key dimensions:

  • Model Version: Comparing different iterations of state-of-the-art LLMs.

  • Information Access: Testing models with and without the ability to use live web search.

  • Prompting Strategy: Comparing how models perform when asked for specific scorelines versus general outcome probabilities.

  • Forecast Horizon: Measuring performance at different time intervals, such as at the start of a tournament versus just two hours before a match.

Key Findings and Performance

The researchers demonstrated the platform by evaluating seven different LLMs during the 2026 FIFA World Cup, covering 104 matches and 15 tournament-related questions. One notable finding from this initial large-scale test was that while giving models access to the web did improve their forecasting performance, the margin of improvement was relatively small—specifically, a 0.023 improvement in the Brier score. This suggests that while web search is helpful, the underlying reasoning capabilities of the models play a significant role in their ability to navigate uncertainty.

A Flexible, Open-Source Future

LLM-SoccerArena is designed to be a permanent, evolving resource. Because it is fully open-source and automated, it is not limited to one tournament. The platform is built to be applied to ongoing league competitions, such as the English Premier League or the German Bundesliga, providing a continuous stream of new data. By making the entire process—from the prompts used to the final evaluation—public and auditable, the researchers have created a framework that allows the community to track how AI forecasting capabilities evolve over time.

Comments (0)

No comments yet

Be the first to share your thoughts!