LLM-SoccerArena: Benchmarking LLMs on Real-World Predictions in Sports
Large language models are increasingly used to help make decisions about the future, but it is difficult to test how well they actually predict real-world events. Most existing benchmarks are static, meaning they test models on questions where the answers are already known, which can lead to models simply "remembering" facts rather than truly forecasting. LLM-SoccerArena addresses this by creating a live, prospective benchmark that forces models to make predictions about sports events before they happen, ensuring that the models are genuinely synthesizing information under uncertainty.
A Live, Prospective Testing Ground
The core of LLM-SoccerArena is its ability to record forecasts for unresolved events. By using a standardized protocol, the platform registers upcoming soccer matches and tournament questions, then triggers models to provide predictions at specific times before kickoff. Because the outcomes are unknown at the time of the forecast, the benchmark prevents "data leakage"—where a model might accidentally rely on information it shouldn't have. The platform records everything from the model’s final prediction and confidence level to its reasoning process and the specific tools it used to gather information.
Factorial Design for Deeper Insights
To understand what makes a model a better forecaster, the researchers used a "factorial design." This means they tested the same models under different conditions to see how specific variables influence accuracy. They varied four key dimensions:
Model Version: Comparing different iterations of state-of-the-art LLMs.
Information Access: Testing models with and without the ability to use live web search.
Prompting Strategy: Comparing how models perform when asked for specific scorelines versus general outcome probabilities.
Forecast Horizon: Measuring performance at different time intervals, such as at the start of a tournament versus just two hours before a match.
Key Findings and Performance
The researchers demonstrated the platform by evaluating seven different LLMs during the 2026 FIFA World Cup, covering 104 matches and 15 tournament-related questions. One notable finding from this initial large-scale test was that while giving models access to the web did improve their forecasting performance, the margin of improvement was relatively small—specifically, a 0.023 improvement in the Brier score. This suggests that while web search is helpful, the underlying reasoning capabilities of the models play a significant role in their ability to navigate uncertainty.
A Flexible, Open-Source Future
LLM-SoccerArena is designed to be a permanent, evolving resource. Because it is fully open-source and automated, it is not limited to one tournament. The platform is built to be applied to ongoing league competitions, such as the English Premier League or the German Bundesliga, providing a continuous stream of new data. By making the entire process—from the prompts used to the final evaluation—public and auditable, the researchers have created a framework that allows the community to track how AI forecasting capabilities evolve over time.
Comments (0)
to join the discussion
No comments yet
Be the first to share your thoughts!