WorldCupArena is a dynamic benchmark designed to evaluate how well language models and deep-research agents can forecast future events. Unlike traditional benchmarks that test models on questions with already-known answers, this project focuses on football matches, requiring systems to analyze changing information and make predictions before a game begins. By evaluating models on the 2026 FIFA World Cup, the researchers created a framework that tests not just the final result, but also detailed predictions like scores, player lineups, match statistics, and the progression of the entire tournament.
How the Benchmark Works
The evaluation process follows a strict protocol for every match. Twenty-four hours before kickoff, models are provided with a standard evidence package (such as team form and news) or are tasked with searching for their own information. They then submit a comprehensive forecast, including win probabilities, expected scores, and specific match events. Once the match concludes, these predictions are compared against official records. The benchmark uses five distinct layers of evaluation: core match results, player lineups, specific events (like goals or cards), tactical statistics, and the overall outcome of the competition.
Evaluating Model Performance
The study tested 13 different systems across all 104 matches of the 2026 World Cup. A key finding is that looking only at "result accuracy"—who won or lost—is insufficient for distinguishing between high-performing models. While many models achieved similar success rates in predicting the winner, they showed significant differences in their ability to predict precise scorelines and detailed match statistics. The researchers also implemented a "Scoreline" metric, which provides partial credit for close misses, offering a more nuanced view of a model's predictive capability than a simple right-or-wrong score.
Search and Baseline Comparisons
One of the study's goals was to determine if allowing models to search the web for real-time information improves their forecasting accuracy. Interestingly, the results showed no consistent performance gain from adding search capabilities compared to models provided with a standardized evidence package. Furthermore, when compared to betting markets and human-fan predictions, the best-performing AI systems showed only minor improvements in result accuracy. However, these AI models did demonstrate a clearer advantage in the "Scoreline" category, suggesting they are better at narrowing down the range of likely outcomes even when they do not predict the exact final score.
Future Applications
Because the benchmark is designed to be modular, it is not limited to the 2026 World Cup. The same evaluation structure can be applied to future leagues, cups, or other sports, allowing researchers to continuously test new models as they are released. By publishing the code, prompts, and evaluation scripts, the authors have created a reusable tool that ensures future models can be judged on their ability to handle real-world uncertainty rather than relying on static, historical data.
Comments (0)
to join the discussion
No comments yet
Be the first to share your thoughts!