What the paper is about
We introduce Kaggle Game Arena, an open and ever-expanding platform to evaluate large language models (LLMs) through competitive games. Different from static benchmarks, game arena enables models to play head-to-head matchups in structured environments where the gameplay strength naturally increases as models evolve, preventing performance saturation. This technical report details the infrastructure behind Game Arena and describes the three pilot game environments: Chess, Poker, and Werewolf. These environments span perfect information, imperfect information, and multiplayer game settings, enabling a systematic study of models' strategic planning, adaptation, and robustness under uncertainty. For each game, we provide a detailed description of the environment, evaluation metrics, and results from running full competitions across models. Through robust infrastructure and large-scale ground-truth based evaluation, Game Arena ensures reproducibility, transparency and generalizability to new games and variants over time.
What it covers
Game Arena: Strategic LLM Evaluation in Competitive Environments Bovard Doerschuk-Tiberi † † thanks: Equal contribution Yao Yan 1 1 footnotemark: 1 Justin Chiu 1 1 footnotemark: 1 Hann Wang 1 1 footnotemark: 1 Timothy Chung 1 1 footnotemark: 1 Affiliation: Martyna Plomecka 1 1 footnotemark: 1 John Schultz 1 1 footnotemark: 1 Jon Lipovetz 1 1 footnotemark: 1 Clayton Drazner Yuchen Zhuang Affiliation: Jaimie Hwang Nate Keating Riley Jones Andrew Lee Oran Kelly Ian Gemp Affiliation: Michael Aaron Laurel Prince Kate Larson Jeff Moser Harrison Jobe Chad Woodford Affiliation: Siqi Liu Andrew Wang Bo Chang Christopher D’Mello Diane Chaleff Addison Howard Affiliation: Johnny Yip Chuck Sugnet Antonio Gulli Meghan O’Connell Will Cukierski Nenad Tomasev Affiliation: Dima Yeroshenko Kinjal Parekh Roxanne Daniel Marc Lanctot Domino Weir Elsa Dong Affiliation: Daniel Hennes Melissa Nalubwama Robert Fraser Ryan Trostle Jun Peng Affiliation: Tom Mason Lloyd Hightower Chiamaka Chukwuka Yuexiang Zhai Phoebe Kirk Yi Su Affiliation: Yuting Han Jie Ren Chris Prichard Sahand Sharifzadeh Karim Hakimzadeh DJ Sterling Meg Risdal Kate Olszewska Ya Xu Orhan Firat Minmin Chen February 2026 Abstract We introduce Kaggle Game Arena, an open and ever-expanding platform to evaluate large language models (LLMs) through competitive games. Different from static benchmarks, game arena enables models to play head-to-head matchups in structured environments where the gameplay strength naturally increases as models evolve, preventing performance saturation. This technical report details the infrastructure behind Game Arena and describes the three pilot game environments: Chess, Poker, and Werewolf. These environments span perfect information, imperfect information, and multiplayer game settings, enabling a systematic study of models’ strategic planning, adaptation, and robustness under uncertainty. For each game, we provide a detailed description of the environment, evaluation metrics, and results from running full competitions across models. Through robust infrastructure and large-scale ground-truth based evaluation, Game Arena ensures reproducibility, transparency and generalizability to new games and variants over time. 1 Introduction A major challenge facing modern AI research is reliably evaluating the performance of LLMs. Standardized static benchmarks such as MMLU ( Hendrycks et al., 2021 ) , GSM8K ( Cobbe et al., 2021 ) , and HellaSwag ( Zellers et al., 2019 ) have long been important tools for measuring model progress in knowledge retrieval, mathematical reasoning, and commonsense reasoning. However, as the performance of frontier models on these fixed test sets approaches saturation, their ability to distinguish between different systems decreases ( White et al., 2024 ) . Furthermore, the static nature of these benchmarks makes them susceptible to data contamination, where test data may leak into the training corpus and lead to overestimation of performance ( White et al., 2024 ; Kapoor et al., 2024 ) . As an alternative, the community has turned to dynamic, preference-based evaluation. Chatbot Arena ( Chiang et al., 2024 ) piloted pairwise comparisons using crowdsourced human votes. MT-bench ( Zheng et al., 2023 ) introduced the LLM-as-a-judge paradigm, using strong models to evaluate weaker models through multi-turn conversations. These dynamic approaches allow for the generation of fresh evaluation instances to resist saturation. However, these approaches have limitations. Human and LLM judges are inherently subjective and often apply inconsistent standards; furthermore, they are susceptible to biases such as verbosity preference and sycophancy, which ultimately yields highly noisy evaluation data. A more reliable alternative is to anchor on ground-truth outcomes. Games offer a compelling option to address both saturation and subjectivity. Unlike static question-answer pairs, each game requires the players to adapt and make strategic decisions based on opponents’ decisions. As the models improve and their strength increases, the games played evolve and are less subject to saturation. Meanwhile, the performance of the players is objectively measurable based on outcomes like wins, losses and draws. There is a long track record of using games to evaluate earlier machine learning models, from Deep Blue’s landmark victory in Chess ( Campbell et al., 2002 ) to AlphaGo’s breakthrough in Go ( Silver et al., 2016 ; Silver et al., 2017 ) , AlphaZero’s superhuman play across Chess, Go and shogi ( Silver et al., 2018 ) , and Cicero’s human-level performance in the natural-language negotiation game Diplomacy ( Bakhtin et al., 2022 ) . AlphaStar ( Vinyals et al., 2019 ) rivaled top professional players in the real-time strategy game StarCraft II. Libratus ( Brown and Sandholm, 2018 ) and Pluribus ( Brown and Sandholm, 2019 ) demonstrated superhuman poker play. In recent years, an increasing number of studies have begun using games to evaluate the capabilities of LLMs. AgentBench ( Liu et al., 2024 ) evaluates LLM agents in eight interactive environments, including game-like tasks, revealing significant performance gaps between commercial and open-source models. GTBench ( Duan et al., 2024 ) evaluated LLMs from a game-theory perspective, covering ten strategic tasks ranging from complete to incomplete information. PokerBench ( Zhuang et al., 2025 ) focuses on poker decision-making, compiling 11,000 scenarios with game-theory-optimal solutions to assess LLMs’ grasp of strategic play under uncertainty. SmartPlay ( Wu et al., 2024 ) tests six distinct games, covering a range of difficulty from Rock-Paper-Scissors to Minecraft. Game Reasoning Arena ( Cipolina-Kun et al., 2025 ) leverages the OpenSpiel framework to capture detailed reasoning traces and benchmark LLMs across multiple classical games. PokerBattle ( PokerBattle.ai, 2025 ) , an online experiment in which nine LLMs competed in approximately 3,800 hands of no-limit hold’em, offered early empirical evidence that model scale and reasoning capability correlate with poker performance. These studies consistently show that even strong models struggle in perfect-information deterministic games and even more in imperfect-information games requiring sophisticated belief modeling. Games also serve as evergreen benchmarks. The complexity and strategic depth of gameplay precludes rote memorization and assesses a model’s ability to handle out-of-distribution scenarios. Moreover, different classes of games demand distinct cognitive capabilities: perfect-information games such as Chess test strategic planning and search, while imperfect-information games such as poker require probabilistic reasoning, opponent modeling and adaptation, and risk management. These capabilities are directly applicable to real-world decision-making under uncertainty (e.g., financial strategy, supply-chain planning, disaster response, etc). However, existing game benchmarks are typically released as fixed datasets or isolated evaluations. They are not designed for continuous evaluation, flexible onboarding of new games, or large-scale head-to-head competition across heterogeneous environments ( Cipolina-Kun et al., 2025 ) . Many also lack the statistical rigor needed for reliable conclusions: benchmarks may report results from only a few thousand interactions, far below the threshold required for significance in high-variance games. They, as a result, offer limited support for longitudinal measurement of progress. More recently, MindGames Wang et al. (2026) , a competition run at NeurIPS 2025, proposed as a live arena but with a focused scope of four theory-of-mind style card games. We introduce Kaggle Game Arena 1 1 1 https://www.kaggle.com/game-arena . The open-source implementation is available at https://github.com/[google](/news/7KrD5J28Wq8dVeaPDgX3)-deepmind/game_arena . , an ever-expanding evaluation platform that uses standardized harnesses and rigorous evaluation. Models interact with their opponents in a dynamic fashion in predefined game environments. The trajectories of actions, states, outcomes, and reasoning traces are recorded and used for evaluation and analysis. In this paper, we release three benchmark environments covering diverse information and interaction paradigms: Chess, Poker and Werewolf. For each environment, we provide detailed descriptions of game rules, harness design, dataset structure, and evaluation metrics, along with empirical results from large-scale gameplay for selected frontier foundation models. Across environments, we prioritize statistical rigor by employing variance-reduction techniques, large sample sizes (e.g., 900,000 hands in the poker benchmark alone), and domain-expert consultation to ensure that reported rankings reflect real capability differences. In Game Arena, new games, variants, and evaluation methodologies can be added without redefining the core framework or invalidating prior results. This data-centric design enables longitudinal analysis of model behavior and supports research into benchmark design itself to actively combat memorization and contamination, and improve strategic generalization. 2 Methods The first set of game environments released is designed around a shared protocol. In every game, we use a uniform, text-based harness: at each decision point, models receive a natural language description of both the current game state and its history, and must return a single action in a prescribed format. When a model produces an invalid response, the harness permits a small number of retries with minimal feedback to attempt a new valid response. Full prompt templates are given in Appendix B ). Second, in all games, we use decisive outcomes including wins, losses, draws or chip counts to compute metrics and build leaderboards, which are not influenced by subjective interpretation. Third, we employed variance reduction techniques and provide bootstrapped confidence intervals for evaluation. Figure 1 demonstrates the Game Arena framework. Figure 1 : Game Arena infrastructure. The three released environments Chess, Poker and Werewolf cover differences in information structure and interaction complexity, as shown in the Table 1 . The remainder of this section will present each environment, the common evaluation framework, and the set of models evaluated. Table 1 : Comparison of Game Arena environments along key evaluation dimensions. Chess Poker Werewolf Information structure Perfect Imperfect Imperfect + asymmetric Stochasticity Deterministic Stochastic Stochastic Players per game 2 2 8 Communication channel None None Natural language Primary cognitive demand Planning, search Belief updating, risk Deception, social inference Primary metric Elo (Bradley–Terry) BB/100 Game-theoretic eval. Evaluation scale 40 games/pair 20,000 hands/pair ∼ \sim 31,500 total games 2.1 Chess Chess is a classic perfect-information game. All games are recorded in PGN (Portable Game Notation) format and follow the standard FIDE (International Chess Federation) rules, including pawn capture, castling, the 50-move rule, etc. Legality is enforced by the environment at every play. If a model continuously outputs illegal moves after the predefined number of retries, the game is counted as a loss for the model. Each pair of models plays 40 games while maintaining color balance (20 games with white pieces and 20 games with black pieces) to neutralize the advantage of the first move. In each turn, the model receives the current position in Forsyth-Edwards Notation (FEN) and a complete move history in PGN format, and the model is required to return a single valid move in standard algebraic notation (SAN). The harness does NOT include the list of valid moves as hints, so the model must independently infer the validity of their moves given the current board state. Minor variations in SAN are normalized, e.g., ambiguous moves, missing check or checkmate symbol. If the model proposes an invalid move, the harness simply responds that the move is invalid and asks the model to retry. Diagnostic information (e.g., why the move is invalid and what alternatives exist) is intentionally hidden, requiring the model to recover solely based on reasoning within the allotted limit of three retries. If the model fails after four attempts, the game ends with a loss. We refer to each retry event as a “rethink” and track the frequency of rethinking as a diagnostic metric for the reliability of the interaction (Section A.2 ). To reduce reliance on narrow openings and explore adaptability in opening game structure, we also introduce a second evaluation mode called ”Chess Opening”. Each game starts with one of the 20 most popular two-move openings from the Lichess database, requiring the model to adapt and make inferences across a wider range of opening scenarios. All other rules and interface details are identical to the free-form setting (hereafter Chess Text ). 2.2 Poker Models competed in heads-up no-limit Texas Hold’em (HU-NLHE), a standard variant for rigorous evaluation in computer poker research ( Bard et al., 2013 ) where superhuman AI has already been demonstrated ( Brown and Sandholm, 2018 ) . In our configuration, blinds are set to 1-2. Players start each hand with 100 big blinds, and chip stacks are reset between hands to isolate decision-making quality. At each decision point, the model receives a text description of the current hand state, including player position, chip stacks, private cards, community cards, and betting history. Illegal actions trigger a retry, and a second consecutive illegal action defaults to the most conservative action: a check (if there is no pending bet) or fold. Each action has a 60-minute time limit. Professional poker players were consulted in developing the system prompt, which explicitly identifies the objective as maximizing overall expected value. The model is instructed to default to a game-theoretic optimal (GTO) strategy, deviating only to exploit opponent tendencies. Without such instruction, models could reasonably adopt alternative objectives that do not reflect their true playing strength, such as playing conservatively to protect their bankroll or optimizing for entertainment value. To aid in post-game analysis, the model must provide a clear reasoning trace utilizing fundamental concepts like range advantage, pot odds, and fold equity (see Appendix B.2 for a full example). Poker is widely considered a game of adapting to and exploiting opponent tendencies. This aspect of the game, often referred to in the literature as opponent modeling, remains an active and challenging area of game theory research. Traditional computer poker competitions largely bypassed this dynamic by treating each hand independently, thereby avoiding the difficulties associated with effectively representing and processing extensive game histories. However, leveraging the ability of LLMs to naturally ingest long contexts of arbitrary text, we sought to explicitly evaluate how well models adapt to their opponents over time. To achieve this, pairwise matchups are divided into independent 100-hand episodes. This batching provides sufficient interaction for meaningful adaptation while managing context length limits and enabling parallelization. During an episode, the model receives the complete text history of all previous hands. Furthermore, to accelerate the feedback loop for opponent modeling, both players’ hole cards are revealed after each hand, regardless of whether the hand went to showdown. To mitigate the high variance inherent to poker, we adopted a duplicate poker format (hand-mirroring). Decks for the 100-hand episodes are pre-shuffled, and each 100-hand sequence is played twice, swapping the players’ seats and cards. Although this does not eliminate luck, it significantly reduces variance and allows us to directly compare how models handle the same sequence of deals. Notably, while the underlying cards are identical, the precise situations will naturally diverge over the course of an episode as models take different actions and develop unique reads on their opponents based on diverging histories. Duplicate poker is standard practice in computer poker tournaments, but to our knowledge, it has never been used in LLM poker benchmarking. Each pairwise matchup consists of 20,000 hands (10,000 individual deals × \times 2 for hand mirroring). Across a full ten-model round-robin tournament ( ( 10 2 ) = 45 \binom{10}{2}=45 matchups), this yields 900,000 hands in total, or 180,000 hands per model. This scale vastly exceeds recent benchmarks: PokerBattle ( PokerBattle.ai, 2025 ) reports approximately ∼ 3,800 {\sim}3{,}800 hands per model, while PokerBench ( Zhuang et al., 2025 ) reports fewer than 2,000 hands in its evaluation. 2.3 Werewolf Werewolf is a multiplayer, general-sum social-deduction game driven by information asymmetry. Our implementation of Werewolf features eight players. Each player is secretly assigned one of four roles, two Werewolves, one Seer, one Doctor, and four Villagers, creating two opposing teams with different strategic incentives. While Werewolf has countless variations in the literature, Game Arena deliberately utilizes this classic rule set (Appendix A.4 ). This setup ensures that complexity naturally emerges from the players’ mixed strategies and social dynamics rather than from intricate rule additions. By navigating this ambiguity, models must exercise “soft skills” such as negotiation, coalition building, and the capacity to detect or engage in strategic deception. Beyond strategic evaluation, Werewolf provides a secure sandbox for agentic safety research. Mastery requires inhabiting opposing roles—the truth-seeking Villager and the deceptive Werewolf—forcing models to both generate and detect manipulation in a controlled setting. This dual-sided dynamic allows researchers to red-team a model’s deceptive capabilities while simultaneously assessing its robustness as a safeguard against bad actors, all without the high stakes of real-world deployment. A moderator orchestrates the game by alternating between night phases for private role-specific actions and day phases for public discussion and majority-vote eliminations. Internally, the environment tracks all state transitions and player interactions, such as votes, messages, and ability usages, via an append-only log of Event records. Each event object strictly defines its own visibility permissions. A centralized event bus enforces information asymmetry by filtering and dispatching these records to individual models based on their access privileges, separating public dialogue from private communications. To facilitate ablation studies on game rules, a plug-in protocol system implements discussion formats, such as circular or parallel, and voting rules, such as sequential or simultaneous, as interchangeable components. Finally, we validate the fairness of our baseline rule configuration through a comprehensive game balance analysis, detailed in Appendix A.5 . At each decision point, models receive their complete historical context, integrating all available public and private observations. The model harness employs a ReAct ( Yao et al., 2023 ) framework, prompting models with a chronological event log, a phase-specific task definition, and a system directive to prioritize team victory over individual survival. Models subsequently generate private reasoning alongside actions. Parser stability is enforced via robust parsing ( pyjson5 ), exponential backoff over endpoint failure, iterative truncation (preserve latest 75% context each time) for context overflows, and a deterministic rule-based filter to intercept inappropriate language. Upon reaching a maximum retry threshold for parsing failures, the model harness forces a turn forfeiture, which is always a valid action within the Werewolf ruleset. 2.4 Evaluation Framework As the three environments differ in outcome structure, we uses a domain-appropriate primary metric for each. For Chess, model strength is quantified via Elo-style ratings derived from the Bradley–Terry model ( Bradley and Terry, 1952 ) fitted to all pairwise match outcomes (draws scored as 0.5). Since Elo is identifiable only up to an additive constant, we anchor the scale by setting the lowest-rated model to 0. We report 95% confidence intervals from bootstrap resampling and provide an approximate external calibration by matching models against multiple Stockfish skill levels and interpolating against reference CCRL Elo mappings; this calibration is less reliable outside the engine calibration range. For Poker, performance is measured in big blinds won per 100 hands (BB/100), the standard metric that normalizes returns across stack sizes and game length. Each model’s BB/100 aggregates net winnings across all ∼ 180,000 {\sim}180{,}000 hands, with confidence intervals obtained via block bootstrap over episodes to account for within-episode correlation. Bootstrapped win-rate distributions are additionally reported for all pairwise matchups. For Werewolf, raw win rates conflate individual skill with role-assignment luck because the game is team-based and asymmetric. We therefore employ a game-theoretic evaluation (GTE) framework ( Liu et al., 2025 ) that decomposes each model’s overall skill into role-specific contributions (Werewolf, Seer, Doctor, Villager), estimating per-model, per-role parameters from the distribution of outcomes across ∼ 31,000 {\sim}31{,}000 games. Confidence intervals are computed via bootstrap; the full GTE formulation is given in Appendix A.7 . In all three environments, every model pair is evaluated under identical conditions in a full round-robin, with outcomes logged at the finest available granularity (per-move for Chess, per-hand for poker, per-game for Werewolf). While summary rankings and bracketed tournaments are provided for public communication, all analyses in this paper are grounded in the comprehensive round-robin data. 2.5 Model selection To ensure we can run sufficient matches between each model pair to reach robust conclusions, we limit model selection to the top 10 models from five frontier labs (as of February 2026) to demonstrate the Game Arena framework. GPT-5.2, GPT-5 mini, and o3 (OpenAI); Claude Opus 4.5, Claude Sonnet 4.5, and Claude Haiku 4.5 (Anthropic); Gemini 3 Pro Preview and Gemini 3 Flash Preview (Google); Grok 4 and Grok 4.1 Fast Reasoning (xAI); and DeepSeek V3.2 (DeepSeek). All models are accessed through public API endpoints using each provider’s default sampling settings and generation limits, without fine-tuning, tool augmentation, or retrieval. Per-turn token counts and inference costs are reported alongside gameplay results to support cost-performance analysis. 3 Results 3.1 Chess Figure 2 : Chess Text Benchmarks. Left: Pairwise win-rate heatmap across models. Right: Bootstrapped win-rate distribution (95% CI). Internal Game Arena Elo ratings for Chess Text are also summarized in Figure 2 and Appendix Table 1 . There is a clear gap in performance between models: Gemini 3 Pro Preview (Game Arena Elo 1325) and Gemini 3 Flash Preview (1297) make up the top tier, matching their consistently high win rates shown in the heatmap in Figure 2 . The second tier includes o3 (1009) and GPT-5.2 (933), which remain competitive but show greater performance variability compared to the Gemini series models. The remaining models lagged significantly behind: Grok 4 (773), Grok 4.1 Fast Reasoning (632) and GPT-5 mini (525) were in the third tier, and the Claude 4.5 series end up in the rear: its Opus variant scored 236, Sonnet 189, and Haiku 122. Figure 3 : Analysis of Stockfish engine evaluations Left: average win probablity over the course of game(Chess Text benchmark). Right: Gemini 3 Pro preview win probability v.s. each opponent(Chess Text benchmark). We used the Stockfish engine to identify where models’ performance gap arise. Specifically, for each position, Stockfish’s centipawn loss predictions were converted into estimated win probabilities, and after normalizing for game progress, the average of these paths was calculated. Figure 3 (left) shows that in the early stages, the models are roughly balanced, but in the middle of the game, their velocity paths quickly diverge. The strongest models, Gemini 3 Pro and Gemini 3 Flash, gradually gain an advantage, achieving a high probability of winning in the endgame. In contrast, weaker models, such as DeepSeek V3.2 and Claude 4.5 variants, show a uniform decline, indicating a systematic tendency to lose position as the models get deeper into the game. Figure 3 (right) further illustrates that Gemini 3 Pro’s advantage grows against almost all opponents during the game, approaching a near-certain victory in the final stages of most matches. The main exception is Gemini 3 Flash, where the probability of winning remains close to equilibrium with greater volatility. Overall, this temporal dynamic is consistent with the Elo ratings given in Appendix Table 1 and shows that the main differences between the models appear after the opening phase, due to the higher quality of decisions made in the middle and endgame compared to opening move selection. Additionally, we measured reliability of these models through ‘rethinking’, i.e., forced retries caused by choosing illegal moves or erroneous move formatting. See Appendix section A.2.2 . The analysis suggests that errors are concentrated in later stage of the games and this pattern is consistent with the analysis of Stockfish’s path shown in Figure 3 : weaker models not only get stuck in unfavorable positions, but also show a significant increase in making illegal moves as the game progresses. Figure 4 : Cost–performance trade-offs across four Game Arena environments. Each panel displays model performance against inference cost (per turn for Chess and Poker; per game for Werewolf). The dashed blue curve denotes the Pareto frontier within that environment, consisting of non-dominated models (no alternative achieves higher performance at equal or lower cost). Orange points indicate models that are Pareto-efficient in at least one Game Arena environment (union across panels), highlighting globally competitive models while preserving environment-specific frontiers. While running chess games in the regular Chess Text setup, we observed the models unanimously favor the Sicilian Defense opening. To test early-game adaptability, we introduce a Chess Opening variant of the benchmark. Each game starts from one of 20 popular two-ply opening positions taken from Lichess. This approach reduces dependence on a limited default repertoire, like repeatedly using the same defense, and encourages models to handle a wider variety of early-game positions. As shown in Figure Figure 4 , Panel A and B, the overall ranking and cost patterns are similar to those in Chess Text. Detailed analysis of Chess Opening is provided in the appendix. 3.2 Poker Aggregate performance by model (sorted by BB/100) is summarized in Table 2 . Figure 5 presents head-to-head match win rates alongside bootstrapped distributions of BB/100 to quantify uncertainty in the rankings. Additional metrics are detailed in 2.2 . The ten models separate into three performance tiers. The top tier comprises GPT-5.2 (+46.6), o3 (+29.7), and Grok 4 (+27.1), all of which achieved win rates substantially above the break-even line. The middle tier includes Claude Opus 4.5 (+17.6), Claude Sonnet 4.5 (+12.8), and Gemini 3 Flash Preview (+5.5), which were still profitable but at moderate rates. The bottom tier consists of four models with negative returns: DeepSeek V3.2 ( − - 10.1), Gemini 3 Pro Preview ( − - 15.2), Grok 4.1 Fast Reasoning ( − - 19.2), and GPT-5 mini ( − - 94.9). To verify the statistical robustness of the overall ranking, bootstrapped confidence intervals were plotted (Figure 5 , right panel). GPT-5.2 is distinctly separated from the field, with no overlap in 95% confidence intervals between itself, o3, and the middle- or bottom-tier models. Grok 4’s 95% interval is largely separated from the middle tier; however, its lower tail approaches the upper tail of Claude Opus 4.5, indicating that the boundary between the top and middle tiers is most distinct at the top. Within the middle tier, the distributions of Claude Opus 4.5 and Claude Sonnet 4.5 overlap in their tails, implying that the performance difference between these models, although consistent in point estimates, is less definitive. Among the bottom-tier models, DeepSeek V3.2, Gemini 3 Pro Preview, and Grok 4.1 Fast Reasoning exhibit partially overlapping distributions, suggesting uncertainty in their relative rankings. GPT-5 mini’s distribution is completely isolated, confirming its status as a clear outlier. Table 2 : Aggregate poker statistics for all models, sorted by win rate (BB/100). VPIP = voluntarily put money in pot. Att To Steal = button open-raise frequency. Call/Fold/3Bet BB v SB = big blind response frequencies when facing a button raise. Player Hands BB/100 VPIP Att To Steal (SB) Call BB v SB Fold BB v SB 3Bet BB v SB GPT-5.2 180,000 46.56 96.42 91.75 57.04 7.59 35.37 o3 180,000 29.69 88.90 87.36 43.83 21.91 34.25 Grok 4 180,000 27.11 92.04 95.02 30.79 12.53 56.68 Claude Opus 4.5 179,936 17.62 82.92 87.71 64.91 22.79 12.30 Claude Sonnet 4.5 180,000 12.83 87.29 88.61 69.26 14.44 16.30 Gemini 3 Flash Prev. 180,000 5.49 69.08 72.39 49.84 34.18 15.98 DeepSeek V3.2 179,936 − - 10.05 78.59 79.43 47.05 24.18 28.78 Gemini 3 Pro Prev. 180,000 − - 15.17 55.16 52.78 50.22 43.78 6.00 Grok 4.1 Fast Reas. 180,000 − - 19.19 81.18 as detailed in the full paper on Arxiv
Comments (0)
to join the discussion
No comments yet
Be the first to share your thoughts!