A three-level chessboard changes what an AI needs to value. In Dragonchess, pieces can interact across vertically stacked boards, and familiar chess heuristics need adjustment. A new study tests two ways to make that adjustment: evolving transferred evaluation weights and learning an evaluation function through self-play.
The authors of Temporal-Difference Learning for Dragonchess report that both adaptive approaches outperform handcrafted and default material evaluations in their tournament. The learned agent matches the evolved agent head-to-head at the tested search depth. That result concerns this game and experimental setup, rather than a general ranking of learning and evolutionary methods.
Different ways to evaluate the same position
The evolutionary agent uses CMA-ES to adapt a 14-dimensional piece-value vector for transferred Stockfish heuristics. The self-play agent learns a 40-feature linear evaluation with temporal-difference learning, TD(lambda). Its features include material and Dragonchess-specific signals such as frozen pieces and proximity to kings.
The learned agent trains on game outcomes without external supervision or search labels. The authors ran 22,000 training games, mixing games against the current weights with games against frozen previous policies. They selected the checkpoint using held-out matches against a depth-two alpha-beta opponent.
Training uses eligibility traces, which let an outcome influence evaluations of earlier positions. Exploration is limited in this implementation: there are no randomized openings, temperature sampling or epsilon-greedy moves. Random tie-breaking and the mixed opponent schedule supply its stochasticity. Those choices define the learner being tested and matter when comparing it with other self-play systems.
Search effort stays fixed
All tournament agents use alpha-beta search at two plies. Fixing depth lets the authors compare evaluation functions without giving one agent a deeper search budget. The field includes the evolved and learned evaluators, default material evaluation, Jackman's handcrafted Dragonchess weights and a random baseline.
Each pair plays 1,000 games, split between colors, for 10,000 games overall. The authors rewrote their earlier PyGame engine in C++, adding a transposition table and typed move generation. They report that the full tournament takes about six minutes on their workstation, making a larger sample and confidence intervals feasible.
The evolved evaluator has the highest overall score, 0.690, compared with 0.653 for the learner. Those scores count draws as half a point. Head-to-head, however, the learner wins 50.3% of decisive games against the evolved evaluator, with a 95% Wilson interval of 46.3% to 54.4%. The authors find no significant difference in that matchup.
What the result does and does not establish
The learner wins 60.1% of decisive games against Jackman's weights and 56.0% against default material evaluation. Both adaptive agents clear those baselines, supporting the paper's narrower conclusion that adapting evaluation helps in this Dragonchess setting.
Their comparison also has two explicit limits. Strong-agent matches produce many draws at this shallow search depth. And the methods use different representations: 40 learned features versus 14 evolved piece values. The experiment therefore cannot isolate optimization method from feature design.
A representation-matched follow-up could evolve the same 40 features or initialize evolution from learned weights. For now, the study offers a reproducible game-engine setting and evidence that self-play can reach the evolved evaluator's performance tier without establishing superiority over it.
Comments