Back to AI Research

AI Research

Self-Play Pretraining with Zero Data | AI Research

Key Takeaways

  • What the paper is about Advances in language modeling have been driven by scaling pretraining on ever more data.
  • Advances in language modeling have been driven by scaling pretraining on ever more data.
  • Yet, the training data is still largely curated on the model's behalf.
  • A more general approach to pretraining would let the model learn to generate the data most useful for its own improvement.
  • This would provide an effectively unbounded source of training data, limited by compute rather than human knowledge.
Paper AbstractExpand

Advances in language modeling have been driven by scaling pretraining on ever more data. Yet, the training data is still largely curated on the model's behalf. A more general approach to pretraining would let the model learn to generate the data most useful for its own improvement. This would provide an effectively unbounded source of training data, limited by compute rather than human knowledge. We introduce Self-Play Pretraining with Zero Data, an initial proof-of-concept towards realizing this vision. Our procedure casts synthetic data generation as a search over the space of all computable structure, taking inspiration from Solomonoff induction. Starting from random initialization, two models learn in tandem: a generator proposes programs interpreted by a universal Turing machine, generating byte sequences, while a learner autoregressively predicts these byte sequences. The learner is trained with standard cross-entropy, while the generator is trained with reinforcement learning to produce sequences at the frontier of the learner's capabilities, yielding an adaptive curriculum. A universal Turing machine gives us a search space over all computable data-generating processes, imposing little domain-specific structure, and self-play searches over this space for useful training data. We test whether zero-shot performance on natural data improves predictably with self-play compute; this is a clean test of transfer since neither generator nor learner is trained on natural data. Across several natural datasets, zero-shot loss exhibits predictable scaling in compute. The models also exhibit in-context learning, and discover recognizable mathematical sequences during training.

What the paper is about

Advances in language modeling have been driven by scaling pretraining on ever more data. Yet, the training data is still largely curated on the model's behalf. A more general approach to pretraining would let the model learn to generate the data most useful for its own improvement. This would provide an effectively unbounded source of training data, limited by compute rather than human knowledge. We introduce Self-Play Pretraining with Zero Data, an initial proof-of-concept towards realizing this vision. Our procedure casts synthetic data generation as a search over the space of all computable structure, taking inspiration from Solomonoff induction. Starting from random initialization, two models learn in tandem: a generator proposes programs interpreted by a universal Turing machine, generating byte sequences, while a learner autoregressively predicts these byte sequences. The learner is trained with standard cross-entropy, while the generator is trained with reinforcement learning to produce sequences at the frontier of the learner's capabilities, yielding an adaptive curriculum. A universal Turing machine gives us a search space over all computable data-generating processes, imposing little domain-specific structure, and self-play searches over this space for useful training data. We test whether zero-shot performance on natural data improves predictably with self-play compute; this is a clean test of transfer since neither generator nor learner is trained on natural data. Across several natural datasets, zero-shot loss exhibits predictable scaling in compute. The models also exhibit in-context learning, and discover recognizable mathematical sequences during training. The same large language models question is explored in LimiX-2, which adds a research perspective.

What it covers

Self-Play Pretraining with Zero Data Thanks: G. Bruno De Luca, Aditya Cowsik, and Kfir Dolev began this work while affiliated with the Stanford Institute for Theoretical Physics. Aditya Cowsik † † thanks: Equal contribution; authors listed alphabetically. Correspondence to [email protected] , [email protected] , [email protected] , and [email protected] . Affiliation: Independent Researcher Kfir Dolev 1 1 footnotemark: 1 Affiliation: Tel Aviv University Michael Y. Li 1 1 footnotemark: 1 Affiliation: Stanford University G. Bruno De Luca Affiliation: LAPTh, USMB Nourya Cohen Affiliation: Tel Aviv University Noah D. Goodman Affiliation: Stanford University Yoav Levine Affiliation: Tel Aviv University Abstract Advances in language modeling have been driven by scaling pretraining on ever more data. Yet, the training data is still largely curated on the model’s behalf. A more general approach to pretraining would let the model learn to generate the data most useful for its own improvement. This would provide an effectively unbounded source of training data, limited by compute rather than human knowledge. We introduce Self-Play Pretraining with Zero Data , an initial proof-of-concept towards realizing this vision. Our procedure casts synthetic data generation as a search over the space of all computable structure, taking inspiration from Solomonoff induction. Starting from random initialization, two models learn in tandem: a generator proposes programs interpreted by a universal Turing machine, generating byte sequences, while a learner autoregressively predicts these byte sequences. The learner is trained with standard cross-entropy, while the generator is trained with reinforcement learning to produce sequences at the frontier of the learner’s capabilities, yielding an adaptive curriculum. A universal Turing machine gives us a search space over all computable data-generating processes, imposing little domain-specific structure, and self-play searches over this space for useful training data. We test whether zero-shot performance on natural data improves predictably with self-play compute; this is a clean test of transfer since neither generator nor learner is trained on natural data. Across several natural datasets, zero-shot loss exhibits predictable scaling in compute. The models also exhibit in-context learning, and discover recognizable mathematical sequences during training. “By teaching, we learn.” Seneca “So much from so little, almost everything from almost nothing.” John Archibald Wheeler 1 Introduction Figure 1: Self-Play Pretraining with Zero Data . Starting from randomly initialized models and using only synthetically generated data, self-play produces predictable scaling on held-out natural data. (Top) Our procedure casts synthetic data generation as search over the space of computable structure. A generator proposes programs, which are executed to produce byte sequences used to train a learner by next-token prediction. The generator is trained via reinforcement learning to propose programs near the frontier of the learner’s capabilities, measured by how strongly the learner’s gradients align with the learner’s recent learning trajectory with respect to an AdamW-preconditioned inner product. (Lower left) Self-play produces predictable improvements in validation loss with compute across natural datasets. We show the compute-optimal frontier over model sizes, number of self-play rounds, and ensemble sizes. (Lower right) The resulting learner exhibits in-context learning on held-out tasks without any gradient updates. We report empirical success under greedy decoding as a function of the number of in-context examples m m , averaged over independently sampled task instances. Advances in language modeling have been driven by pretraining on Internet data. Yet, the training data is still largely curated and constructed on the model’s behalf through large-scale data curation efforts ( Li et al., 2025 ; Soldaini et al., 2024 ; Penedo et al., 2023 ) , data mixture design ( Xie et al., 2023 ; Chen et al., 2026 ) , and hand-designed synthetic generators ( Gunasekar et al., 2023 ; Yang et al., 2025 ) . A more generic approach to pretraining would let the model itself learn to generate training data that is most useful for its own improvement. Such an approach would extend the Bitter Lesson ( Sutton, 2019 ) to the training data itself, minimizing hand-engineered inductive bias and creating a self-contained path to scaling, where compute alone can sustain continued improvement ( Kim et al., 2026b ; Silver and Sutton, ) . As an initial step towards this vision, we introduce a self-play algorithm for pretraining from zero data . Starting from random initialization, two autoregressive transformers co-evolve: a generator proposes programs for a minimal universal Turing machine whose execution produces byte sequences, while a learner is trained on those sequences via next-token prediction. We use a universal Turing machine to make the space of all synthetic training data as expressive as possible while imposing minimal domain-specific structure: any computable data-generating process can, in principle, be represented as a program. However, the space of computable structure is enormous, and only a small subset of programs produce sequences that are useful for the learner. Moreover, whether a sequence is useful changes as the learner improves. This is why a self-play approach that adapts to the learner is natural: rather than specifying useful structure in advance, we let the generator discover which programs are most useful to the learner as training progresses. To encourage this behavior, the generator is trained with reinforcement learning using a learning-progress reward, shifting probability toward programs whose outputs lie near the frontier of the learner’s current capabilities. These design choices place our approach in the lineage of classical universal prediction ( Solomonoff, 1964 ; Grau-Moya et al., 2024 ; Hutter, 2000 ; Bloem, 2025 ; Merhav and Feder, 1998 ) , which formalizes how induction can be possible when the hypothesis class contains all computable data-generating processes and the learner has unbounded compute: we aim at an efficient computable approximation to universal prediction. The resulting self-generated data is only valuable insofar as what the learner learns transfers to natural data. Our key hypothesis is that self-play over this space of computable data-generating processes can discover generic predictive regularities—such as copying, recursion, and hierarchical composition—that improve prediction on natural data. Importantly, we hypothesize that these regularities capture structure independent of contingent information —the particular facts, symbols, or modality of any one dataset—that can therefore transfer across data-generating processes. Indeed, prior work on formal, algorithmic, and non-linguistic pretraining distributions provides evidence that such cross-distribution transfer is possible ( Grau-Moya et al., 2024 ; Papadimitriou and Jurafsky, 2020 ; Hu et al., 2025 ; Lee et al., 2026 ) . We test this hypothesis through a compute-optimal scaling law analysis. Concretely, we train randomly-initialized transformers at various scales via self-play and evaluate the resulting learners zero-shot on held-out datasets spanning natural language, images, speech, melodies, DNA, and mathematical sequences. For each dataset, we construct a compute-optimal frontier over model size, self-play rounds, and ensemble size. Across these diverse domains, a single family of self-play models exhibits predictable power-law improvements in zero-shot loss with compute, with scaling exponents comparable to those obtained by training directly on natural data. Importantly, this is a clean test of transfer since we deliberately run this process tabula rasa : both models are randomly initialized and all learner training data is generated through self-play. Thus, our experiments isolate the effect of our self-play procedure and test whether useful predictive structure can emerge ex nihilo . We also find that the learner develops in-context learning on completely held-out tasks, and the generator discovers known mathematical sequences. 2 Self-Play Pretraining with Zero Data We introduce a self-play formulation of pretraining that involves searching over the space of computable structure. Every training sequence is the output of a program executed on a fixed universal Turing machine U U and these programs are produced by a learned generator. Our method has two components: (1) a Learner π θ \pi_{\theta} , an autoregressive language model trained to predict program outputs, and (2) a Generator g ϕ g_{\phi} , an autoregressive language model defined over programs for U U . Both are transformers with the same architecture, trained from random initialization ; the generator is trained via reinforcement learning (RL) and the learner is trained via next-token prediction. Importantly, we consider this tabula rasa setup to cleanly test whether self-play can generate structure that transfers to natural data. Each round of self-play proceeds as follows: 1. Program generation: Sample N N programs from the generator { x i } i = 1 N ∼ g ϕ {x_{i}}{i=1}^{N}\sim g{\phi} . 2. Execution: Run each program on U U to obtain output sequences y i = U ⁡ ( x i , ω i ) y_{i}=U(x_{i},\omega_{i}) , where ω i \omega_{i} is a random input tape. 3. Learner and Generator update: The learner takes one gradient step on the output sequences, optimizing the standard next-token loss. The generator takes a policy gradient step with a learning-progress reward that encourages the generator to propose programs at the frontier of the learner’s capabilities. In addition, the generator is updated via a supervised fine-tuning objective on existing programs to mitigate catastrophic forgetting and on mutated programs to promote exploration. 2.1 Program space We would like the generator’s search space to be as expressive as possible while imposing little domain-specific structure. We therefore use programs for a minimal universal Turing machine as the substrate for generating synthetic data. Specifically, we use a Brainfck-like Turing-complete language, following Grau-Moya et al. (2024) . Its primitive instructions manipulate a byte-valued tape, implement loops, and read or emit bytes. Because the language is universal, any computable data-generating process can, in principle, be represented as a program. Let x ∈ 𝒜 ≤ L x\in\mathcal{A}^{\leq L} denote a program generated over the machine’s instruction alphabet. Executing x x on the universal machine U U with random input tape ω \omega produces a byte sequence y = U ⁡ ( x , ω ) ∈ { 0 , … , 255 } T . y=U(x,\omega)\in{0,\ldots,255}^{T}. The random input tape allows a single program to represent a distribution over output sequences. These output bytes, rather than the programs themselves, constitute the learner’s training data. We design the execution semantics so that every generated string is executable: programs cannot fail through syntax or memory errors, and execution always produces a bounded-length output. We defer the precise execution semantics and resource limits to Appendix E . 2.2 Objectives Program pool. At each self-play round e e , we construct a pool of programs ℬ e = ℬ e fresh ∪ ˙ ℬ e mut ∪ ˙ ℬ e replay , \mathcal{B}{e}=\mathcal{B}{e}^{\mathrm{fresh}}\mathbin{\dot{\cup}}\mathcal{B}{e}^{\mathrm{mut}}\mathbin{\dot{\cup}}\mathcal{B}{e}^{\mathrm{replay}}, containing fresh samples from the current generator, local mutations of previously high-reward programs, and programs replayed from earlier rounds. Fresh samples provide global exploration, mutations refine promising regions of program space, and replay preserves useful structures discovered earlier in training. We write M e = | ℬ e | M_{e}=|\mathcal{B}{e}| for the number of programs in the pool. Details for mutation, replay, and the program bank are offered in Appendix G . Learner Update. The learner is trained by standard next-token prediction on program outputs. For an output sequence y y , define the per-sequence loss ℒ ⁡ ( y , θ ) \mathcal{L}(y;\theta) as the mean cross-entropy over its content tokens. If y i = U ⁡ ( x i , ω i ) y{i}=U(x_{i},\omega_{i}) is the output obtained by executing program x i x_{i} , the learner objective in round e e is ℒ learner ​ ( θ , ℬ e ) = 1 M e ​ ∑ i ∈ ℬ e ℒ ⁡ ( y i , θ ) . \mathcal{L}{\mathrm{learner}}(\theta;\mathcal{B}{e})=\frac{1}{M_{e}}\sum_{i\in\mathcal{B}{e}}\mathcal{L}(y{i};\theta). (1) Thus fresh, mutated, and replay programs all train the learner. Generator Reward. The generator’s reward must be capable of identifying programs with useful structure from programs without any external feedback. Initially, we considered a reward based on how difficult a sequence is to predict, motivated by prior work on self-play ( Dong and Ma, 2025b ; Bailey et al., 2026a ) . However, this has a fundamental failure mode: a program can be made arbitrarily difficult without containing useful structure—for example, by injecting random bytes into an otherwise predictable sequence. To avoid this degeneracy, we evaluate a new program based on whether it builds on what the learner has actually been able to learn. Intuitively, the learner’s change in parameters summarizes this: learning signals arising from reusable structure accumulate, whereas we expect that idiosyncratic effects that are not learnable do not. Concretely, we reward the generator for producing programs whose learner gradients align with the learner’s current learning trajectory. Let p ⁡ ( e ) = ⌊ e / 2 ⌋ , δ ​ θ e = θ p ⁡ ( e ) − θ e , p(e)=\lfloor e/2\rfloor,\qquad\delta\theta_{e}=\theta_{p(e)}-\theta_{e}, where θ p ⁡ ( e ) \theta_{p(e)} is the learner checkpoint at the lookback horizon. The reward for program x i x_{i} with output y i y_{i} is r i = | ⟨ ∇ θ ℒ ​ ( y i , θ e ) , P e ⊙ δ ​ θ e ⟩ | , r_{i}=|\left\langle\nabla_{\theta}\mathcal{L}(y_{i};\theta_{e}),P_{e}\odot\delta\theta_{e}\right\rangle|, (2) where P e = lr v ^ e + ϵ P_{e}=\frac{\mathrm{lr}}{\sqrt{\hat{v}{e}}+\epsilon} is the diagonal AdamW step operator obtained from the learner’s optimizer state. We use the lookback window of ⌈ e / 2 ⌉ \lceil e/2\rceil so that signals which take a long time to appear can be measured. The growing window helps average over short-term fluctuations and produce a signal which becomes more stable as training progresses, but allows early mistakes to eventually be forgotten. We provide some additional intuition below, but we emphasize that we selected this reward after searching through several possibilities at small scale. Detailed analysis can be found in table 5 . Formally, this reward is a preconditioned gradient-alignment score, between the learner’s gradient on a program’s output and the learner’s parameter movement over the lookback window, with the diagonal AdamW preconditioner defining the inner product; we found that using this pre-conditioning was important in line with Thrush et al. (2026) . Intuitively, this reward favors programs that are not yet mastered, but whose structure extends what the learner has already shown it can learn. We expect that programs that are already mastered induce nearly zero gradients, while programs containing unrelated or unlearnable structure induce gradients that do not align with the learner’s parameter movement. Both receive little reward. Instead, high reward is assigned to programs that induce substantial learning, but only in directions congruent with the learner’s recent progress; this concentrates the generator on the frontier of the learner’s current capabilities. We avoid materializing full gradients by using forward mode automatic differentiation ( Griewank and Walther, 2008 ) to calculate Equation 2 . 1 1 1 Forward mode kernel implemented in https://github.com/amorehead/jvp_flash_attention Policy-Gradient RL Objective. We train the generator using a KL-regularized expected reward J RL ( ϕ ) = 𝔼 x ∼ g ϕ [ r ( x ) ] − β KL ( g ϕ ∥ g 0 ) , J{\mathrm{RL}}(\phi)=\mathbb{E}{x\sim g{\phi}}[r(x)]-\beta,\mathrm{KL}(g_{\phi}|g_{0}), (3) where β \beta is the KL regularization coefficient and g 0 g_{0} is the fixed uniform program prior defined as g 0 ​ ( x ) = | 𝒜 | − ℓ ⁡ ( x ) . g_{0}(x)=|\mathcal{A}|^{-\ell(x)}. Here ℓ ⁡ ( x ) \ell(x) is the number of tokens up to and including the terminating F . This is the natural analog of the Solomonoff prior 2 − | p | 2^{-|p|} ( Solomonoff, 1964 ) , which weights programs according to description length. The generator is initialized near g 0 g_{0} and regularized toward it throughout training. Since the vanilla policy gradient estimator is high variance, we consider a GRPO (batch-level) based estimator ( Guo et al., 2025 ) . Let r ¯ e \bar{r}{e} and σ r , e \sigma{r,e} denote the mean and standard deviation, respectively, of the rewards { r i } i ∈ ℬ e {r_{i}}{i\in\mathcal{B}{e}} over the full round pool. We define the advantage of program i i as A i = r i − r ¯ e σ r , e + ϵ − β ⁡ ( log ⁡ g ϕ ​ ( x i ) − log ⁡ g 0 ​ ( x i ) ) . A_{i}=\frac{r_{i}-\bar{r}{e}}{\sigma{r,e}+\epsilon}-\beta\left(\log g_{\phi}(x_{i})-\log g_{0}(x_{i})\right). Since our bank consists of off-policy samples, we use a sequence-level importance ratio correction ρ i = g ϕ ​ ( x i ) g ϕ old , i ​ ( x i ) \rho_{i}=\frac{g_{\phi}(x_{i})}{g_{\phi_{\mathrm{old},i}}(x_{i})} ( Zheng et al., 2025 ) . It is equal to one for fresh on-policy samples; for replay samples, its denominator is the sampling probability stored when the program originally entered the replay bank. The policy-gradient term is then ℒ PG ( ϕ ) = − 1 | ℬ e ∖ ℬ e mut | ∑ i ∈ ℬ e ∖ ℬ mut stopgrad ( ρ i ) log g ϕ ( x i ) stopgrad ( A i ) . \mathcal{L}{\mathrm{PG}}(\phi)=-\frac{1}{\left|\mathcal{B}{e}\setminus\mathcal{B}^{\text{mut}}{e}\right|}\sum{i\in\mathcal{B}{e}\setminus\mathcal{B}^{\text{mut}}}\operatorname{stopgrad}!\left(\rho{i}\right),\log g_{\phi}(x_{i}),\operatorname{stopgrad}!\left(A_{i}\right). (4) Mutation rows are excluded because they were not sampled from a proposal distribution with a well-defined log probability. We clip the sequence-level importance ratio to ensure ρ i ∈ [ e − 20 , e 20 ] \rho_{i}\in[e^{-20},e^{20}] . Expert Iteration. To prevent forgetting, we replay previous programs by distilling high-reward programs back into the generator using reward-weighted supervised fine-tuning over the full pool ℬ e \mathcal{B}{e} . This is a technique used to mitigate forgetting in pretraining and RL ( Ibrahim et al., 2024 ; Schaul et al., 2016 ) . We assign each program a normalized sequence-level weight w i = [ r i ] + ∑ j ∈ ℬ e [ r j ] + , [ r ] + ≡ max ⁡ ( r , 0 ) , w{i}=\frac{[r_{i}]{+}}{\sum{j\in\mathcal{B}{e}}[r{j}]{+}},\qquad[r]{+}\equiv\max(r,0), and optimize ℒ EI ( ϕ ; ℬ e ) = − ∑ i ∈ ℬ e w i log g ϕ ( x i ) . \mathcal{L}{\mathrm{EI}}(\phi;\mathcal{B}{e})=-\sum_{i\in\mathcal{B}{e}}w{i}\log g_{\phi}(x_{i}). The generator’s full training objective is ℒ generator ​ ( ϕ ) = ℒ PG ​ ( ϕ ) + λ EI ​ ℒ EI ​ ( ϕ , ℬ e ) , \mathcal{L}{\mathrm{generator}}(\phi)=\mathcal{L}{\mathrm{PG}}(\phi)+\lambda_{\mathrm{EI}},\mathcal{L}{\mathrm{EI}}(\phi;\mathcal{B}{e}), (5) where λ EI = 1.0 \lambda_{\mathrm{EI}}=1.0 controls the strength of the reward-weighted SFT term. 2.3 Architecture and Tokenization The learner and generator are independently parameterized decoder-only Llama transformers with identical architecture ( Touvron et al., 2023 ) . We use byte-level tokenization with a fixed vocabulary of 256 byte values because the universal machine produces raw bytes, and because it enables a clean, modality-agnostic evaluation of universal prediction: how well the model predicts the next byte in arbitrary sequences encoding text, images, audio, or other data. Programs and outputs are prefixed by the bytes S and O , respectively. During program generation, logits are restricted to the eight Brainfck instructions, ten canonical single-byte macro instructions, and the end-of-program token F , whereas output sequences may contain any byte value. 3 Empirical Results 3.1 Universal zero-shot transfer scaling laws Figure 2: Self-play learns predictive structure that transfers across modalities. We compare self-play to two fixed synthetic pretraining distributions: programs sampled from a universal prior over Brainf*ck programs and probabilistic context-free grammars (PCFGs). Sampling from the universal prior scales substantially more slowly than self-play, showing that access to a universal space of programs alone is insufficient without the adaptive curriculum. Pretraining on PCFG transfers strongly to language-like domains, but its benefits are less consistent across non-language modalities. In contrast, self-play exhibits predictable scaling across images, melody, audio, speech, text, and code despite using minimal inductive bias. The compute-optimal frontier ends where we do not find further models with smaller loss than the largest-compute model shown. In standard pretraining, scaling compute typically entails both increasing model size and training the model on more natural data. Our scaling experiments study whether increasing compute via self-play, without any natural data, produces predictable improvements in zero-shot performance on held-out natural data. Scaling recipe. We follow the scaling methodology of Kim et al. (2026b) ; Wen et al. (2026) . At each model scale, we tune hyperparameters to local optimality , in the sense defined in Kim et al. (2026b) , via coordinate descent on a geometrically-spaced grid. Because our training objective is entirely synthetic, it does not provide an obvious criterion for selecting hyperparameters; poorly tuned hyperparameters could prevent clean scaling laws. We therefore use validation loss averaged across DCLM and DNA as a model-selection signal. This introduces limited leakage through hyperparameter selection, but natural data are never used for gradient updates. Because tuning every possible hyperparameter at every scale is infeasible, we restrict the search to the most important hyperparameters based on preliminary experiments: the learner learning rate, the generator-to-learner learning-rate ratio , the batch size, and the generator KL regularization coefficient β \beta . Due to compute constraints, we fix a large maximum training budget of 34.36B tokens rather than separately tuning the number of self-play rounds. Because we use only a short fixed warmup followed by a constant learning rate, every intermediate checkpoint is equivalent to a run stopped at that token budget. Thus, a single long run allows us to optimize over training duration retrospectively when constructing the scaling laws. When tuning batch size, we hold the total number of training tokens fixed; when the batch size changes, the number of rounds is adjusted accordingly to preserve the total token budget. Across all scales, we use a context length of 4096 tokens. We find that ensembling across randomly initialized models is useful. Therefore, at the locally-optimal hyperparameters, we train K K independent seeds per scale and form ensembles averaging the models’ predictive distributions. For each dataset and algorithm, we construct a compute-optimal frontier ( Kaplan et al., 2020 ) over model sizes, checkpoints, and ensemble sizes. A point is on the frontier if and only if it achieves lower validation loss than every observed configuration with less than or equal compute. We fit each compute-optimal frontier with the asymptotic power law L ⁡ ( C ) = E + A ​ C − α , L(C)=E+AC^{-\alpha}, where L L is the observed loss in bits per byte and E E is a fitted asymptotic loss floor. We fit the model independently for each dataset. Self-play exhibits universal zero-shot power-law scaling in compute. As shown in Figure 1 , we observe scaling laws over a diverse range of modalities: text, images, and music. We present additional results in Figure 7 . The scaling laws over these modalities are broadly similar, as seen in Table 2 , (with DNA as the exceptional case). We discuss the implications of this in Section 4 . In brief, we expect this to be the case when learning universal structure rather than contingent knowledge is the bottleneck to scaling. Importantly, these scaling results are entirely zero-shot: that is, they arise without any gradient steps on any of the evaluation datasets. Pretraining on a fixed universal program prior exhibits slow scaling. To isolate the value of self-play, we compare self-play against a non-adaptive baseline over exactly the same program space; we use the same mixture of validation loss on DCLM and DNA. Instead of learning a distribution over programs, the baseline samples programs from a fixed Solomonoff-style prior ( Solomonoff, 1964 ) : instruction tokens are drawn i.i.d. until termination, giving full support to every finite program while favoring shorter descriptions. Thus, both methods have access to the same universal space of computable structure; they differ only in whether the sampling distribution adapts to the learner. Figure 2 shows that fixed sampling scales substantially more slowly, demonstrating that access to a universal program space alone is not enough—self-play must learn where in that space to allocate training compute. We offer additional evaluation datasets in Figure 7 . Qualitatively, we see that self-play’s improvement is because it discovers programs whose outputs exhibit recognizable mathematical structure ( Table 1 ) far earlier than we would expect under uniform sampling from the universal prior: across 1.64 × 10 8 1.64\times 10^{8} programs drawn from the uniform prior, we find no instances of any family except arithmetic sequences. Because such structures are common in mathematical modeling, this shows that our self-play algorithm can efficiently identify universal data, and that this is one mechanism driving the faster scaling observed in Figure 2 . Family (mod 256) Example program Its output Earliest round 𝔼 ⁡ [ first round ] \mathbb{E}[\text{first round}] (univ. prior) Arithmetic S+[.++] 1 , 3 , 5 , 7 , 9 , … 1,3,5,7,9,\ldots 0 ≈ 105 \approx 105 Fibonacci S,[[.C>.C>] 1 , 1 , 2 , 3 , 5 , … 1,1,2,3,5,\ldots 512 > 53,000 >53{,}000 Geometric S+[.L>] 1 , 3 , 9 , 27 , 81 , … 1,3,9,27,81,\ldots 256 > 53,000 >53{,}000 Quadratic S,.[>VX 53,000 >53{,}000 Cubic S+[[-.L>L>-]-] 0,254,236 , 74 , … 0,254,236,74,\ldots 512 > 53,000 >53{,}000 Table 1: Program families with recognizable mathematical structure discovered by the generator during self-play. Earliest round gives the earliest round in which a member of the family first appears during training, while univ. prior gives the expected first appearance round if programs are drawn from the universal prior, including the added primitives. See appendix C for additional details. PCFG pretraining is effective on language-like domains but lacks broad cross-domain transfer. In Figure 2 , we also compare against pretraining on probabilistic context-free grammars (PCFGs), which provide a hand-designed source of hierarchical and compositional structure particularly well suited for language; for details of PCFG data generation see Appendix H . We expect pretraining on PCFG to be highly competitive on language-like domains, but its inductive bias is specialized to a particular class of structure. In contrast, our self-play procedure is, in principle, universal. Consistent with this interpretation, PCFG pretraining is stronger on text and code, where its inductive bias is well matched, while self-play substantially outperforms it on images, music, audio, and speech. Thus, self-play does not always match the performance of a specialized prior on domains where that prior is particularly well suited; however, it learns structure that transfers more broadly across modalities. We find that models trained on PCFG and the universal prior fail on our ICL evaluations in Figure 4 . Figure 3: Later generator checkpoints provide additional, non-redundant training value. For each generator-training endpoint T > 0 T>0 we build a fixed corpus of programs by sampling from 16 generator snapshots spaced T / 16 T/16 apart and ending at T T ; g 0 g_{0} uses the untrained generator. A 1M-parameter learner is then trained from scratch on each corpus, four seeds per endpoint. (a) Epiplexity, the excess training loss the learner accumulates before converging, grows steadily with T T : corpora written by later checkpoints contain more structure. (b) Out-of-distribution validation BPB of 4 ensembled seeds on text, audio, and images falls with T T . Together, the two panels show that later checkpoints enrich the curriculum rather than repeating earlier material, and that this additional data improves transfer to unseen datasets. The generator produces increasingly useful training data. We next ask whether the generator improves over the course of self-play. For each endpoint T T , we construct a fixed corpus 𝒟 T \mathcal{D}{T} by sampling 4.19M programs uniformly across 16 generator checkpoints up to T T (with g 0 g{0} as untrained), and train a fresh 1M-parameter learner for one epoch on each corpus under a fixed token budget. We evaluate each corpus by its epiplexity —the amount of structure extractable by a compute-bounded learner ( Finzi et al., 2026 ) —and by zero-shot performance on held-out text, audio, and images (Figure 3 ). Both improve steadily with T T : later generators produce data containing more learnable structure and yielding The same ai evaluation question is explored in What Should We Ask Next? Retrieval-Aware..., which adds a research perspective. as detailed in the full paper on Arxiv The same large language models question is explored in Infinite-Parameter LLMs, which adds a research perspective.

Comments (0)

No comments yet

Be the first to share your thoughts!