Back to AI Research

AI Research

Thomson: Continual Learning of Frontier Models for... | AI Research

Key Takeaways

  • Thomson: Continual Learning of Frontier Models for SovereignAI introduces a method for institutions to develop high-performance, private AI models without re...
  • We argue that frontier performance is achievable by a wide range of institutions through Continual Learning on readily available open-weight models.
  • We demonstrate this with Thomson, a general-purpose frontier model trained with an enhanced focus on high-stakes professional work.
  • Thomson performs competitively with recent frontier models across agentic tasks, safety, legal, tax & multilingualism, and large-scale Deep Research.
  • The authors argue that frontier AI performance is not exclusive to heavily funded players.
Paper AbstractExpand

The development of frontier models is commonly perceived to be the exclusive remit of a small number of heavily funded players, creating an information, economic and power asymmetry between developers and the diverse user base of modern AI. Recent public discourse acknowledges this concern, calling for SovereignAI (an organisation's capability to independently build, deploy and govern AI use), but offers little concrete advice on how this can be achieved in the short term under a diversity of funding settings. We argue that frontier performance is achievable by a wide range of institutions through Continual Learning on readily available open-weight models. Unlike limited approaches such as small-scale fine-tuning, prompt engineering, or tool-augmentation of a frozen model, our approach exploits a modern mid- & post-training stack while introducing safeguards that preserve both plasticity and stability at each stage, making the minimal number of high-impact interventions on the parameters. This yields gains comparable to those typically seen across multiple successive model generations, at compute and personnel budgets substantially lower than commonly thought, making ownership of large parts of the SovereignAI stack (model, tool infrastructure, values & data privacy) viable for far more actors. We demonstrate this with Thomson, a general-purpose frontier model trained with an enhanced focus on high-stakes professional work. Thomson performs competitively with recent frontier models across agentic tasks, safety, legal, tax & multilingualism, and large-scale Deep Research. Evaluations show a distinctive $\pi$-shaped pattern: distinct improvements across a wide range of capabilities, including those not explicitly targeted, while almost completely eliminating the forgetting problem common to narrow domain adaptation.

Thomson: Continual Learning of Frontier Models for SovereignAI introduces a method for institutions to develop high-performance, private AI models without relying on the massive resources of a few dominant tech companies. By using a "continual learning" approach on existing open-weight models, the authors argue that organizations can achieve frontier-level performance while maintaining control over their own data, infrastructure, and governance—a concept they define as SovereignAI.

Achieving Frontier Performance

The authors argue that frontier AI performance is not exclusive to heavily funded players. Instead of training models from scratch, which is economically inefficient, or relying on limited fine-tuning, the team repurposed open-weight models (Qwen3.5-397B and Qwen3.6-35B). By applying full-weight updates and a specialized mid- and post-training stack, they improved these models across a wide range of domains. The development of the Thomson family was completed by a team of fewer than 36 engineers in three months, with the final training run for the largest model costing less than $450,000.

The Continual Learning Approach

The core of the Thomson methodology is a "π-shaped" performance pattern. This refers to the model's ability to achieve distinct improvements across targeted capabilities—such as legal, tax, and deep research tasks—while simultaneously avoiding the "forgetting" problem common in narrow domain adaptation. The authors emphasize three pillars:

  • Continual Learning: Using algorithms that preserve existing model stability while allowing for the acquisition of new skills.

  • Data-Centric Machine Learning: Focusing on high-quality data curation, including re-aligning models with institutional values and using Bayesian optimization to calibrate training mixtures.

  • Agentic Training: Developing models capable of using tools and deep research harnesses, with reward structures designed to reduce hallucinations and improve citation accuracy.

Performance and System Results

In evaluations, Thomson-1.0-Large performed competitively with flagship models released through mid-2026, including Gemini 3.1 Pro and Sonnet 5. In a blind preference study involving over 3,000 conversations, a system built using Thomson-1.0-Large—equipped with access to proprietary legal and news databases—was preferred by subject-matter experts over systems from major providers like OpenAI and Anthropic. The authors note that while the model excels in professional domains, it shows a clear advantage when paired with an institution's own private data.

Limitations and Considerations

The authors identify specific boundaries to their results. Coding performance fell below frontier levels, which they attribute to it not being a primary target domain. While the models successfully protected against forgetting in general reasoning and mathematics, they did not surpass the strongest proprietary models in those areas. Additionally, the authors acknowledge that while they have achieved sovereignty over model, data, and tool infrastructure, they remain dependent on external hardware providers. They also note that fair adversarial safety testing against proprietary models is difficult because of the undisclosed guardrails used by those providers.

Comments (0)

No comments yet

Be the first to share your thoughts!