Thomson: Continual Learning of Frontier Models for SovereignAI introduces a method for institutions to develop high-performance, private AI models without relying on the massive resources of a few dominant tech companies. By using a "continual learning" approach on existing open-weight models, the authors argue that organizations can achieve frontier-level performance while maintaining control over their own data, infrastructure, and governance—a concept they define as SovereignAI.
Achieving Frontier Performance
The authors argue that frontier AI performance is not exclusive to heavily funded players. Instead of training models from scratch, which is economically inefficient, or relying on limited fine-tuning, the team repurposed open-weight models (Qwen3.5-397B and Qwen3.6-35B). By applying full-weight updates and a specialized mid- and post-training stack, they improved these models across a wide range of domains. The development of the Thomson family was completed by a team of fewer than 36 engineers in three months, with the final training run for the largest model costing less than $450,000.
The Continual Learning Approach
The core of the Thomson methodology is a "π-shaped" performance pattern. This refers to the model's ability to achieve distinct improvements across targeted capabilities—such as legal, tax, and deep research tasks—while simultaneously avoiding the "forgetting" problem common in narrow domain adaptation. The authors emphasize three pillars:
Continual Learning: Using algorithms that preserve existing model stability while allowing for the acquisition of new skills.
Data-Centric Machine Learning: Focusing on high-quality data curation, including re-aligning models with institutional values and using Bayesian optimization to calibrate training mixtures.
Agentic Training: Developing models capable of using tools and deep research harnesses, with reward structures designed to reduce hallucinations and improve citation accuracy.
Performance and System Results
In evaluations, Thomson-1.0-Large performed competitively with flagship models released through mid-2026, including Gemini 3.1 Pro and Sonnet 5. In a blind preference study involving over 3,000 conversations, a system built using Thomson-1.0-Large—equipped with access to proprietary legal and news databases—was preferred by subject-matter experts over systems from major providers like OpenAI and Anthropic. The authors note that while the model excels in professional domains, it shows a clear advantage when paired with an institution's own private data.
Limitations and Considerations
The authors identify specific boundaries to their results. Coding performance fell below frontier levels, which they attribute to it not being a primary target domain. While the models successfully protected against forgetting in general reasoning and mathematics, they did not surpass the strongest proprietary models in those areas. Additionally, the authors acknowledge that while they have achieved sovereignty over model, data, and tool infrastructure, they remain dependent on external hardware providers. They also note that fair adversarial safety testing against proprietary models is difficult because of the undisclosed guardrails used by those providers.
Comments (0)
to join the discussion
No comments yet
Be the first to share your thoughts!