arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

Thomson:面向主权人工智能的前沿模型持续学习

Thomson: Continual Learning of Frontier Models for SovereignAI

Shengzhuang Chen, Jerrod Parker, Yejin Bang, Andrew M. Bean, Nabeel Seedat, Stefan Winzeck, Daniil Glazko, Jannik Zgraggen, Fangyi Yu, Scott Arnott, Dietrich Trautmann, Luca Ciuffreda, Guglielmo Bonifazi, Davide Romano, Bradley Bell, Kirsty Fielding, Daniele Giofrè, Tom Zielund, Ipshita Chatterjee, Sneha Murthy Ghantasala, Manpreet Nanreh, John Scoville, Maciej Sakowicz, Wassim Seifeddine, Lukas Thede, Jonathan Richard Schwarz

arXiv 2608.27147首次发表:更新:

发表机构

Imperial College London; DatologyAI; Lambda(帝国理工学院; DatologyAI; Lambda)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

该研究提出Thomson,通过对开放权重模型的持续学习,以低预算实现前沿模型性能,可帮助更多机构构建主权人工智能,且在多任务表现优异,几乎消除了遗忘问题。

AI 中文摘要

前沿模型的开发通常被认为是少数资金雄厚机构的专属领域,这在开发者与现代AI的多样化用户群体之间造成了信息、经济和权力的不对称。近期的公开讨论已认识到这一问题,呼吁实现主权人工智能(SovereignAI,即一个机构独立构建、部署和管控AI应用的能力),但几乎未提供在不同资金环境下短期内如何实现这一目标的具体建议。我们认为,众多机构可通过对现成的开放权重模型进行持续学习来获得前沿性能。与小规模微调、提示工程或冻结模型的工具增强等有限方法不同,我们的方法利用了现代的训练中及训练后技术栈,同时引入了在每个阶段同时保持可塑性和稳定性的保障措施,仅对参数进行最少的高影响力干预。这产生的增益与通常在多个连续模型代际间看到的增益相当,且计算和人员预算远低于普遍认为的水平,使更多主体能够拥有主权人工智能栈的大部分内容(模型、工具基础设施、价值观及数据隐私)。我们通过Thomson展示了这一点,Thomson是一款经过优化以聚焦高风险专业工作的通用前沿模型。Thomson在智能体任务、安全性、法律、税务、多语言能力以及大规模深度研究方面的表现可与近期的前沿模型相媲美。评估显示其呈现出独特的π形模式:在广泛的能力上取得了显著提升,包括那些未被明确针对的能力,同时几乎完全消除了窄域适应中常见的遗忘问题。

英文摘要

The development of frontier models is commonly perceived to be the exclusive remit of a small number of heavily funded players, creating an information, economic and power asymmetry between developers and the diverse user base of modern AI. Recent public discourse acknowledges this concern, calling for SovereignAI (an organisation's capability to independently build, deploy and govern AI use), but offers little concrete advice on how this can be achieved in the short term under a diversity of funding settings. We argue that frontier performance is achievable by a wide range of institutions through Continual Learning on readily available open-weight models. Unlike limited approaches such as small-scale fine-tuning, prompt engineering, or tool-augmentation of a frozen model, our approach exploits a modern mid- & post-training stack while introducing safeguards that preserve both plasticity and stability at each stage, making the minimal number of high-impact interventions on the parameters. This yields gains comparable to those typically seen across multiple successive model generations, at compute and personnel budgets substantially lower than commonly thought, making ownership of large parts of the SovereignAI stack (model, tool infrastructure, values & data privacy) viable for far more actors. We demonstrate this with Thomson, a general-purpose frontier model trained with an enhanced focus on high-stakes professional work. Thomson performs competitively with recent frontier models across agentic tasks, safety, legal, tax & multilingualism, and large-scale Deep Research. Evaluations show a distinctive $π$-shaped pattern: distinct improvements across a wide range of capabilities, including those not explicitly targeted, while almost completely eliminating the forgetting problem common to narrow domain adaptation.

CommentsOpen-weight model: https://huggingface.co/thomsonreuters/Thomson-1.0-Small

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑