arXivDaily arXiv每日学术速递 周一至周五更新

高校专区

University of Cambridge(剑桥大学)

2026-01-16 至 2026-01-16 共收录 3
2601.09421 2026-01-16 cs.CL cs.AI

Bias Dynamics in BabyLMs: Towards a Compute-Efficient Sandbox for Democratising Pre-Training Debiasing

BabyLMs中的偏见动态:朝着更高效的计算 sandbox 以民主化预训练去偏研究

Filip Trhlik, Andrew Caines, Paula Buttery

机构 * Department of Computer Science & Technology, University of Cambridge(计算机科学与技术系,剑桥大学) ALTA Institute, University of Cambridge(ALTA研究所,剑桥大学)

AI总结 通过低成本BabyLMs研究偏见动态,降低预训练成本,促进公平语言模型的民主化研究。

Comments 21 pages, 18 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.22310 2026-01-16 cs.LG cs.AI cs.CV

From Dormant to Deleted: Tamper-Resistant Unlearning Through Weight-Space Regularization

从沉睡到删除:通过权重空间正则化实现抗篡改的遗忘

Shoaib Ahmed Siddiqui, Adrian Weller, David Krueger, Gintare Karolina Dziugaite, Michael Curtis Mozer, Eleni Triantafillou

机构 * University of Cambridge(剑桥大学) The Alan Turing Institute(艾伦·图灵研究所) Mila(Mila研究所) Google DeepMind(谷歌DeepMind)

AI总结 通过权重空间正则化方法,提升大型语言模型对重新学习攻击的抗性,实现从沉睡到删除的高效遗忘机制。

详情

展开后加载摘要…

URL PDF HTML 收藏
2410.03055 2026-01-16 cs.LG cs.AI

Permissive Information-Flow Analysis for Large Language Models

宽松的信息流分析用于大型语言模型

Shoaib Ahmed Siddiqui, Radhika Gaonkar, Boris Köpf, David Krueger, Andrew Paverd, Ahmed Salem, Shruti Tople, Lukas Wutschitz, Menglin Xia, Santiago Zanella-Béguelin

机构 * University of Cambridge(剑桥大学) Microsoft(微软公司) Mila

AI总结 本文提出了一种更宽松的信息流分析方法,通过传播对模型输出有影响的样本标签来提高大型语言模型的安全性和隐私保护效果。

详情

展开后加载摘要…

URL PDF HTML 收藏