arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

大模型对齐与安全

大模型对齐、安全、越狱、红队、提示注入和可信评测。

2025-10-31 至 2025-10-31 共收录 33 信号源:cs.CL, cs.AI, cs.CY, cs.LG

1. 其他安全 9 篇

2510.23127 2025-10-31 cs.AI 57%

Lost in Tokenization: Context as the Key to Unlocking Biomolecular Understanding in Scientific LLMs

Kai Zhuang, Jiawei Zhang, Yumou Liu, Hanqun Cao, Chunbin Gu, Mengdi Liu, Zhangyang Gao, Zitong Jerry Wang, Xuanhe Zhou, Pheng-Ann Heng, Lijun Wu, Conghui He, Cheng Tan

机构 * Shanghai Artificial Intelligence Laboratory(上海人工智能实验室) Westlake University(西交大学) Shanghai Innovation Institute(上海创新研究院) Shanghai Jiaotong University(上海交通大学) The Chinese University of Hong Kong(香港中文大学) Institute of Computing Technology, Chinese Academy of Sciences(中国科学院计算技术研究所)

专题命中 其他安全 :alignment(abstract);分类 cs.AI

Comments 38 pages, under review

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.16629 2025-10-31 cs.LG 57%

On the Impossibility of Retrain Equivalence in Machine Unlearning

Jiatong Yu, Yinghui He, Anirudh Goyal, Sanjeev Arora

机构 * Princeton Language and Intelligence(普林斯顿语言与智能)

专题命中 其他安全 :alignment(abstract);分类 cs.LG

Comments Code available at https://princeton-pli.github.io/impossibility-unlearning/

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.15811 2025-10-31 cs.LG 57%

On the creation of narrow AI: hierarchy and nonlocality of neural network skills

Eric J. Michaud, Asher Parker-Sartori, Max Tegmark

机构 * Department of Physics, Massachusetts Institute of Technology(物理学系,麻省理工学院) Department of EECS, Massachusetts Institute of Technology(电子工程与计算机科学系,麻省理工学院) The NSF AI Institute for Artificial Intelligence and Fundamental Interactions(国家科学基金会人工智能与基本相互作用研究所)

专题命中 其他安全 :safety(abstract);分类 cs.LG

Comments NeurIPS 2025; 20 pages, 13 figures

详情

展开后加载摘要…

URL PDF HTML 收藏