arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

AI 大模型

大模型对齐与安全

大模型对齐、安全、越狱、红队、提示注入和可信评测。

2025-10-13 至 2025-10-13 共收录 10 信号源:cs.CL, cs.AI, cs.CY, cs.LG

1. 安全训练 10 篇

2510.09004 2025-10-13 cs.CL 89%

Decoupling Safety into Orthogonal Subspace: Cost-Efficient and Performance-Preserving Alignment for Large Language Models

Yutao Mou, Xiaoling Zhou, Yuxiao Luo, Shikun Zhang, Wei Ye

机构 * National Engineering Research Center for Software Engineering, Peking University, China(软件工程国家工程研究中心,北京大学,中国)

专题命中 安全训练 :alignment(title,abstract);safety(title,abstract);trustworthy(abstract);分类 cs.CL

Comments Work in Progress

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.06944 2025-10-13 cs.LG cs.AI cs.CL cs.CV 67%

AMFT: Aligning LLM Reasoners by Meta-Learning the Optimal Imitation-Exploration Balance

Lixuan He, Jie Feng, Yong Li

专题命中 安全训练 :alignment(abstract);分类 cs.CL、cs.AI、cs.LG

Comments The paper is currently under investigation regarding concerns of potential academic misconduct. While the investigation is ongoing, the authors have voluntarily requested to withdraw the manuscript

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.08776 2025-10-13 cs.CL cs.AI 62%

Measuring Moral LLM Responses in Multilingual Capacities

Kimaya Basu, Savi Kolari, Allison Yu

专题命中 安全训练 :safety(abstract);分类 cs.CL、cs.AI

Comments 10 pages, 5 figures; referenced articles: arXiv:2303.08774, arXiv:2303.12528, arXiv:2308.14132, arXiv:2505.12201, arXiv:2406.04428, arXiv:2407.02273, arXiv:2404.01268, arXiv:2502.09747, arXiv:2507.13474, arXiv:2505.21479, arXiv:2306.05685

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.08648 2025-10-13 cs.LG cs.AI 62%

Inverse-Free Wilson Loops for Transformers: A Practical Diagnostic for Invariance and Order Sensitivity

Edward Y. Chang, Ethan Y. Chang

专题命中 安全训练 :safety(abstract);分类 cs.AI、cs.LG

Comments 24 pages, 10 figures, 2 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.04503 2025-10-13 cs.CR cs.AI cs.CL 62%

P2P: A Poison-to-Poison Remedy for Reliable Backdoor Defense in LLMs

Shuai Zhao, Xinyi Wu, Shiqian Zhao, Xiaobao Wu, Zhongliang Guo, Yanhao Jia, Anh Tuan Luu

机构 * Nanyang Technological University, Singapore(南洋理工大学) Shanghai Jiao Tong University, Shanghai, China(上海交通大学)

专题命中 安全训练 :trustworthy(abstract);分类 cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.11411 2025-10-13 cs.LG 57%

Detecting and Filtering Unsafe Training Data via Data Attribution with Denoised Representation

Yijun Pan, Taiwei Shi, Jieyu Zhao, Jiaqi W. Ma

机构 * University of Michigan(密歇根大学) University of Southern California(南加州大学) University of Illinois Urbana-Champaign(伊利诺伊大学厄巴纳-香槟分校)

专题命中 安全训练 :trustworthy(abstract);分类 cs.LG

Comments 14 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2412.08428 2025-10-13 cs.RO cs.AI cs.SY eess.SY 57%

SwarmGPT: Combining Large Language Models with Safe Motion Planning for Drone Swarm Choreography

Martin Schuck, Dinushka Orrin Dahanaggamaarachchi, Ben Sprenger, Vedant Vyas, Siqi Zhou, Angela P. Schoellig

机构 * Learning Systems and Robotics Lab(学习系统与机器人实验室) Munich Institute of Robotics and Machine Intelligence(慕尼黑机器人与机器智能研究所) Technical University of Munich(慕尼黑技术大学)

专题命中 安全训练 :safety(abstract);分类 cs.AI

Comments Accepted at RA-L 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.09200 2025-10-13 cs.CV cs.AI cs.HC 57%

Towards Safer and Understandable Driver Intention Prediction

Mukilan Karuppasamy, Shankar Gangisetty, Shyam Nandan Rai, Carlo Masone, C V Jawahar

机构 * IIIT Hyderabad(海得拉巴印度理工学院) Politecnico di Torino(托里诺理工学院)

专题命中 安全训练 :safety(abstract);分类 cs.AI

Comments 10 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.13336 2025-10-13 cs.RO cs.LG 57%

Maximizing UAV Cellular Connectivity with Reinforcement Learning for BVLoS Path Planning

Mehran Behjati, Rosdiadee Nordin, Nor Fadzilah Abdullah

专题命中 安全训练 :safety(abstract);分类 cs.LG

Comments Submitted to an IEEE Conference

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.08917 2025-10-13 cs.HC 50%

"I know it's not right, but that's what it said to do": Investigating Trust in AI Chatbots for Cybersecurity Policy

Brandon Lit, Edward Crowder, Daniel Vogel, Hassan Khan

专题命中 安全训练 :prompt injection(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏