arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

AI 大模型

大模型对齐与安全

大模型对齐、安全、越狱、红队、提示注入和可信评测。

2025-10-07 至 2025-10-07 共收录 10 信号源:cs.CL, cs.AI, cs.CY, cs.LG

1. 安全训练 10 篇

2504.20924 2025-10-07 cs.AI 88%

Domain-Agnostic Scalable AI Safety Ensuring Framework

Beomjun Kim, Kangyeon Kim, Sunwoo Kim, Yeonsang Shin, Heejin Ahn

机构 * Massachusetts Institute of Technology(麻省理工学院) Korea Advanced Institute of Science and Technology(韩国科学技术院) Seoul National University(首尔国立大学)

专题命中 安全训练 :safety(title,abstract);AI safety(title,abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.03283 2025-10-07 cs.LG cs.AI cs.CL cs.DC 82%

MACE: A Hybrid LLM Serving System with Colocated SLO-aware Continuous Retraining Alignment

Yufei Li, Yu Fu, Yue Dong, Cong Liu

专题命中 安全训练 :alignment(title,abstract);分类 cs.CL、cs.AI、cs.LG

Comments 14 pages, 15 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.16856 2025-10-07 cs.CV cs.AI 80%

SIA: Enhancing Safety via Intent Awareness for Vision-Language Models

Youngjin Na, Sangheon Jeong, Youngwan Lee, Jian Lee, Dawoon Jeong, Youngman Kim

机构 * VLM Safety LAB, MODULABS(视觉语言模型安全实验室,MODULABS) ETRI(电子技术研究院) KAIST(韩国科学技术院)

专题命中 安全训练 :safety(title,abstract);分类 cs.AI;trustworthy(comments)

Comments Accepted to Safe and Trustworthy Multimodal AI Systems(SafeMM-AI) Workshop at ICCV2025, Non-archival track

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.04392 2025-10-07 cs.CL cs.AI cs.CY cs.LG 70%

Improving Consistency in Retrieval-Augmented Systems with Group Similarity Rewards

Faisal Hamman, Chenyang Zhu, Anoop Kumar, Xujun Peng, Sanghamitra Dutta, Daben Liu, Alfy Samuel

机构 * University of Maryland, College Park(马里兰大学学院公园分校) Capital One

专题命中 安全训练 :safety(abstract);分类 cs.CL、cs.AI、cs.CY

Comments Accepted at NeurIPS 2025 Workshop on Reliable ML from Unreliable Data

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.17601 2025-10-07 cs.CL 70%

Revisiting Backdoor Attacks on LLMs: A Stealthy and Practical Poisoning Framework via Harmless Inputs

Jiawei Kong, Hao Fang, Xiaochen Yang, Kuofeng Gao, Bin Chen, Shu-Tao Xia, Ke Xu, Han Qiu

机构 * Tsinghua Shenzhen International Graduate School, Tsinghua University(清华大学深圳国际研究生院,清华大学) Department of Software Engineering, Harbin Institute of Technology(哈尔滨工业大学软件工程系) School of Computer Science and Technology, Harbin Institute of Technology, Shenzhen(哈尔滨工业大学深圳校区计算机科学与技术学院) Institute for Network Sciences and Cyberspace, Tsinghua University(清华大学网络科学与空间研究院)

专题命中 安全训练 :alignment(abstract);safety(abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.03892 2025-10-07 cs.AI cs.CL 62%

Kantian-Utilitarian XAI: Meta-Explained

Zahra Atf, Peter R. Lewis

机构 * Faculty of Business and Information Technology(商业与信息技术学院) Ontario Tech University(安大略技术大学)

专题命中 安全训练 :alignment(abstract);分类 cs.CL、cs.AI

Comments Accepted for presentation as a poster at the 35th IEEE International Conference on Collaborative Advances in Software and Computing, 2025. Conference website:https://conf.researchr.org/details/cascon-2025/posters-track/1/Kantian-Utilitarian-XAI-Meta-Explained

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.03527 2025-10-07 cs.CL 57%

Sample, Align, Synthesize: Graph-Based Response Synthesis with ConGrs

Sayan Ghosh, Shahzaib Saqib Warraich, Dhruv Tarsadiya, Gregory Yauney, Swabha Swayamdipta

机构 * University of Southern California(南加州大学)

专题命中 安全训练 :alignment(abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.21292 2025-10-07 cs.SE 50%

Semantic Clustering of Civic Proposals: A Case Study on Brazil's National Participation Platform

Ronivaldo Ferreira, Guilherme da Silva, Carla Rocha, Gustavo Pinto

专题命中 安全训练 :alignment(abstract)

Comments 12 pages, in Portuguese language

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.04076 2025-10-07 cs.RO cs.SY eess.SY 50%

From Shadow to Light: Toward Safe and Efficient Policy Learning Across MPC, DeePC, RL, and LLM Agents

Amin Vahidi-Moghaddam, Sayed Pedram Haeri Boroujeni, Iman Jebellat, Ehsan Jebellat, Niloufar Mehrabi, Zhaojian Li

专题命中 安全训练 :safety(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2409.00882 2025-10-07 cs.SE 50%

SAFE: Advancing Large Language Models in Leveraging Semantic and Syntactic Relationships for Software Vulnerability Detection

Van Nguyen, Surya Nepal, Tingmin Wu, Xingliang Yuan, Carsten Rudolph

专题命中 安全训练 :safety(abstract)

Journal ref Proceedings of the 20th ACM Asia Conference on Computer and Communications Security (ASIA CCS), 2025

详情

展开后加载摘要…

URL PDF HTML 收藏