arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

大模型对齐与安全

大模型对齐、安全、越狱、红队、提示注入和可信评测。

2025-10-28 至 2025-10-28 共收录 84 信号源:cs.CL, cs.AI, cs.CY, cs.LG

1. 偏好对齐 15 篇

2510.21794 2025-10-28 cs.CV cs.AI 83%

Token-Level Inference-Time Alignment for Vision-Language Models

Kejia Chen, Jiawen Zhang, Jiacong Hu, Kewei Gao, Jian Lou, Zunlei Feng, Mingli Song

机构 * Zhejiang University(浙江大学) Sun Yat-sen University(中山大学)

专题命中 偏好对齐 :alignment(title,abstract);DPO(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.05465 2025-10-28 cs.CL cs.AI cs.LG 82%

ComPO: Preference Alignment via Comparison Oracles

Peter Chen, Xi Chen, Wotao Yin, Tianyi Lin

机构 * Columbia University(哥伦比亚大学) Stern School of Business(斯特恩商学院) New York University(纽约大学) DAMO Academy, Alibaba Group US(阿里云达摩院)

专题命中 偏好对齐 :alignment(title,abstract);分类 cs.CL、cs.AI、cs.LG

Comments Accepted to NeurIPS 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.07193 2025-10-28 cs.LG stat.ML 79%

Provably Efficient Online RLHF with One-Pass Reward Modeling

Long-Fei Li, Yu-Yang Qian, Peng Zhao, Zhi-Hua Zhou

机构 * National Key Laboratory for Novel Software Technology, Nanjing University, China(国家新型软件技术实验室,南京大学) School of Artificial Intelligence, Nanjing University, China(人工智能学院,南京大学)

专题命中 偏好对齐 :RLHF(title,abstract);分类 cs.LG

Comments NeurIPS 2025; The first two authors contributed equally

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.19158 2025-10-28 cs.CL cs.AI 73%

When Personalization Meets Reality: A Multi-Faceted Analysis of Personalized Preference Learning

Yijiang River Dong, Tiancheng Hu, Yinhong Liu, Ahmet Üstün, Nigel Collier

机构 * University of Cambridge(剑桥大学) Cohere For AI

专题命中 偏好对齐 :RLHF(abstract);safety(abstract);分类 cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.07426 2025-10-28 cs.LG cs.AI 73%

RePO: Understanding Preference Learning Through ReLU-Based Optimization

Junkang Wu, Kexin Huang, Xue Wang, Jinyang Gao, Bolin Ding, Jiancan Wu, Xiangnan He, Xiang Wang

机构 * University of Science and Technology of China(中国科学技术大学) Alibaba Group(阿里巴巴集团) Institute of Dataspace, Hefei Comprehensive National Science Center(数据空间研究所,合肥综合性国家科学中心) MoE Key Lab of BIPC, University of Science and Technology of China(BIPC联合实验室,中国科学技术大学)

专题命中 偏好对齐 :RLHF(abstract);DPO(abstract);分类 cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.13177 2025-10-28 cs.LG cs.AI 73%

KL Penalty Control via Perturbation for Direct Preference Optimization

Sangkyu Lee, Janghoon Han, Hosung Song, Stanley Jungkyu Choi, Honglak Lee, Youngjae Yu

专题命中 偏好对齐 :alignment(abstract);DPO(abstract);分类 cs.AI、cs.LG

Comments Published as a main conference track paper at NeurIPS 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.23590 2025-10-28 cs.LG 70%

Lightweight Robust Direct Preference Optimization

Cheol Woo Kim, Shresth Verma, Mauricio Tec, Milind Tambe

机构 * School of Engineering and Applied Sciences, Harvard University(哈佛大学工程与应用科学学院)

专题命中 偏好对齐 :alignment(abstract);DPO(abstract);分类 cs.LG

Comments arXiv admin note: substantial text overlap with arXiv:2509.02709

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.22954 2025-10-28 cs.CL 70%

Artificial Hivemind: The Open-Ended Homogeneity of Language Models (and Beyond)

Liwei Jiang, Yuanjun Chai, Margaret Li, Mickel Liu, Raymond Fok, Nouha Dziri, Yulia Tsvetkov, Maarten Sap, Alon Albalak, Yejin Choi

专题命中 偏好对齐 :safety(abstract);AI safety(abstract);分类 cs.CL

Comments NeurIPS 2025 D&B Paper (Oral); Camera-Ready Version

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.22084 2025-10-28 cs.CL 70%

Compositional Bias Control in Large Language Models: Preference Learning Fails, Supervision Succeeds

Atij Mahesh

机构 * University of California, Los Angeles(加州大学洛杉矶分校)

专题命中 偏好对齐 :alignment(abstract);DPO(abstract);分类 cs.CL

Comments 20 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.19601 2025-10-28 cs.LG cs.AI cs.CL 67%

Preference Optimization by Estimating the Ratio of the Data Distribution

Yeongmin Kim, Heesun Bae, Byeonghu Na, Il-Chul Moon

机构 * Korea Advanced Institute of Science and Technology (KAIST)(韩国科学技术院)

专题命中 偏好对齐 :DPO(abstract);分类 cs.CL、cs.AI、cs.LG

Comments NeurIPS 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.10637 2025-10-28 cs.CL cs.AI cs.LG 67%

Better Estimation of the Kullback--Leibler Divergence Between Language Models

Afra Amini, Tim Vieira, Ryan Cotterell

机构 * ETH Zürich(苏黎世联邦理工学院)

专题命中 偏好对齐 :RLHF(abstract);分类 cs.CL、cs.AI、cs.LG

Comments NeurIPS 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.23217 2025-10-28 cs.CL cs.AI 62%

Process Reward Models for Sentence-Level Verification of LVLM Radiology Reports

Alois Thomas, Maya Varma, Jean-Benoit Delbrouck, Curtis P. Langlotz

机构 * AIMI Center Stanford University(AIMI中心 斯坦福大学)

专题命中 偏好对齐 :safety(abstract);分类 cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.21847 2025-10-28 cs.LG 57%

SynCast: Synergizing Contradictions in Precipitation Nowcasting via Diffusion Sequential Preference Optimization

Kaiyi Xu, Junchao Gong, Wenlong Zhang, Ben Fei, Lei Bai, Wanli Ouyang

机构 * University of Science and Technology of China(中国科学技术大学) Shanghai Artificial Intelligence Laboratory(上海人工智能实验室) Shanghai Jiao Tong University(上海交通大学) Chinese University of Hong Kong(香港中文大学)

专题命中 偏好对齐 :alignment(abstract);分类 cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.05818 2025-10-28 cs.AI 57%

Chatbot To Help Patients Understand Their Health

Won Seok Jang, Hieu Tran, Manav Mistry, SaiKiran Gandluri, Yifan Zhang, Sharmin Sultana, Sunjae Kown, Yuan Zhang, Zonghai Yao, Hong Yu

机构 * Center for Healthcare Organization and Implementation Research, VA Bedford Health Care(VA贝德福德医疗中心健康组织与实施研究中心) Miner School of Computer and Information Sciences, University of Massachusetts Lowell(马萨诸塞大学洛厄尔分校计算机与信息科学学院) Manning College of Information and Computer Sciences, University of Massachusetts Amherst(马萨诸塞大学阿姆赫斯特分校信息与计算机科学学院)

专题命中 偏好对齐 :alignment(abstract);分类 cs.AI

Comments Accepted in EMNLP 2025 Findings

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.21723 2025-10-28 cs.HC 50%

Recognizing internal states in AI: evidence from patterned preferences in large language models

Annika Hedberg

专题命中 偏好对齐 :alignment(abstract)

Comments 14 pages, 1 table. Includes test protocol transcript and multi-system response data

详情

展开后加载摘要…

URL PDF HTML 收藏

2. 安全训练 11 篇

2510.14301 2025-10-28 cs.AI 83%

A Guardrail for Safety Preservation: When Safety-Sensitive Subspace Meets Harmful-Resistant Null-Space

Bingjie Zhang, Yibo Yang, Zhe Ren, Dandan Guo, Jindong Gu, Philip Torr, Bernard Ghanem

机构 * School of Artificial Intelligence, Jilin University(吉林大学人工智能学院) King Abdullah University of Science and Technology(卡塔尔国王大学科学与技术研究院) University of Oxford(牛津大学)

专题命中 安全训练 :safety(title,abstract);alignment(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.00971 2025-10-28 cs.LG cs.AI 81%

Reasoning as an Adaptive Defense for Safety

Taeyoun Kim, Fahim Tajwar, Aditi Raghunathan, Aviral Kumar

机构 * Carnegie Mellon University(卡内基梅隆大学)

专题命中 安全训练 :safety(title,abstract);分类 cs.AI、cs.LG

Comments 44 pages, 10 Figures, 7 Tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.12520 2025-10-28 cs.CV 78%

SafeEraser: Enhancing Safety in Multimodal Large Language Models through Multimodal Machine Unlearning

Junkai Chen, Zhijie Deng, Kening Zheng, Yibo Yan, Shuliang Liu, PeiJun Wu, Peijie Jiang, Jia Liu, Xuming Hu

机构 * The Hong Kong University of Science and Technology (Guangzhou)(香港科技大学(广州)) The Hong Kong University of Science and Technology(香港科技大学) Southeast University(东南大学) Ant Group, Alibaba(蚂蚁集团)

专题命中 安全训练 :safety(title,abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.23659 2025-10-28 cs.CL cs.AI 62%

Aligning LLMs for Multilingual Consistency in Enterprise Applications

Amit Agarwal, Hansa Meghwani, Hitesh Laxmichand Patel, Tao Sheng, Sujith Ravi, Dan Roth

机构 * Oracle AI

专题命中 安全训练 :alignment(abstract);分类 cs.CL、cs.AI

Comments Accepted at EMNLP 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.17873 2025-10-28 cs.CL cs.AI cs.CE 62%

MOOSE-Chem3: Toward Experiment-Guided Hypothesis Ranking via Simulated Experimental Feedback

Wanhao Liu, Zonglin Yang, Jue Wang, Lidong Bing, Di Zhang, Dongzhan Zhou, Yuqiang Li, Houqiang Li, Erik Cambria, Wanli Ouyang

机构 * University of Science and Technology of China(中国科学技术大学) Shanghai Artificial Intelligence Laboratory(上海人工智能实验室) Nanyang Technological University(南洋理工大学) MiroMind

专题命中 安全训练 :alignment(abstract);分类 cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.19599 2025-10-28 cs.AI cs.LG 62%

GVPO: Group Variance Policy Optimization for Large Language Model Post-Training

Kaichen Zhang, Yuzhong Hong, Junwei Bao, Hongfei Jiang, Yang Song, Dingqian Hong, Hui Xiong

机构 * Thrust of Artificial Intelligence, Hong Kong University of Science and Technology (Guangzhou)(香港科技大学(广州)人工智能研究所) Zuoyebang Education Technology(佐业邦教育科技) Department of Computer Science and Engineering, HKUST(香港科技大学计算机科学与工程系)

专题命中 安全训练 :alignment(abstract);分类 cs.AI、cs.LG

Comments Accepted by NeurIPS 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.21885 2025-10-28 cs.CL cs.AI 62%

Preventing Catastrophic Forgetting: Behavior-Aware Sampling for Safer Language Model Fine-Tuning

Anh Pham, Mihir Thalanki, Michael Sun, Aditya Chaloo, Ankita Gupta, Tian Xia, Aditya Mate, Ehimwenma Nosakhare, Soundararajan Srinivasan

专题命中 安全训练 :safety(abstract);分类 cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.18253 2025-10-28 cs.RO cs.AI 57%

Depth-Constrained ASV Navigation with Deep RL and Limited Sensing

Amirhossein Zhalehmehrabi, Daniele Meli, Francesco Dal Santo, Francesco Trotti, Alessandro Farinelli

机构 * Department of Computer Science, University of Verona(计算机科学系,威尼斯大学)

专题命中 安全训练 :safety(abstract);分类 cs.AI

Comments 8 pages, 8 figures, Accepted to IEEE Robotics and Automation Letters (this is not the final version)

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.21853 2025-10-28 cs.CL 57%

Policy Optimization Prefers The Path of Least Resistance

Debdeep Sanyal, Aakash Sen Sharma, Dhruv Kumar, Saurabh Deshpande, Murari Mandal

机构 * Birla AI Labs(Birla人工智能实验室) InvideoAI BITS Pilani(比斯·皮尔尼学院) Kalinga Institute of Industrial Technology(卡林加工业技术学院)

专题命中 安全训练 :alignment(abstract);分类 cs.CL

Comments 21 pages, 8 figures, 2 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.04654 2025-10-28 cs.AI 57%

E-bike agents: Large Language Model-Driven E-Bike Accident Analysis and Severity Prediction

Zhichao Yang, Jiashu He, Mohammad B. Al-Khasawneh, Darshan Pandit, Cirillo Cinzia

机构 * Civil and Environmental Engineering, University of Maryland(大学工程学院) Computer and Information Science, University of Pennsylvania(大学计算机与信息科学学院)

专题命中 安全训练 :safety(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.13939 2025-10-28 cs.CV 50%

Med-R1: Reinforcement Learning for Generalizable Medical Reasoning in Vision-Language Models

Yuxiang Lai, Jike Zhong, Ming Li, Shitian Zhao, Yuheng Li, Konstantinos Psounis, Xiaofeng Yang

机构 * Department of Computer Science and Informatics, Emory University(计算机科学与信息学系,埃默里大学) Department of Computer Science and Department of Electrical and Computer Engineering, University of Southern California(计算机科学系和电气与计算机工程系,南加州大学) Department of Computer Science, University of Tokyo(计算机科学系,东京大学) Department of Computer Science, Johns Hopkins University(计算机科学系,约翰霍普金斯大学) Department of Biomedical Engineering, Georgia Institute of Technology and Emory University(生物医学工程系,佐治亚理工学院和埃默里大学)

专题命中 安全训练 :alignment(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏

3. 越狱攻击 5 篇

2510.22085 2025-10-28 cs.CR cs.AI cs.CL cs.LG 87%

Jailbreak Mimicry: Automated Discovery of Narrative-Based Jailbreaks for Large Language Models

Pavlos Ntais

机构 * University of Athens(雅典大学)

专题命中 越狱攻击 :jailbreak(title,abstract);alignment(abstract);safety(abstract);分类 cs.CL、cs.AI、cs.LG

Comments 18 pages, 5 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.10281 2025-10-28 cs.CR cs.AI cs.CL cs.CV cs.LG 85%

ArtPerception: ASCII Art-based Jailbreak on LLMs with Recognition Pre-test

Guan-Yan Yang, Tzu-Yu Cheng, Ya-Wen Teng, Farn Wanga, Kuo-Hui Yeh

机构 * Department of Electrical Engineering, National Taiwan University(国立台湾大学电子工程系) GARMIN (ASIA) CORPORATION(GARMIN(亚洲)公司) Institute of Artificial Intelligence Innovation, National Yang Ming Chiao Tung University(国家阳明交通大学人工智能创新研究所) Department of Information Management, National Dong Hwa University(国立东吴大学资讯管理系)

专题命中 越狱攻击 :jailbreak(title,abstract);safety(abstract);分类 cs.CL、cs.AI、cs.LG

Comments 30 pages, 22 figures. This preprint has been accepted for publication in Elsevier JOURNAL OF NETWORK AND COMPUTER APPLICATIONS (JNCA)

Journal ref Journal of Network and Computer Applications, Vol. 244, (2025) 104356

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.21983 2025-10-28 cs.CL cs.AI 79%

Uncovering the Persuasive Fingerprint of LLMs in Jailbreaking Attacks

Havva Alizadeh Noughabi, Julien Serbanescu, Fattane Zarrinkalam, Ali Dehghantanha

机构 * Cyber Science Lab, University of Guelph(圭尔夫大学网络安全实验室) College of Engineering, University of Guelph(圭尔夫大学工程学院)

专题命中 越狱攻击 :alignment(abstract);safety(abstract);jailbreak(abstract);分类 cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.03728 2025-10-28 cs.AI cs.HC 57%

PersonaTeaming: Exploring How Introducing Personas Can Improve Automated AI Red-Teaming

Wesley Hanwen Deng, Sunnie S. Y. Kim, Akshita Jha, Ken Holstein, Motahhare Eslami, Lauren Wilcox, Leon A Gatys

机构 * Carnegie Mellon University(卡内基梅隆大学) Apple(苹果公司)

专题命中 越狱攻击 :safety(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏