arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

大模型对齐与安全

大模型对齐、安全、越狱、红队、提示注入和可信评测。

2025-07-29 至 2025-07-29 共收录 57 信号源:cs.CL, cs.AI, cs.CY, cs.LG

1. 偏好对齐 11 篇

2504.02193 2025-07-29 cs.AI 92%

More is Less: The Pitfalls of Multi-Model Synthetic Preference Data in DPO Safety Alignment

Yifan Wang, Runjin Chen, Bolian Li, David Cho, Yihe Deng, Ruqi Zhang, Tianlong Chen, Zhangyang Wang, Ananth Grama, Junyuan Hong

机构 * Purdue University(普渡大学) The University of Texas at Austin(德克萨斯大学奥斯汀分校) The University of North Carolina at Chapel Hill(北卡罗来纳大学教堂山分校) University of California, Los Angeles(加州大学洛杉矶分校)

专题命中 偏好对齐 :alignment(title,abstract);DPO(title,abstract);safety(title,abstract);RLHF(abstract)

Comments This version includes updated results and expanded discussion

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.19672 2025-07-29 cs.AI cs.LG stat.ML 88%

Alignment and Safety in Large Language Models: Safety Mechanisms, Training Paradigms, and Emerging Challenges

Haoran Lu, Luyang Fang, Ruidong Zhang, Xinliang Li, Jiazhang Cai, Huimin Cheng, Lin Tang, Ziyu Liu, Zeliang Sun, Tao Wang, Yingchuan Zhang, Arif Hassan Zidan, Jinwen Xu, Jincheng Yu, Meizhi Yu, Hanqi Jiang, Xilin Gong, Weidi Luo, Bolun Sun, Yongkai Chen, Terry Ma, Shushan Wu, Yifan Zhou, Junhao Chen, Haotian Xiang, Jing Zhang, Afrar Jahin, Wei Ruan, Ke Deng, Yi Pan, Peilong Wang, Jiahui Li, Zhengliang Liu, Lu Zhang, Lin Zhao, Wei Liu, Dajiang Zhu, Xin Xing, Fei Dou, Wei Zhang, Chao Huang, Rongjie Liu, Mengrui Zhang, Yiwen Liu, Xiaoxiao Sun, Qin Lu, Zhen Xiang, Wenxuan Zhong, Tianming Liu, Ping Ma

机构 * Department of Statistics, University of Georgia(统计学系,佐治亚大学) School of Computing, University of Georgia(计算学院,佐治亚大学) Department of Biostatistics, Boston University(生物统计学系,波士顿大学) Department of Epidemiology & Biostatistics, University of Georgia(流行病学与生物统计学系,佐治亚大学) School of Computer and Cyber Sciences, Augusta University(计算机与网络科学学院,奥古斯塔大学) School of Electrical and Computer Engineering, University of Georgia(电气与计算机工程学院,佐治亚大学) Department of Statistics & Data Science, University of Arizona(统计学与数据科学系,亚利桑那大学) Kellogg School of Management, Northwestern University(凯洛格管理学院,西北大学) Department of Statistics, Harvard University(统计学系,哈佛大学) School of Computer Science, Carnegie Mellon University(计算机科学学院,卡内基梅隆大学)

专题命中 偏好对齐 :alignment(title,abstract);safety(title);DPO(abstract);分类 cs.AI、cs.LG

Comments 119 pages, 10 figures, 7 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.20335 2025-07-29 cs.LG cs.AI 81%

Cultivating Helpful, Personalized, and Creative AI Tutors: A Framework for Pedagogical Alignment using Reinforcement Learning

Siyu Song, Wentao Liu, Ye Lu, Ruohua Zhang, Tao Liu, Jinze Lv, Xinyun Wang, Aimin Zhou, Fei Tan, Bo Jiang, Hao Hao

机构 * Shanghai Innavation Institute(上海创新研究院) Shanghai Institute of AI for Education(上海人工智能教育研究院) School of Computer Science and Technology(计算机科学与技术学院) Department of Educational Information Technology(教育信息技术系)

专题命中 偏好对齐 :alignment(title,abstract);分类 cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.12854 2025-07-29 cs.CL 79%

Enhancing LLM Reasoning with Iterative DPO: A Comprehensive Empirical Investigation

Songjun Tu, Jiahao Lin, Xiangyu Tian, Qichao Zhang, Linjing Li, Yuqian Fu, Nan Xu, Wei He, Xiangyuan Lan, Dongmei Jiang, Dongbin Zhao

机构 * Institute of Automation, Chinese Academy of Sciences(中国科学院自动化研究所) Pengcheng Laboratory(鹏城实验室) School of Artificial Intelligence, University of Chinese Academy of Sciences(中国科学院大学人工智能学院) Wenge Technology(文生科技) Fudan University(复旦大学)

专题命中 偏好对齐 :DPO(title,abstract);分类 cs.CL

Comments 23pages

Journal ref COLM2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.20181 2025-07-29 cs.CL cs.AI 73%

SGPO: Self-Generated Preference Optimization based on Self-Improver

Hyeonji Lee, Daejin Jo, Seohwan Yun, Sungwoong Kim

机构 * Korea University(韩国大学)

专题命中 偏好对齐 :alignment(abstract);DPO(abstract);分类 cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.21560 2025-07-29 cs.CL cs.AI 73%

Reinforcement learning fine-tuning of language model for instruction following and math reasoning

Yifu Han, Geo Zhang

机构 * Stanford University(斯坦福大学)

专题命中 偏好对齐 :alignment(abstract);DPO(abstract);分类 cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2406.00222 2025-07-29 cs.CL cs.AI cs.LG 67%

Learning to Clarify: Multi-turn Conversations with Action-Based Contrastive Self-Training

Maximillian Chen, Ruoxi Sun, Tomas Pfister, Sercan Ö. Arık

机构 * Google(谷歌) Columbia University(哥伦比亚大学)

专题命中 偏好对齐 :DPO(abstract);分类 cs.CL、cs.AI、cs.LG

Comments ICLR 2025; Code: https://github.com/google-research/google-research/tree/master/learning_to_clarify

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.20956 2025-07-29 cs.CL cs.AI 62%

Mind the Gap: Conformative Decoding to Improve Output Diversity of Instruction-Tuned Large Language Models

Max Peeperkorn, Tom Kouwenhoven, Dan Brown, Anna Jordanous

机构 * School of Computing University of Kent(肯特大学计算机学院) Leiden Institute of Advanced Computer Science Universiteit Leiden(莱顿先进计算机科学研究所) Cheriton School of Computer Science University of Waterloo(滑铁卢大学查里顿计算机科学学院)

专题命中 偏好对齐 :DPO(abstract);分类 cs.CL、cs.AI

Comments 9 pages, 3 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2501.06645 2025-07-29 cs.CL cs.AI 62%

FocalPO: Enhancing Preference Optimizing by Focusing on Correct Preference Rankings

Tong Liu, Xiao Yu, Wenxuan Zhou, Jindong Gu, Volker Tresp

专题命中 偏好对齐 :DPO(abstract);分类 cs.CL、cs.AI

Comments ACL 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.20000 2025-07-29 cs.AI cs.DL 57%

Matching Game Preferences Through Dialogical Large Language Models: A Perspective

Renaud Fabre, Daniel Egret, Patrice Bellot

专题命中 偏好对齐 :trustworthy(abstract);分类 cs.AI

Comments 28 pages, 1 figure. Published in Applied Sciences

Journal ref Applied Sciences, 2025, 15(15), 8307

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.19679 2025-07-29 cs.CV cs.AI 57%

Efficient Learning for Product Attributes with Compact Multimodal Models

Mandar Kulkarni

机构 * Flipkart Data Science(Flipkart数据科学)

专题命中 偏好对齐 :DPO(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏

2. 安全训练 4 篇

2507.20150 2025-07-29 cs.AI cs.CL cs.LG 83%

The Policy Cliff: A Theoretical Analysis of Reward-Policy Maps in Large Language Models

Xingcheng Xu

机构 * Shanghai Artificial Intelligence Laboratory(上海人工智能实验室)

专题命中 安全训练 :alignment(abstract);RLHF(abstract);safety(abstract);trustworthy(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.17275 2025-07-29 eess.SY cs.AI cs.LG cs.SY 81%

Conformal Safety Shielding for Imperfect-Perception Agents

William Scarbro, Calum Imrie, Sinem Getir Yaman, Kavan Fatehi, Corina S. Pasareanu, Radu Calinescu, Ravi Mangal

机构 * Colorado State University, USA(科罗拉多州立大学) University of York, UK(约克大学) Carnegie Mellon University, USA(卡内基梅隆大学)

专题命中 安全训练 :safety(title,abstract);分类 cs.AI、cs.LG

Comments 32 pages; Equal contribution by W. Scarbro and C. Imrie; Accepted at 25th International Conference on Runtime Verification, 2025 (RV25)

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.20685 2025-07-29 eess.SY cs.SY 78%

What's Really Different with AI? -- A Behavior-based Perspective on System Safety for Automated Driving Systems

Marcus Nolte, Nayel Fabian Salem, Olaf Franke, Jan Heckmann, Christoph Höhmann, Georg Stettinger, Markus Maurer

专题命中 安全训练 :safety(title,abstract)

Comments 8 pages, 1 figure, 1 table, to be published in 2025 IEEE International Automated Vehicle Validation Conference (IAVVC)

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.20614 2025-07-29 cs.CL 57%

Before the Outrage: Challenges and Advances in Predicting Online Antisocial Behavior

Anaïs Ollagnier

专题命中 安全训练 :safety(abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏

3. 越狱攻击 4 篇

2507.20333 2025-07-29 cs.AI cs.LG stat.ML 88%

The Blessing and Curse of Dimensionality in Safety Alignment

Rachel S. Y. Teo, Laziz U. Abdullaev, Tan M. Nguyen

机构 * Department of Mathematics National University of Singapore(数学系新加坡国立大学)

专题命中 越狱攻击 :alignment(title,abstract);safety(title,abstract);分类 cs.AI、cs.LG

Comments Published as a conference paper at COLM 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.15606 2025-07-29 cs.LG cs.AI cs.CL 85%

LoX: Low-Rank Extrapolation Robustifies LLM Safety Against Fine-tuning

Gabriel J. Perin, Runjin Chen, Xuxi Chen, Nina S. T. Hirata, Zhangyang Wang, Junyuan Hong

机构 * University of São Paulo(圣保罗大学) University of Texas at Austin(德克萨斯大学奥斯汀分校)

专题命中 越狱攻击 :safety(title,abstract);alignment(abstract);分类 cs.CL、cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.19880 2025-07-29 cs.CR cs.AI 57%

Trivial Trojans: How Minimal MCP Servers Enable Cross-Tool Exfiltration of Sensitive Data

Nicola Croce, Tobin South

机构 * Pivotal Research(Pivotal研究机构) Stanford University(斯坦福大学)

专题命中 越狱攻击 :trustworthy(abstract);分类 cs.AI

Comments Abstract submitted to the Technical AI Governance Forum 2025 (https://www.techgov.ai/)

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.19609 2025-07-29 cs.CR 50%

Securing the Internet of Medical Things (IoMT): Real-World Attack Taxonomy and Practical Security Measures

Suman Deb, Emil Lupu, Emm Mic Drakakis, Anil Anthony Bharath, Zhen Kit Leung, Guang Rui Ma, Anupam Chattopadhyay

专题命中 越狱攻击 :safety(abstract)

Comments Submitted as a book chapter in 'Handbook of Industrial Internet of Things' to be published by Springer Nature

详情

展开后加载摘要…

URL PDF HTML 收藏

4. 提示注入 1 篇

2504.14348 2025-07-29 cs.CV 85%

Manipulating Multimodal Agents via Cross-Modal Prompt Injection

Le Wang, Zonghao Ying, Tianyuan Zhang, Siyuan Liang, Shengshan Hu, Mingchuan Zhang, Aishan Liu, Xianglong Liu

机构 * Beihang University(北洋大学) National University of Singapore(新加坡国立大学) Huazhong University of Science and Technology(华中科技大学) Henan University of Science and Technology(河南科技大学)

专题命中 提示注入 :prompt injection(title,abstract);alignment(abstract);safety(abstract)

Comments 16 pages, 5 figures

详情

展开后加载摘要…

URL PDF HTML 收藏

5. 幻觉与事实性 1 篇

2507.19548 2025-07-29 cs.CY cs.AI 81%

Justifications for Democratizing AI Alignment and Their Prospects

André Steingrüber, Kevin Baum

机构 * German Research Center for Artificial Intelligence (DFKI)(德国人工智能研究中心) Center for European Research in Trusted Artificial Intelligence (CERTAIN)(欧洲可信人工智能研究中心)

专题命中 幻觉与事实性 :alignment(title,abstract);分类 cs.AI、cs.CY

Comments accepted for the LNCS on-site proceedings of the AISoLA 2025 conference

详情

展开后加载摘要…

URL PDF HTML 收藏

6. 隐私与版权 1 篇

2502.15278 2025-07-29 cs.CV cs.AI 57%

CopyJudge: Automated Copyright Infringement Identification and Mitigation in Text-to-Image Diffusion Models

Shunchang Liu, Zhuan Shi, Lingjuan Lyu, Yaochu Jin, Boi Faltings

机构 * EPFL(苏黎世联邦理工学院) Mila - Quebec AI Institute Mcgill University(蒙特利尔麦吉尔大学人工智能研究所) Westlake University(西湖大学)

专题命中 隐私与版权 :alignment(abstract);分类 cs.AI

Comments Accepted by ACM MM 2025

详情

展开后加载摘要…

URL PDF HTML 收藏

7. 安全评测 16 篇

2507.20526 2025-07-29 cs.AI cs.CL cs.CY 67%

Security Challenges in AI Agent Deployment: Insights from a Large Scale Public Competition

Andy Zou, Maxwell Lin, Eliot Jones, Micha Nowak, Mateusz Dziemian, Nick Winter, Alexander Grattan, Valent Nathanael, Ayla Croft, Xander Davies, Jai Patel, Robert Kirk, Nate Burnikell, Yarin Gal, Dan Hendrycks, J. Zico Kolter, Matt Fredrikson

专题命中 安全评测 :red teaming(abstract);分类 cs.CL、cs.AI、cs.CY

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.19598 2025-07-29 cs.CL cs.AI cs.CR cs.LG 67%

MOCHA: Are Code Language Models Robust Against Multi-Turn Malicious Coding Prompts?

Muntasir Wahed, Xiaona Zhou, Kiet A. Nguyen, Tianjiao Yu, Nirav Diwan, Gang Wang, Dilek Hakkani-Tür, Ismini Lourentzou

机构 * University of Illinois Urbana-Champaign(伊利诺伊大学厄巴纳-香槟分校)

专题命中 安全评测 :safety(abstract);分类 cs.CL、cs.AI、cs.LG

Comments Winner Defender Team at Amazon Nova AI Challenge 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.20280 2025-07-29 cs.AI cs.CL 62%

SciToolAgent: A Knowledge Graph-Driven Scientific Agent for Multi-Tool Integration

Keyan Ding, Jing Yu, Junjie Huang, Yuchen Yang, Qiang Zhang, Huajun Chen

机构 * College of Computer Science and Technology(计算机科学与技术学院) Zhejiang University(浙江大学) Zhejiang Key Laboratory of Intelligent Manufacturing for Functional Chemicals(功能化学品智能制造重点实验室) ZJU-Hangzhou Global Scientific and Technological Innovation Center(浙大杭州国际科学技术创新中心) ZJU-UIUC Institute(浙大UIUC研究院) The Polytechnic Institute(技术学院) State Key Laboratory of Ocean Sensing(海洋感知国家重点实验室)

专题命中 安全评测 :safety(abstract);分类 cs.CL、cs.AI

Comments 21 pages, 6 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.16802 2025-07-29 cs.CL cs.LG 62%

Agentar-Fin-R1: Enhancing Financial Intelligence through Domain Expertise, Training Efficiency, and Advanced Reasoning

Yanjun Zheng, Xiyang Du, Longfei Liao, Xiaoke Zhao, Zhaowen Zhou, Jingze Song, Bo Zhang, Jiawei Liu, Xiang Qi, Zhe Li, Zhiqiang Zhang, Wei Wang, Peng Zhang

机构 * Ant Group(蚂蚁集团)

专题命中 安全评测 :trustworthy(abstract);分类 cs.CL、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.19803 2025-07-29 cs.LG cs.AI 62%

AI-Based Clinical Rule Discovery for NMIBC Recurrence through Tsetlin Machines

Saram Abbas, Naeem Soomro, Rishad Shafik, Rakesh Heer, Kabita Adhikari

机构 * Newcastle University, UK Freeman Hospital, UK Imperial College London \& Newcastle University Centre for Care School of Engineering Newcastle University, UK

专题命中 安全评测 :trustworthy(abstract);分类 cs.AI、cs.LG

Comments Submitted to ISTM 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.19556 2025-07-29 cs.CY cs.AI 62%

PEMUTA: Pedagogically-Enriched Multi-Granular Undergraduate Thesis Assessment

Jialu Zhang, Qingyang Sun, Qianyi Wang, Weiyi Zhang, Zunjie Xiao, Xiaoqing Zhang, Jianfeng Ren, Jiang Liu

机构 * Research Institute of Trustworthy Autonomous Systems and Department of Computer Science and Engineering(可信自主系统研究 institute 和计算机科学与工程系) Southern University of Science and Technology(南方科技大学) School of Computer Science, University of Nottingham Ningbo China(宁波大学计算机学院) School of Ophthalmology and Optometry, Wenzhou Medical University(温州医学院眼视光学院)

专题命中 安全评测 :alignment(abstract);分类 cs.AI、cs.CY

详情

展开后加载摘要…

URL PDF HTML 收藏
2410.15956 2025-07-29 cs.CL cs.AI 62%

Do Large Language Models Have an English Accent? Evaluating and Improving the Naturalness of Multilingual LLMs

Yanzhu Guo, Simone Conia, Zelin Zhou, Min Li, Saloni Potdar, Henry Xiao

机构 * Apple(苹果公司) Inria Paris(巴黎国家信息与自动化研究所) École Polytechnique(巴黎高等理工学院) Sapienza University of Rome(罗马萨皮恩扎大学)

专题命中 安全评测 :alignment(abstract);分类 cs.CL、cs.AI

Comments ACL 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.20850 2025-07-29 cs.RO cs.AI 57%

Free Energy-Inspired Cognitive Risk Integration for AV Navigation in Pedestrian-Rich Environments

Meiting Dang, Yanping Wu, Yafei Wang, Dezong Zhao, David Flynn, Chongfeng Wei

专题命中 安全评测 :safety(abstract);分类 cs.AI

Comments 14 pages, 5 figures

详情

展开后加载摘要…

URL PDF HTML 收藏