arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

AI 大模型

大模型对齐与安全

大模型对齐、安全、越狱、红队、提示注入和可信评测。

2025-08-26 至 2025-08-26 共收录 77 信号源:cs.CL, cs.AI, cs.CY, cs.LG

1. 偏好对齐 8 篇

2508.16982 2025-08-26 cs.CL 83%

Decoding Alignment: A Critical Survey of LLM Development Initiatives through Value-setting and Data-centric Lens

Ilias Chalkidis

机构 * Department of Computer Science, University of Copenhagen(哥本哈根大学计算机科学系)

专题命中 偏好对齐 :alignment(title,abstract);RLHF(abstract);分类 cs.CL

Comments This is a working paper and will be updated with new information or corrections based on community feedback

详情

展开后加载摘要…

URL PDF HTML 收藏
2405.16455 2025-08-26 stat.ML cs.LG stat.ME 83%

On the Algorithmic Bias of Aligning Large Language Models with RLHF: Preference Collapse and Matching Regularization

Jiancong Xiao, Ziniu Li, Xingyu Xie, Emily Getzen, Cong Fang, Qi Long, Weijie J. Su

机构 * University of Pennsylvania(宾夕法尼亚大学) The Chinese University of Hong Kong, Shenzhen(香港中文大学(深圳)) National University of Singapore(新加坡国立大学) Peking University(北京大学) Joint corresponding authors(联合通讯作者)

专题命中 偏好对齐 :RLHF(title,abstract);alignment(abstract);分类 cs.LG

Comments Accepted for publication in the Journal of the American Statistical Association

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.17000 2025-08-26 cs.CL cs.LG 81%

KL-Regularised Q-Learning: A Token-level Action-Value perspective on Online RLHF

Jason R Brown, Lennie Wells, Edward James Young, Sergio Bacallado

机构 * Computational and Biological Learning Group, Department of Engineering, University of Cambridge, Cambridge, UK(计算生物学学习组,工程系,剑桥大学,剑桥,英国) Department of Computer Science and Technology, University of Cambridge, Cambridge, UK(计算机科学与技术系,剑桥大学,剑桥,英国) Statistics Laboratory, Department of Pure Mathematics and Mathematical Statistics, University of Cambridge, UK(统计实验室,纯粹数学与数学统计系,剑桥大学,英国)

专题命中 偏好对齐 :RLHF(title,abstract);分类 cs.CL、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.17718 2025-08-26 cs.CV cs.AI 79%

Instant Preference Alignment for Text-to-Image Diffusion Models

Yang Li, Songlin Yang, Xiaoxuan Han, Wei Wang, Jing Dong, Yueming Lyu, Ziyu Xue

机构 * New Laboratory of Pattern Recognition, CASIA(模式识别新实验室,中国科学院自动化研究所) The Hong Kong University of Science and Technology(香港科技大学) Nanjing university(南京大学) Academy of Broadcasting Science, NRTA(广播科学研究院,国家广播电视总局)

专题命中 偏好对齐 :alignment(title,abstract);分类 cs.AI

Comments 17 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.16741 2025-08-26 cs.LG cs.AI 73%

WST: Weak-to-Strong Knowledge Transfer via Reinforcement Learning

Haosen Ge, Shuo Li, Lianghuan Huang

机构 * Wharton AI & Analytics Initiative(沃顿人工智能与分析倡议) University of Pennsylvania(宾夕法尼亚大学) Department of Computer and Information Science(计算机与信息科学系) Department of Physics and Astronomy(物理学与天文学系)

专题命中 偏好对齐 :alignment(abstract);RLHF(abstract);分类 cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.17637 2025-08-26 cs.CL cs.AI 62%

Weights-Rotated Preference Optimization for Large Language Models

Chenxu Yang, Ruipeng Jia, Mingyu Zheng, Naibin Gu, Zheng Lin, Siyuan Chen, Weichong Yin, Hua Wu, Weiping Wang

机构 * Institute of Information Engineering, Chinese Academy of Sciences(中国科学院信息工程研究所) School of Cyber Security, University of Chinese Academy of Sciences(中国科学院大学网络安全学院) Baidu Inc.(百度公司)

专题命中 偏好对齐 :DPO(abstract);分类 cs.CL、cs.AI

Comments EMNLP 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.15652 2025-08-26 cs.AI cs.IT cs.LG cs.MA math.IT 62%

Understanding Action Effects through Instrumental Empowerment in Multi-Agent Reinforcement Learning

Ardian Selmonaj, Miroslav Strupl, Oleg Szehr, Alessandro Antonucci

机构 * Istituto Dalle Molle di Studi sull’Intelligenza Artificiale (IDSIA), USI-SUPSI(日内瓦人工智能研究所(IDSIA)、USI-SUPSI)

专题命中 偏好对齐 :alignment(abstract);分类 cs.AI、cs.LG

Comments European Conference on Artificial Intelligence (ECAI) 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.17703 2025-08-26 cs.CL 57%

EMPOWER: Evolutionary Medical Prompt Optimization With Reinforcement Learning

Yinda Chen, Yangfan He, Jing Yang, Dapeng Zhang, Zhenlong Yuan, Muhammad Attique Khan, Jamel Baili, Por Lip Yee

机构 * MoE Key Laboratory of Brain-inspired Intelligent Perception and Cognition, University of Science and Technology of China(脑启发智能感知与认知国家重点实验室,中国科学技术大学) Department of Computer Science, University of Minnesota-Twin Cities(计算机科学系,明尼苏达大学双城分校) Center of Research for Cyber Security and Network (CSNET), Faculty of Computer Science and Information Technology, Universiti Malaya(网络安全与网络研究中心(CSNET),马来亚大学计算机科学与信息技术学院) DSLAB, School of Information Science & Engineering, Lanzhou University(信息科学与工程学院,兰州大学) Institute of Computing Technology, Chinese Academy of Sciences(计算技术研究所,中国科学院) Department of AI, Prince Mohammad bin Fahd University(人工智能系,普林姆·法赫德大学) Department of Computer Engineering, College of Computer Science, King Khalid University(计算机工程系,国王·卡利德大学)

专题命中 偏好对齐 :safety(abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏

2. 安全训练 7 篇

2504.13203 2025-08-26 cs.CR cs.AI cs.CL cs.LG cs.MA 80%

X-Teaming: Multi-Turn Jailbreaks and Defenses with Adaptive Multi-Agents

Salman Rahman, Liwei Jiang, James Shiffer, Genglin Liu, Sheriff Issaka, Md Rizwan Parvez, Hamid Palangi, Kai-Wei Chang, Yejin Choi, Saadia Gabriel

专题命中 安全训练 :alignment(abstract);safety(abstract);jailbreak(abstract);分类 cs.CL、cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.18556 2025-08-26 cs.CL cs.AI 73%

Exploring the Vulnerability of the Content Moderation Guardrail in Large Language Models via Intent Manipulation

Jun Zhuang, Haibo Jin, Ye Zhang, Zhengjian Kang, Wenbin Zhang, Gaby G. Dagher, Haohan Wang

机构 * Boise State University(博伊州立大学) University of Illinois Urbana-Champaign(伊利诺伊大学厄巴纳-香槟分校) University of Pittsburgh(匹兹堡大学) New York University(纽约大学) Florida International University(佛罗里达国际大学)

专题命中 安全训练 :safety(abstract);jailbreak(abstract);分类 cs.CL、cs.AI

Comments Accepted for EMNLP'25 Findings. TL;DR: We propose a new two-stage intent-based prompt-refinement framework, IntentPrompt, that aims to explore the vulnerability of LLMs' content moderation guardrails by refining prompts into benign-looking declarative forms via intent manipulation for red-teaming purposes

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.09600 2025-08-26 cs.MA cs.AI cs.CL cs.CR 62%

Effective Red-Teaming of Policy-Adherent Agents

Itay Nakash, George Kour, Koren Lazar, Matan Vetzler, Guy Uziel, Ateret Anaby-Tavor

专题命中 安全训练 :jailbreak(abstract);分类 cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.18397 2025-08-26 cs.MA cs.AI cs.ET cs.LG 62%

An Outlook on the Opportunities and Challenges of Multi-Agent AI Systems

Fangqiao Tian, An Luo, Jin Du, Xun Xian, Robert Specht, Ganghua Wang, Xuan Bi, Jiawei Zhou, Ashish Kundu, Jayanth Srinivasa, Charles Fleming, Rui Zhang, Zirui Liu, Mingyi Hong, Jie Ding

专题命中 安全训练 :safety(abstract);分类 cs.AI、cs.LG

Comments Corrected references

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.18268 2025-08-26 cs.RO cs.AI 57%

SafeBimanual: Diffusion-based Trajectory Optimization for Safe Bimanual Manipulation

Haoyuan Deng, Wenkai Guo, Qianzhun Wang, Zhenyu Wu, Ziwei Wang

机构 * Nanyang Technological University(南洋理工大学) Beijing University of Posts and Telecommunications(北京邮电大学)

专题命中 安全训练 :safety(abstract);分类 cs.AI

Comments Project website is at: https://denghaoyuan123.github.io/SafeBimanip/

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.17511 2025-08-26 cs.AI 57%

School of Reward Hacks: Hacking harmless tasks generalizes to misaligned behavior in LLMs

Mia Taylor, James Chua, Jan Betley, Johannes Treutlein, Owain Evans

专题命中 安全训练 :alignment(abstract);分类 cs.AI

Comments 42 pages, 26 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.17158 2025-08-26 cs.LG 57%

Towards Safeguarding LLM Fine-tuning APIs against Cipher Attacks

Jack Youstra, Mohammed Mahfoud, Yang Yan, Henry Sleight, Ethan Perez, Mrinank Sharma

专题命中 安全训练 :safety(abstract);分类 cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏

3. 越狱攻击 7 篇

2501.18628 2025-08-26 cs.CR cs.AI cs.CL cs.CY 87%

TombRaider: Entering the Vault of History to Jailbreak Large Language Models

Junchen Ding, Jiahao Zhang, Yi Liu, Ziqi Ding, Gelei Deng, Yuekang Li

机构 * UNSW(新南威尔士大学) NTU(国立大学)

专题命中 越狱攻击 :jailbreak(title,abstract);safety(abstract);red teaming(abstract);分类 cs.CL、cs.AI、cs.CY

Comments Main Conference of EMNLP

详情

展开后加载摘要…

URL PDF HTML 收藏
2408.11308 2025-08-26 cs.AI cs.CL cs.CR 84%

Defending against Jailbreak through Early Exit Generation of Large Language Models

Chongwen Zhao, Zhihao Dou, Kaizhu Huang

机构 * Duke Kunshan University(杜克昆山大学)

专题命中 越狱攻击 :jailbreak(title,abstract);alignment(abstract);分类 cs.CL、cs.AI

Comments ICONIP 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2403.17710 2025-08-26 cs.CR cs.AI 83%

Optimization-based Prompt Injection Attack to LLM-as-a-Judge

Jiawen Shi, Zenghui Yuan, Yinuo Liu, Yue Huang, Pan Zhou, Lichao Sun, Neil Zhenqiang Gong

机构 * Huazhong University of Science and Technology(华中科技大学) University of Notre Dame(圣母大学) Lehigh University(莱文森大学) Duke University(杜克大学)

专题命中 越狱攻击 :prompt injection(title,abstract);jailbreak(abstract);分类 cs.AI

Comments To appear in the Proceedings of The ACM Conference on Computer and Communications Security (CCS), 2024

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.05945 2025-08-26 cs.CL cs.AI 80%

Head-Specific Intervention Can Induce Misaligned AI Coordination in Large Language Models

Paul Darm, Annalisa Riccardi

机构 * University of Strathclyde(斯特拉思克莱德大学)

专题命中 越狱攻击 :alignment(abstract,comments);safety(abstract);jailbreak(abstract);分类 cs.CL、cs.AI

Comments Published at Transaction of Machine Learning Research 08/2025, Large Language Models (LLMs), Interference-time activation shifting, Steerability, Explainability, AI alignment, Interpretability

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.19793 2025-08-26 cs.CR 78%

Prompt Injection Attack to Tool Selection in LLM Agents

Jiawen Shi, Zenghui Yuan, Guiyao Tie, Pan Zhou, Neil Zhenqiang Gong, Lichao Sun

专题命中 越狱攻击 :prompt injection(title,abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.12072 2025-08-26 cs.CR cs.CL 70%

Mitigating Jailbreaks with Intent-Aware LLMs

Wei Jie Yeo, Ranjan Satapathy, Erik Cambria

机构 * Nanyang Technological University(南洋理工大学) Institute of High Performance Computing(高性能计算研究所) Agency for Science, Technology and Research(科技研究局)

专题命中 越狱攻击 :safety(abstract);jailbreak(abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.17244 2025-08-26 cs.AI 57%

L-XAIDS: A LIME-based eXplainable AI framework for Intrusion Detection Systems

Aoun E Muhammad, Kin-Choong Yow, Nebojsa Bacanin-Dzakula, Muhammad Attique Khan

专题命中 越狱攻击 :safety(abstract);分类 cs.AI

Comments This is the authors accepted manuscript of an article accepted for publication in Cluster Computing. The final published version is available at: 10.1007/s10586-025-05326-9

详情

展开后加载摘要…

URL PDF HTML 收藏

4. 红队测试 1 篇

2508.08243 2025-08-26 cs.CL 85%

Jinx: Unlimited LLMs for Probing Alignment Failures

Jiahao Zhao, Liwei Dong

专题命中 红队测试 :alignment(title,abstract);safety(abstract);red teaming(abstract);分类 cs.CL

Comments https://huggingface.co/Jinx-org

详情

展开后加载摘要…

URL PDF HTML 收藏

5. 提示注入 2 篇

2507.07974 2025-08-26 cs.CR 78%

Defending Against Prompt Injection With a Few DefensiveTokens

Sizhe Chen, Yizhu Wang, Nicholas Carlini, Chawin Sitawarin, David Wagner

专题命中 提示注入 :prompt injection(title,abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.17155 2025-08-26 cs.CR cs.AI 77%

Mind the Gap: Time-of-Check to Time-of-Use Vulnerabilities in LLM-Enabled Agents

Derek Lilienthal, Sanghyun Hong

专题命中 提示注入 :safety(abstract);prompt injection(abstract);AI safety(abstract);分类 cs.AI

Comments Pre-print

详情

展开后加载摘要…

URL PDF HTML 收藏

6. 幻觉与事实性 6 篇

2506.01881 2025-08-26 cs.AI cs.CL 81%

WHEN TO ACT, WHEN TO WAIT: Modeling the Intent-Action Alignment Problem in Dialogue

Yaoyao Qian, Jindan Huang, Yuanli Wang, Simon Yu, Kyrie Zhixuan Zhou, Jiayuan Mao, Mingfu Liang, Hanhan Zhou

专题命中 幻觉与事实性 :alignment(title,abstract);分类 cs.CL、cs.AI

Comments Project website: https://nanostorm.netlify.app/

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.17753 2025-08-26 cs.RO cs.AI cs.CL cs.HC 62%

Talking to Robots: A Practical Examination of Speech Foundation Models for HRI Applications

Theresa Pekarek Rosin, Julia Gachot, Henri-Leon Kordt, Matthias Kerzel, Stefan Wermter

机构 * Knowledge Technology, Department of Informatics, University of Hamburg(知识技术,信息学院,汉堡大学)

专题命中 幻觉与事实性 :safety(abstract);分类 cs.CL、cs.AI

Comments Accepted at the workshop on Foundation Models for Social Robotics (FoMoSR) at ICSR 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2401.10266 2025-08-26 cs.LG cs.AI cs.SY eess.SP eess.SY 62%

Intelligent Condition Monitoring of Industrial Plants: An Overview of Methodologies and Uncertainty Management Strategies

Maryam Ahang, Todd Charter, Mostafa Abbasi, Maziyar Khadivi, Oluwaseyi Ogunfowora, Homayoun Najjaran

机构 * Department of Electrical and Computer Engineering, University of Victoria(电气与计算机工程系,维多利亚大学) Department of Mechanical Engineering, University of Victoria(机械工程系,维多利亚大学)

专题命中 幻觉与事实性 :safety(abstract);分类 cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.12964 2025-08-26 cs.CL 57%

Trust Me, I'm Wrong: LLMs Hallucinate with Certainty Despite Knowing the Answer

Adi Simhi, Itay Itzhak, Fazl Barez, Gabriel Stanovsky, Yonatan Belinkov

机构 * Technion – Israel Institute of Technology(技术ion-以色列理工学院) University of Oxford and WhiteBox(牛津大学和WhiteBox) School of Computer Science and Engineering, The Hebrew University of Jerusalem(耶路撒冷希伯来大学计算机科学与工程学院)

专题命中 幻觉与事实性 :safety(abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.18001 2025-08-26 cs.LG stat.ML 57%

A Novel Framework for Uncertainty Quantification via Proper Scores for Classification and Beyond

Sebastian G. Gruber

专题命中 幻觉与事实性 :trustworthy(abstract);分类 cs.LG

Comments PhD Thesis (cumulative, spanning 6 peer-reviewed publications)

详情

展开后加载摘要…

URL PDF HTML 收藏