arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

AI 大模型

大模型对齐与安全

大模型对齐、安全、越狱、红队、提示注入和可信评测。

2025-09-23 至 2025-09-23 共收录 69 信号源:cs.CL, cs.AI, cs.CY, cs.LG

1. 偏好对齐 9 篇

2410.16531 2025-09-23 cs.CL cs.AI cs.FL cs.LG 80%

Bayesian scaling laws for in-context learning

Aryaman Arora, Dan Jurafsky, Christopher Potts, Noah D. Goodman

机构 * Stanford University(斯坦福大学)

专题命中 偏好对齐 :alignment(abstract);DPO(abstract);safety(abstract);分类 cs.CL、cs.AI、cs.LG

Comments COLM 2025 camera-ready version; 9 pages main text, 39 pages total

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.13146 2025-09-23 cs.CV cs.LG 77%

Re-Align: Aligning Vision Language Models via Retrieval-Augmented Direct Preference Optimization

Shuo Xing, Peiran Li, Yuping Wang, Ruizheng Bai, Yueqi Wang, Chan-Wei Hu, Chengxuan Qian, Huaxiu Yao, Zhengzhong Tu

机构 * Texas A&M University(德克萨斯大学) University of Michigan(密歇根大学) UIUC(伊利诺伊大学香槟分校) UNC Chapel Hill(北卡罗来纳大学教堂山分校)

专题命中 偏好对齐 :alignment(abstract);RLHF(abstract);DPO(abstract);分类 cs.LG

Comments Published at EMNLP 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.13388 2025-09-23 cs.CL cs.AI cs.LG 67%

R3: Robust Rubric-Agnostic Reward Models

David Anugraha, Zilu Tang, Lester James V. Miranda, Hanyang Zhao, Mohammad Rifqi Farhansyah, Garry Kuwanto, Derry Wijaya, Genta Indra Winata

机构 * Stanford University(斯坦福大学) Boston University(波士顿大学) Columbia University(哥伦比亚大学) University of Toronto(多伦多大学) Institut Teknologi Bandung(Bandung 工程技术大学) Monash University Indonesia(墨尔本大学印尼分校) Capital One

专题命中 偏好对齐 :alignment(abstract);分类 cs.CL、cs.AI、cs.LG

Comments Preprint

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.04039 2025-09-23 cs.CV cs.AI cs.CL 62%

Mitigating Hallucinations in Large Vision-Language Models via Entity-Centric Multimodal Preference Optimization

Jiulong Wu, Zhengliang Shi, Shuaiqiang Wang, Jizhou Huang, Dawei Yin, Lingyong Yan, Min Cao, Min Zhang

机构 * Soochow University(苏州大学) Baidu Inc.(百度公司) Shandong University(山东大学)

专题命中 偏好对齐 :alignment(abstract);分类 cs.CL、cs.AI

Comments This paper is accepted by EMNLP2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.16475 2025-09-23 cs.LG cs.CL 62%

Towards Universal Debiasing for Language Models-based Tabular Data Generation

Tianchun Li, Tianci Liu, Xingchen Wang, Rongzhe Wei, Pan Li, Lu Su, Jing Gao

机构 * Purdue University(普渡大学) Georgia Institute of Technology(佐治亚理工学院)

专题命中 偏好对齐 :DPO(abstract);分类 cs.CL、cs.LG

Comments EMNLP 2025 Findings

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.16965 2025-09-23 cs.CL 57%

Preference Distillation via Value based Reinforcement Learning

Minchan Kwon, Junwon Ko, Kangil Kim, Junmo Kim

机构 * Korea Advanced Institute of Science and Technology (KAIST)(韩国先进科学技术研究院) Gwangju Institute of Science and Technology (GIST)(全州科学技术院)

专题命中 偏好对齐 :DPO(abstract);分类 cs.CL

Comments 20 page

Journal ref NIPS 2025 Poster

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.16679 2025-09-23 cs.CL 57%

Reinforcement Learning Meets Large Language Models: A Survey of Advancements and Applications Across the LLM Lifecycle

Keliang Liu, Dingkang Yang, Ziyun Qian, Weijie Yin, Yuchi Wang, Hongsheng Li, Jun Liu, Peng Zhai, Yang Liu, Lihua Zhang

机构 * Fudan University(复旦大学) The Chinese University of Hong Kong, MMLab(香港中文大学MMLab) Lancaster University(兰卡斯特大学) Tongji University(同济大学) The University of Toronto(多伦多大学)

专题命中 偏好对齐 :alignment(abstract);分类 cs.CL

Comments A Survey of Reinforcement Learning for Large Language Models

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.16895 2025-09-23 cs.IR 50%

Temporal-Aware User Behaviour Simulation with Large Language Models for Recommender Systems

Xinye Wanyan, Danula Hettiachchi, Chenglong Ma, Ziqi Xu, Jeffrey Chan

专题命中 偏好对齐 :alignment(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.16560 2025-09-23 cs.CV 50%

Captioning for Text-Video Retrieval via Dual-Group Direct Preference Optimization

Ji Soo Lee, Byungoh Ko, Jaewon Cho, Howoong Lee, Jaewoon Byun, Hyunwoo J. Kim

机构 * Korea University(韩国大学) Hanwha Vision(翰威英航) KAIST(韩国科学技术院)

专题命中 偏好对齐 :DPO(abstract)

Comments EMNLP 2025 Findings

详情

展开后加载摘要…

URL PDF HTML 收藏

2. 安全训练 5 篇

2509.16457 2025-09-23 cs.CL cs.AI cs.CY 85%

Implicit Behavioral Alignment of Language Agents in High-Stakes Crowd Simulations

Yunzhe Wang, Gale M. Lucas, Burcin Becerik-Gerber, Volkan Ustun

机构 * University of Southern California(南加州大学) USC Institute for Creative Technologies(USC创意技术研究所)

专题命中 安全训练 :alignment(title,abstract);trustworthy(abstract);分类 cs.CL、cs.AI、cs.CY

Comments Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing (EMNLP 2025), Main Conference

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.09996 2025-09-23 cs.CL cs.CY 82%

From Judgment to Interference: Early Stopping LLM Harmful Outputs via Streaming Content Monitoring

Yang Li, Qiang Sheng, Yehan Yang, Xueyao Zhang, Juan Cao

机构 * Yang Li Media Synthesis and Forensics Lab, Institute of Computing Technology, Chinese Academy of Sciences University of Chinese Academy of Sciences(中国科学院计算技术研究所、中国科学院自动化研究所、中国科学院大学) Qiang Sheng Media Synthesis and Forensics Lab, Institute of Computing Technology, Chinese Academy of Sciences(中国科学院计算技术研究所、中国科学院自动化研究所) Yehan Yang Media Synthesis and Forensics Lab, Institute of Computing Technology, Chinese Academy of Sciences(中国科学院计算技术研究所、中国科学院自动化研究所) Xueyao Zhang The Chinese University of Hong Kong, Shenzhen(香港中文大学(深圳)) Juan Cao Media Synthesis and Forensics Lab, Institute of Computing Technology, Chinese Academy of Sciences University of Chinese Academy of Sciences(中国科学院计算技术研究所、中国科学院自动化研究所、中国科学院大学)

专题命中 安全训练 :alignment(abstract);DPO(abstract);safety(abstract);harmlessness(abstract)

Comments NeurIPS 2025 Accepted Paper

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.16861 2025-09-23 cs.CR cs.AI cs.SE 79%

AdaptiveGuard: Towards Adaptive Runtime Safety for LLM-Powered Software

Rui Yang, Michael Fu, Chakkrit Tantithamthavorn, Chetan Arora, Gunel Gulmammadova, Joey Chua

机构 * Monash University(墨尔本大学) The University of Melbourne(墨尔本大学) Transurban

专题命中 安全训练 :safety(title);jailbreak(abstract);分类 cs.AI

Comments Accepted to the ASE 2025 International Conference on Automated Software Engineering, Industry Showcase Track

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.16490 2025-09-23 cs.LG 57%

Revisiting Broken Windows Theory

Ziyao Cui, Erick Jiang, Nicholas Sortisio, Haiyan Wang, Eric Chen, Cynthia Rudin

专题命中 安全训练 :safety(abstract);分类 cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.17448 2025-09-23 cs.LG 57%

Rectified Robust Policy Optimization for Model-Uncertain Constrained Reinforcement Learning without Strong Duality

Shaocong Ma, Ziyi Chen, Yi Zhou, Heng Huang

专题命中 安全训练 :safety(abstract);分类 cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏

3. 越狱攻击 8 篇

2502.12970 2025-09-23 cs.CL 83%

Reasoning-to-Defend: Safety-Aware Reasoning Can Defend Large Language Models from Jailbreaking

Junda Zhu, Lingyong Yan, Shuaiqiang Wang, Dawei Yin, Lei Sha

机构 * Beihang University(北京航空航天大学) Baidu Inc.(百度公司) Zhongguancun Laboratory(中关村实验室)

专题命中 越狱攻击 :safety(title,abstract);jailbreak(abstract);分类 cs.CL

Comments EMNLP 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.16870 2025-09-23 cs.SE cs.CR 82%

DecipherGuard: Understanding and Deciphering Jailbreak Prompts for a Safer Deployment of Intelligent Software Systems

Rui Yang, Michael Fu, Chakkrit Tantithamthavorn, Chetan Arora, Gunel Gulmammadova, Joey Chua

专题命中 越狱攻击 :jailbreak(title,abstract);safety(abstract)

Comments Under Review

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.11630 2025-09-23 cs.CR cs.AI cs.CL cs.CY 82%

Jailbreak-Tuning: Models Efficiently Learn Jailbreak Susceptibility

Brendan Murphy, Dillon Bowen, Shahrad Mohammadzadeh, Tom Tseng, Julius Broomfield, Adam Gleave, Kellin Pelrine

机构 * Berkeley(伯克利) Mila – Quebec AI Institute(魁北克人工智能研究所) Montreal, Quebec, Canada(魁北克省蒙特利尔市) McGill University(麦吉尔大学) Georgia Tech(佐治亚理工学院)

专题命中 越狱攻击 :jailbreak(title,abstract);分类 cs.CL、cs.AI、cs.CY

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.21059 2025-09-23 cs.CV cs.AI cs.CR cs.LG 79%

FC-Attack: Jailbreaking Multimodal Large Language Models via Auto-Generated Flowcharts

Ziyi Zhang, Zhen Sun, Zongmin Zhang, Jihui Guo, Xinlei He

机构 * The Hong Kong University of Science and Technology (Guangzhou)(香港科技大学(广州))

专题命中 越狱攻击 :alignment(abstract);safety(abstract);jailbreak(abstract);分类 cs.AI、cs.LG

Comments Accepted to Findings of EMNLP 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.16530 2025-09-23 cs.CL cs.AI 73%

AIPsychoBench: Understanding the Psychometric Differences between LLMs and Humans

Wei Xie, Shuoyoucheng Ma, Zhenhua Wang, Enze Wang, Kai Chen, Xiaobing Sun, Baosheng Wang

机构 * College of Computer Science and Technology, National University of Defense Technology(国防科技大学计算机科学与技术学院) Institute of High Performance Computing, Agency for Science, Technology and Research (A*STAR)(科学研究院高性能计算研究所) Institute of Information Engineering, Chinese Academy of Sciences(中国科学院信息工程研究所)

专题命中 越狱攻击 :alignment(abstract);jailbreak(abstract);分类 cs.CL、cs.AI

Comments Thank you for your attention. This paper was accepted by the CogSci 2025 conference in April and published in August. The location in the proceedings is: https://escholarship.org/uc/item/39k8f46q

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.05652 2025-09-23 cs.CR cs.CL 70%

Sugar-Coated Poison: Benign Generation Unlocks LLM Jailbreaking

Yu-Hang Wu, Yu-Jie Xiong, Hao Zhang, Jia-Chen Zhang, Zheng Zhou

机构 * Shanghai University of Engineering Science(上海工程技术大学) Institute of Computing Technology, Chinese Academy of Sciences(中国科学院计算技术研究所)

专题命中 越狱攻击 :safety(abstract);jailbreak(abstract);分类 cs.CL

Comments Accepted by EMNLP2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.17265 2025-09-23 cs.LG cs.AI 62%

SUA: Stealthy Multimodal Large Language Model Unlearning Attack

Xianren Zhang, Hui Liu, Delvin Ce Zhang, Xianfeng Tang, Qi He, Dongwon Lee, Suhang Wang

机构 * The Pennsylvania State University(宾夕法尼亚州立大学) Amazon(亚马逊) University of Sheffield(谢菲尔德大学)

专题命中 越狱攻击 :alignment(abstract);分类 cs.AI、cs.LG

Comments EMNLP25

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.16792 2025-09-23 cs.CL cs.AI 62%

MIST: Jailbreaking Black-box Large Language Models via Iterative Semantic Tuning

Muyang Zheng, Yuanzhi Yao, Changting Lin, Caihong Kai, Yanxiang Chen, Zhiquan Liu

机构 * School of Computer Science and Information Engineering, Hefei University of Technology(合肥工业大学计算机科学与信息工程学院) College of Computer Science and Technology, Zhejiang University(浙江大学计算机科学与技术学院) College of Cyber Security, Jinan University(暨南大学网络安全学院)

专题命中 越狱攻击 :jailbreak(abstract);分类 cs.CL、cs.AI

Comments 13 pages, 6 figures

详情

展开后加载摘要…

URL PDF HTML 收藏

4. 红队测试 2 篇

2509.17259 2025-09-23 cs.AI 80%

Mind the Gap: Comparing Model- vs Agentic-Level Red Teaming with Action-Graph Observability on GPT-OSS-20B

Ilham Wicaksono, Zekun Wu, Rahul Patel, Theo King, Adriano Koshiyama, Philip Treleaven

机构 * University College London(伦敦大学学院) Holistic AI(整体AI)

专题命中 红队测试 :red teaming(title,abstract);分类 cs.AI

Comments Winner of the OpenAI GPT-OSS-20B Red Teaming Challenge (Kaggle, 2025)

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.17832 2025-09-23 cs.CR 50%

AEAS: Actionable Exploit Assessment System

Xiangmin Shen, Wenyuan Cheng, Yan Chen, Zhenyuan Li, Yuqiao Gu, Lingzhi Wang, Wencheng Zhao, Dawei Sun, Jiashui Wang

专题命中 红队测试 :alignment(abstract)

Comments AEAS has been implemented in the planning agent of PentestAgent, our LLM-driven automated penetration testing framework. Check out our repository: https://github.com/nbshenxm/pentest-agent

详情

展开后加载摘要…

URL PDF HTML 收藏

5. 幻觉与事实性 7 篇

2409.17407 2025-09-23 cs.AI cs.CL 73%

Post-hoc Reward Calibration: A Case Study on Length Bias

Zeyu Huang, Zihan Qiu, Zili Wang, Edoardo M. Ponti, Ivan Titov

机构 * University of Edinburgh(爱丁堡大学) Alibaba Group(阿里巴巴集团) INF Technology(INF技术) University of Amsterdam(阿姆斯特丹大学)

专题命中 幻觉与事实性 :alignment(abstract);RLHF(abstract);分类 cs.CL、cs.AI

Comments ICLR 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.17671 2025-09-23 cs.CL cs.AI 62%

Turk-LettuceDetect: A Hallucination Detection Models for Turkish RAG Applications

Selva Taş, Mahmut El Huseyni, Özay Ezerceli, Reyhan Bayraktar, Fatma Betül Terzioğlu

机构 * Hidden for Review(保密)

专题命中 幻觉与事实性 :trustworthy(abstract);分类 cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.16696 2025-09-23 cs.CL cs.LG 62%

Decoding Uncertainty: The Impact of Decoding Strategies for Uncertainty Estimation in Large Language Models

Wataru Hashimoto, Hidetaka Kamigaito, Taro Watanabe

机构 * Nara Institute of Science and Technology(奈良科学技术大学)

专题命中 幻觉与事实性 :alignment(abstract);分类 cs.CL、cs.LG

Comments Accepted at EMNLP 2025 Findings

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.16742 2025-09-23 cs.AI 57%

Sycophancy Mitigation Through Reinforcement Learning with Uncertainty-Aware Adaptive Reasoning Trajectories

Mohammad Beigi, Ying Shen, Parshin Shojaee, Qifan Wang, Zichao Wang, Chandan Reddy, Ming Jin, Lifu Huang

机构 * University of California, Davis(加州大学戴维斯分校) University of Illinois Urbana-Champaign(伊利诺伊大学厄巴纳-香槟分校) Virginia Tech(弗吉尼亚理工大学) Meta AI Adobe Research(Adobe研究)

专题命中 幻觉与事实性 :alignment(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.16369 2025-09-23 cs.IR cs.AI cs.CE 57%

Enhancing Financial RAG with Agentic AI and Multi-HyDE: A Novel Approach to Knowledge Retrieval and Hallucination Reduction

Akshay Govind Srinivasan, Ryan Jacob George, Jayden Koshy Joe, Hrushikesh Kant, Harshith M R, Sachin Sundar, Sudharshan Suresh, Rahul Vimalkanth, Vijayavallabh

机构 * Indian Institute of Technology Madras(印度理工学院马德拉斯学院)

专题命中 幻觉与事实性 :trustworthy(abstract);分类 cs.AI

Comments 14 Pages, 8 Tables, 2 Figures. Accepted and to be published in the proceedings of FinNLP, Empirical Methods in Natural Language Processing 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.07755 2025-09-23 cs.CL cs.CR 57%

Factuality Beyond Coherence: Evaluating LLM Watermarking Methods for Medical Texts

Rochana Prih Hastuti, Rian Adam Rajagede, Mansour Al Ghanim, Mengxin Zheng, Qian Lou

机构 * University of Central Florida(中央佛罗里达大学)

专题命中 幻觉与事实性 :safety(abstract);分类 cs.CL

Comments Accepted at EMNLP 2025 Findings. Camera Ready

详情

展开后加载摘要…

URL PDF HTML 收藏