arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

AI 大模型

大模型对齐与安全

大模型对齐、安全、越狱、红队、提示注入和可信评测。

2025-10-07 至 2025-10-07 共收录 82 信号源:cs.CL, cs.AI, cs.CY, cs.LG

1. 偏好对齐 6 篇

2411.04712 2025-10-07 cs.CV cs.LG 79%

SEE-DPO: Self Entropy Enhanced Direct Preference Optimization

Shivanshu Shekhar, Shreyas Singh, Tong Zhang

机构 * Siebel School of Computing and Data Science(计算与数据科学学院) University of Illinois Urbana-Champaign(伊利诺伊大学厄巴纳-香槟分校) Fractal AI Research(Fractal AI研究院)

专题命中 偏好对齐 :DPO(title,abstract);分类 cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.07510 2025-10-07 cs.CV cs.LG 79%

Divergence Minimization Preference Optimization for Diffusion Model Alignment

Binxu Li, Minkai Xu, Jiaqi Han, Meihua Dang, Stefano Ermon

机构 * Stanford University(斯坦福大学)

专题命中 偏好对齐 :alignment(title,abstract);分类 cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2410.12491 2025-10-07 cs.CL 70%

Insights from the Inverse: Reconstructing LLM Training Goals Through Inverse Reinforcement Learning

Jared Joselowitz, Ritam Majumdar, Arjun Jagota, Matthieu Bou, Nyal Patel, Satyapriya Krishna, Sonali Parbhoo

机构 * Imperial College London(伦敦帝国学院) Harvard University(哈佛大学)

专题命中 偏好对齐 :alignment(abstract);RLHF(abstract);分类 cs.CL

Comments Published as a conference paper at COLM 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.19281 2025-10-07 cs.LG 70%

A Snapshot of Influence: A Local Data Attribution Framework for Online Reinforcement Learning

Yuzheng Hu, Fan Wu, Haotian Ye, David Forsyth, James Zou, Nan Jiang, Jiaqi W. Ma, Han Zhao

机构 * UIUC Urbana(伊利诺伊大学 Urbana分校) Stanford University(斯坦福大学)

专题命中 偏好对齐 :RLHF(abstract);safety(abstract);分类 cs.LG

Comments Accepted at NeurIPS 2025 as an oral

详情

展开后加载摘要…

URL PDF HTML 收藏
2310.16810 2025-10-07 cs.CL cs.AI 62%

Can GPT models Follow Human Summarization Guidelines? A Study for Targeted Communication Goals

Yongxin Zhou, Fabien Ringeval, François Portet

机构 * Univ. Grenoble Alpes, CNRS, Inria, Grenoble INP, LIG(格勒诺布尔阿尔卑斯大学、国家科学研究中心、法国国家信息与自动化研究所、格勒诺布尔INP、实验室)

专题命中 偏好对齐 :alignment(abstract);分类 cs.CL、cs.AI

Comments INLG 2025, Hanoi, Vietnam, October 29 - November 2, 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.25361 2025-10-07 cs.AI 57%

Structural Reward Model: Enhancing Interpretability, Efficiency, and Scalability in Reward Modeling

Xiaoyu Liu, Di Liang, Chang Dai, Hongyu Shan, Peiyang Liu, Yonghao Liu, Muling Wu, Yuntao Li, Xianjie Wu, LI Miao, Jiangrong Shen, Minlong Peng

机构 * Northeastern University, Boston(东北大学,波士顿) Independent Developer(独立开发者) Peiking University(北京大学) Jilin University(吉林大学) Beihang University(北航) Xi’an Jiaotong University(西安交通大学) Baidu Inc(百度公司)

专题命中 偏好对齐 :alignment(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏

2. 安全训练 10 篇

2504.20924 2025-10-07 cs.AI 88%

Domain-Agnostic Scalable AI Safety Ensuring Framework

Beomjun Kim, Kangyeon Kim, Sunwoo Kim, Yeonsang Shin, Heejin Ahn

机构 * Massachusetts Institute of Technology(麻省理工学院) Korea Advanced Institute of Science and Technology(韩国科学技术院) Seoul National University(首尔国立大学)

专题命中 安全训练 :safety(title,abstract);AI safety(title,abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.03283 2025-10-07 cs.LG cs.AI cs.CL cs.DC 82%

MACE: A Hybrid LLM Serving System with Colocated SLO-aware Continuous Retraining Alignment

Yufei Li, Yu Fu, Yue Dong, Cong Liu

专题命中 安全训练 :alignment(title,abstract);分类 cs.CL、cs.AI、cs.LG

Comments 14 pages, 15 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.16856 2025-10-07 cs.CV cs.AI 80%

SIA: Enhancing Safety via Intent Awareness for Vision-Language Models

Youngjin Na, Sangheon Jeong, Youngwan Lee, Jian Lee, Dawoon Jeong, Youngman Kim

机构 * VLM Safety LAB, MODULABS(视觉语言模型安全实验室,MODULABS) ETRI(电子技术研究院) KAIST(韩国科学技术院)

专题命中 安全训练 :safety(title,abstract);分类 cs.AI;trustworthy(comments)

Comments Accepted to Safe and Trustworthy Multimodal AI Systems(SafeMM-AI) Workshop at ICCV2025, Non-archival track

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.04392 2025-10-07 cs.CL cs.AI cs.CY cs.LG 70%

Improving Consistency in Retrieval-Augmented Systems with Group Similarity Rewards

Faisal Hamman, Chenyang Zhu, Anoop Kumar, Xujun Peng, Sanghamitra Dutta, Daben Liu, Alfy Samuel

机构 * University of Maryland, College Park(马里兰大学学院公园分校) Capital One

专题命中 安全训练 :safety(abstract);分类 cs.CL、cs.AI、cs.CY

Comments Accepted at NeurIPS 2025 Workshop on Reliable ML from Unreliable Data

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.17601 2025-10-07 cs.CL 70%

Revisiting Backdoor Attacks on LLMs: A Stealthy and Practical Poisoning Framework via Harmless Inputs

Jiawei Kong, Hao Fang, Xiaochen Yang, Kuofeng Gao, Bin Chen, Shu-Tao Xia, Ke Xu, Han Qiu

机构 * Tsinghua Shenzhen International Graduate School, Tsinghua University(清华大学深圳国际研究生院,清华大学) Department of Software Engineering, Harbin Institute of Technology(哈尔滨工业大学软件工程系) School of Computer Science and Technology, Harbin Institute of Technology, Shenzhen(哈尔滨工业大学深圳校区计算机科学与技术学院) Institute for Network Sciences and Cyberspace, Tsinghua University(清华大学网络科学与空间研究院)

专题命中 安全训练 :alignment(abstract);safety(abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.03892 2025-10-07 cs.AI cs.CL 62%

Kantian-Utilitarian XAI: Meta-Explained

Zahra Atf, Peter R. Lewis

机构 * Faculty of Business and Information Technology(商业与信息技术学院) Ontario Tech University(安大略技术大学)

专题命中 安全训练 :alignment(abstract);分类 cs.CL、cs.AI

Comments Accepted for presentation as a poster at the 35th IEEE International Conference on Collaborative Advances in Software and Computing, 2025. Conference website:https://conf.researchr.org/details/cascon-2025/posters-track/1/Kantian-Utilitarian-XAI-Meta-Explained

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.03527 2025-10-07 cs.CL 57%

Sample, Align, Synthesize: Graph-Based Response Synthesis with ConGrs

Sayan Ghosh, Shahzaib Saqib Warraich, Dhruv Tarsadiya, Gregory Yauney, Swabha Swayamdipta

机构 * University of Southern California(南加州大学)

专题命中 安全训练 :alignment(abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.21292 2025-10-07 cs.SE 50%

Semantic Clustering of Civic Proposals: A Case Study on Brazil's National Participation Platform

Ronivaldo Ferreira, Guilherme da Silva, Carla Rocha, Gustavo Pinto

专题命中 安全训练 :alignment(abstract)

Comments 12 pages, in Portuguese language

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.04076 2025-10-07 cs.RO cs.SY eess.SY 50%

From Shadow to Light: Toward Safe and Efficient Policy Learning Across MPC, DeePC, RL, and LLM Agents

Amin Vahidi-Moghaddam, Sayed Pedram Haeri Boroujeni, Iman Jebellat, Ehsan Jebellat, Niloufar Mehrabi, Zhaojian Li

专题命中 安全训练 :safety(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2409.00882 2025-10-07 cs.SE 50%

SAFE: Advancing Large Language Models in Leveraging Semantic and Syntactic Relationships for Software Vulnerability Detection

Van Nguyen, Surya Nepal, Tingmin Wu, Xingliang Yuan, Carsten Rudolph

专题命中 安全训练 :safety(abstract)

Journal ref Proceedings of the 20th ACM Asia Conference on Computer and Communications Security (ASIA CCS), 2025

详情

展开后加载摘要…

URL PDF HTML 收藏

3. 越狱攻击 4 篇

2510.05025 2025-10-07 cs.CL cs.AI cs.CR 73%

Imperceptible Jailbreaking against Large Language Models

Kuofeng Gao, Yiming Li, Chao Du, Xin Wang, Xingjun Ma, Shu-Tao Xia, Tianyu Pang

机构 * Tsinghua University(清华大学) Sea AI Lab, Singapore(新加坡Sea AI实验室) Nanyang Technological University(南洋理工大学) Fudan University(复旦大学) Peng Cheng Laboratory(鹏城实验室)

专题命中 越狱攻击 :jailbreak(abstract);prompt injection(abstract);分类 cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.18564 2025-10-07 cs.CR cs.AI 70%

DualBreach: Efficient Dual-Jailbreaking via Target-Driven Initialization and Multi-Target Optimization

Xinzhe Huang, Kedong Xiu, Tianhang Zheng, Churui Zeng, Wangze Ni, Zhan Qin, Kui Ren, Chun Chen

机构 * State Key Laboratory of Blockchain and Data Security, Zhejiang University, Hangzhou, China(区块链与数据安全国家重点实验室,浙江大学,杭州,中国) Hangzhou High-Tech Zone (Binjiang) Institute of Blockchain and Data Security, Hangzhou, China(杭州高新技术区(滨江)区块链与数据安全研究院,杭州,中国)

专题命中 越狱攻击 :safety(abstract);jailbreak(abstract);分类 cs.AI

Comments 20 pages, 8 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.21947 2025-10-07 cs.LG cs.AI 62%

Active Attacks: Red-teaming LLMs via Adaptive Environments

Taeyoung Yun, Pierre-Luc St-Charles, Jinkyoo Park, Yoshua Bengio, Minsu Kim

机构 * KAIST(韩国科学技术院) Mila – Québec AI Institute(魁北克人工智能研究所) LawZero Omelet Université de Montréal(蒙特利尔大学)

专题命中 越狱攻击 :safety(abstract);分类 cs.AI、cs.LG

Comments 22 pages, 7 figures, 18 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.23882 2025-10-07 cs.AI cs.CR 57%

Quant Fever, Reasoning Blackholes, Schrodinger's Compliance, and More: Probing GPT-OSS-20B

Shuyi Lin, Tian Lu, Zikai Wang, Bo Wen, Yibo Zhao, Cheng Tan

机构 * Northeastern University(东北大学) Shanghai Jiao Tong University(上海交通大学)

专题命中 越狱攻击 :jailbreak(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏

4. 提示注入 6 篇

2510.04885 2025-10-07 cs.CR cs.LG 83%

RL Is a Hammer and LLMs Are Nails: A Simple Reinforcement Learning Recipe for Strong Prompt Injection

Yuxin Wen, Arman Zharmagambetov, Ivan Evtimov, Narine Kokhlikyan, Tom Goldstein, Kamalika Chaudhuri, Chuan Guo

机构 * University of Maryland(马里兰大学) FAIR at Meta(Meta 元宇宙研究院)

专题命中 提示注入 :prompt injection(title,abstract);safety(abstract);分类 cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.04257 2025-10-07 cs.CR cs.AI 79%

AgentTypo: Adaptive Typographic Prompt Injection Attacks against Black-box Multimodal Agents

Yanjie Li, Yiming Cao, Dong Wang, Bin Xiao

机构 * Computing Department of Hong Kong Polytechnic University(香港理工大学计算机系) Computing Department, The Hong Kong Polytechnic University(香港理工大学计算机系)

专题命中 提示注入 :prompt injection(title,abstract);分类 cs.AI

Comments 13 pages, 8 figures. Submitted to IEEE Transactions on Information Forensics & Security

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.04261 2025-10-07 cs.CR 78%

VortexPIA: Indirect Prompt Injection Attack against LLMs for Efficient Extraction of User Privacy

Yu Cui, Sicheng Pan, Yifei Liu, Haibin Zhang, Cong Zuo

专题命中 提示注入 :prompt injection(title,abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.03705 2025-10-07 cs.CR 78%

Backdoor-Powered Prompt Injection Attacks Nullify Defense Methods

Yulin Chen, Haoran Li, Yuan Sui, Yangqiu Song, Bryan Hooi

专题命中 提示注入 :prompt injection(title,abstract)

Comments EMNLP 2025 Findings

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.13686 2025-10-07 cs.CR 78%

TopicAttack: An Indirect Prompt Injection Attack via Topic Transition

Yulin Chen, Haoran Li, Yuexin Li, Yue Liu, Yangqiu Song, Bryan Hooi

专题命中 提示注入 :prompt injection(title,abstract)

Comments EMNLP 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.16580 2025-10-07 cs.CR 78%

Can Indirect Prompt Injection Attacks Be Detected and Removed?

Yulin Chen, Haoran Li, Yuan Sui, Yufei He, Yue Liu, Yangqiu Song, Bryan Hooi

专题命中 提示注入 :prompt injection(title,abstract)

Comments ACL 2025 Main

详情

展开后加载摘要…

URL PDF HTML 收藏

5. 幻觉与事实性 7 篇

2510.04045 2025-10-07 cs.CL cs.LG 81%

Exploring Chain-of-Thought Reasoning for Steerable Pluralistic Alignment

Yunfan Zhang, Kathleen McKeown, Smaranda Muresan

机构 * Columbia University(哥伦比亚大学) Barnard College(巴纳德学院)

专题命中 幻觉与事实性 :alignment(title);safety(abstract);分类 cs.CL、cs.LG

Comments ACL EMNLP 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.17225 2025-10-07 cs.CL cs.AI 73%

SSFO: Self-Supervised Faithfulness Optimization for Retrieval-Augmented Generation

Xiaqiang Tang, Yi Wang, Keyu Hu, Rui Xu, Chuang Li, Weigao Sun, Jian Li, Sihong Xie

机构 * The Hong Kong University of Science and Technology (Guangzhou)(香港科技大学(广州)) Chinese Academy of Sciences(中国科学院) Shanghai AI Lab(上海人工智能实验室) Tencent Hunyuan(腾讯文英)

专题命中 幻觉与事实性 :alignment(abstract);DPO(abstract);分类 cs.CL、cs.AI

Comments Working in progress

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.04933 2025-10-07 cs.CL cs.AI cs.IT cs.LG cs.NE math.IT 67%

The Geometry of Truth: Layer-wise Semantic Dynamics for Hallucination Detection in Large Language Models

Amir Hameed Mir

机构 * Sirraya Labs(Sirraya实验室)

专题命中 幻觉与事实性 :alignment(abstract);分类 cs.CL、cs.AI、cs.LG

Comments Comments: 14 pages, 14 figures, 5 tables. Code available at: https://github.com/sirraya-tech/Sirraya_LSD_Code

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.04439 2025-10-07 cs.CL 57%

On the Role of Unobserved Sequences on Sample-based Uncertainty Quantification for LLMs

Lucie Kunitomo-Jacquin, Edison Marrese-Taylor, Ken Fukuda

机构 * National Institute of Advanced Industrial Science and Technology (AIST)(国家先进工业科学与技术研究院)

专题命中 幻觉与事实性 :safety(abstract);分类 cs.CL

Comments Accepted to UncertaiNLP workshop of EMNLP 2025

详情

展开后加载摘要…

URL PDF HTML 收藏