arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

大模型对齐与安全

大模型对齐、安全、越狱、红队、提示注入和可信评测。

2025-10-22 至 2025-10-22 共收录 48 信号源:cs.CL, cs.AI, cs.CY, cs.LG

1. 偏好对齐 5 篇

2502.00657 2025-10-22 cs.LG cs.AI cs.CY stat.ML 90%

LLM Safety Alignment is Divergence Estimation in Disguise

Rajdeep Haldar, Ziyi Wang, Qifan Song, Guang Lin, Yue Xing

机构 * Department of Statistics, Purdue University(普渡大学统计系) Department of Statistics, Michigan State University(密歇根州立大学统计系)

专题命中 偏好对齐 :alignment(title,abstract);safety(title,abstract);RLHF(abstract);分类 cs.AI、cs.CY、cs.LG

Comments Accepted to NeurIPS 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.10935 2025-10-22 cs.CL 70%

Introducing Spotlight: A Novel Approach for Generating Captivating Key Information from Documents

Ankan Mullick, Sombit Bose, Rounak Saha, Ayan Kumar Bhowmick, Aditya Vempaty, Prasenjit Dey, Ravi Kokku, Pawan Goyal, Niloy Ganguly

机构 * IIT Kharagpur(印度克达尔普大学) Emergence AI

专题命中 偏好对齐 :alignment(abstract);DPO(abstract);分类 cs.CL

Comments Paper accepted in EMNLP 2025 Main Conference (Full Paper)

Journal ref EMNLP 2025 Main Conference (Full Paper)

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.17999 2025-10-22 cs.CY cs.AI cs.HC cs.LG 67%

The Narcissus Hypothesis: Descending to the Rung of Illusion

Riccardo Cadei, Christian Internò

专题命中 偏好对齐 :alignment(abstract);分类 cs.AI、cs.CY、cs.LG

Comments NeurIPS 2025 Workshop on Evaluating the Evolving LLM Lifecycle: Benchmarks, Emergent Abilities, and Scaling

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.18849 2025-10-22 cs.CL cs.AI 62%

Towards Faithful and Controllable Personalization via Critique-Post-Edit Reinforcement Learning

Chenghao Zhu, Meiling Tao, Tiannan Wang, Dongyi Ding, Yuchen Eleanor Jiang, Wangchunshu Zhou

机构 * The Chinese University of Hong Kong, Shenzhen(香港中文大学(深圳)) University of Electronic Science and Technology of China(电子科技大学) South China Agricultural University(华南农业大学)

专题命中 偏好对齐 :RLHF(abstract);分类 cs.CL、cs.AI

Comments work in progress

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.18433 2025-10-22 cs.CV cs.AI cs.IR 57%

ImageGem: In-the-wild Generative Image Interaction Dataset for Generative Model Personalization

Yuanhe Guo, Linxi Xie, Zhuoran Chen, Kangrui Yu, Ryan Po, Guandao Yang, Gordon Wetztein, Hongyi Wen

机构 * NYU(纽约大学) Stanford(斯坦福大学)

专题命中 偏好对齐 :alignment(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏

2. 安全训练 6 篇

2506.19257 2025-10-22 cs.CV cs.CL 89%

MSR-Align: Policy-Grounded Multimodal Alignment for Safety-Aware Reasoning in Vision-Language Models

Yinan Xia, Yilei Jiang, Yingshui Tan, Xiaoyong Zhu, Xiangyu Yue, Bo Zheng

机构 * Future Lab, Alibaba Group(阿里巴巴集团未来实验室)

专题命中 安全训练 :alignment(title,abstract);safety(title,abstract);jailbreak(abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.18081 2025-10-22 cs.LG cs.AI 88%

Any-Depth Alignment: Unlocking Innate Safety Alignment of LLMs to Any-Depth

Jiawei Zhang, Andrew Estornell, David D. Baek, Bo Li, Xiaojun Xu

机构 * University of Chicago(芝加哥大学) University of Illinois Urbana-Champaign(伊利诺伊大学厄巴纳-香槟分校) Massachusetts Institute of Technology(麻省理工学院)

专题命中 安全训练 :alignment(title,abstract);safety(title,abstract);分类 cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.17918 2025-10-22 cs.CL cs.AI 84%

JT-Safe: Intrinsically Enhancing the Safety and Trustworthiness of LLMs

Junlan Feng, Fanyu Meng, Chong Long, Pengyu Cong, Duqing Wang, Yan Zheng, Yuyao Zhang, Xuanchang Gao, Ye Yuan, Yunfei Ma, Zhijie Ren, Fan Yang, Na Wu, Di Jin, Chao Deng

机构 * China Mobile Jiutian Research(中国移动九天研究所)

专题命中 安全训练 :safety(title,abstract);trustworthy(abstract);分类 cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.18053 2025-10-22 cs.LG cs.AI 73%

Adaptive Divergence Regularized Policy Optimization for Fine-tuning Generative Models

Jiajun Fan, Tong Wei, Chaoran Cheng, Yuxin Chen, Ge Liu

机构 * University of Illinois Urbana-Champaign(伊利诺伊大学厄巴纳-香槟分校)

专题命中 安全训练 :alignment(abstract);DPO(abstract);分类 cs.AI、cs.LG

Comments 30 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.15720 2025-10-22 cs.LG cs.AI 62%

ProSh: Probabilistic Shielding for Model-free Reinforcement Learning

Edwin Hamel-De le Court, Gaspard Ohlmann, Francesco Belardinelli

机构 * Imperial College(帝国学院)

专题命中 安全训练 :safety(abstract);分类 cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.17995 2025-10-22 cs.AI 57%

FABRIC: Framework for Agent-Based Realistic Intelligence Creation

Abhigya Verma, Seganrasan Subramanian, Nandhakumar Kandasamy, Naman Gupta

机构 * ServiceNow

专题命中 安全训练 :alignment(abstract);分类 cs.AI

Comments 51 Pages, 38 Listings, 5 Figures

详情

展开后加载摘要…

URL PDF HTML 收藏

3. 越狱攻击 3 篇

2510.18728 2025-10-22 cs.CR cs.AI 79%

HarmNet: A Framework for Adaptive Multi-Turn Jailbreak Attacks on Large Language Models

Sidhant Narula, Javad Rafiei Asl, Mohammad Ghasemigol, Eduardo Blanco, Daniel Takabi

机构 * University of Arizona(亚利桑那大学)

专题命中 越狱攻击 :jailbreak(title,abstract);分类 cs.AI

Comments This paper has been accepted for presentation at the Conference on Applied Machine Learning in Information Security (CAMLIS 2025)

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.03417 2025-10-22 cs.CR cs.AI 70%

NEXUS: Network Exploration for eXploiting Unsafe Sequences in Multi-Turn LLM Jailbreaks

Javad Rafiei Asl, Sidhant Narula, Mohammad Ghasemigol, Eduardo Blanco, Daniel Takabi

机构 * Old Dominion University(旧 Dominion 大学) University of Arizona(亚利桑那大学)

专题命中 越狱攻击 :alignment(abstract);jailbreak(abstract);分类 cs.AI

Comments This paper has been accepted in the main conference proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing (EMNLP 2025). Javad Rafiei Asl and Sidhant Narula are co-first authors

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.07153 2025-10-22 cs.CR cs.AI 57%

Mind the Web: The Security of Web Use Agents

Avishag Shapira, Parth Atulbhai Gandhi, Edan Habler, Asaf Shabtai

机构 * Ben-Gurion University of the Negev, Israel(内盖夫本·古里安大学)

专题命中 越狱攻击 :safety(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏

4. 红队测试 1 篇

2510.18131 2025-10-22 cs.SE 82%

BlueCodeAgent: A Blue Teaming Agent Enabled by Automated Red Teaming for CodeGen AI

Chengquan Guo, Yuzhou Nie, Chulin Xie, Zinan Lin, Wenbo Guo, Bo Li

专题命中 红队测试 :red teaming(title,abstract);safety(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏

5. 幻觉与事实性 3 篇

2510.18454 2025-10-22 cs.CL 79%

Engagement Undermines Safety: How Stereotypes and Toxicity Shape Humor in Language Models

Atharvan Dogra, Soumya Suvra Ghosal, Ameet Deshpande, Ashwin Kalyan, Dinesh Manocha

机构 * Centre for Responsible AI, IIT Madras(负责任人工智能中心,印度理工学院马德拉斯分校) University of Maryland, College Park(马里兰大学 College Park 分校) Princeton University(普林斯顿大学)

专题命中 幻觉与事实性 :safety(title,abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.14008 2025-10-22 cs.MA 78%

Stop Reducing Responsibility in LLM-Powered Multi-Agent Systems to Local Alignment

Jinwei Hu, Yi Dong, Shuang Ao, Zhuoyun Li, Boxuan Wang, Lokesh Singh, Guangliang Cheng, Sarvapali D. Ramchurn, Xiaowei Huang

专题命中 幻觉与事实性 :alignment(title,abstract)

Comments Updated manuscript of our previous version (arXiv:2502.01714). Under review

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.21657 2025-10-22 cs.CL cs.AI cs.LG 67%

Explaining Large Language Models with gSMILE

Zeinab Dehghani, Mohammed Naveed Akram, Koorosh Aslansefat, Adil Khan, Yiannis Papadopoulos

机构 * University of Hull(赫尔大学) Fraunhofer IESE(弗劳恩霍夫研究所)

专题命中 幻觉与事实性 :trustworthy(abstract);分类 cs.CL、cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏

6. 隐私与版权 2 篇

2510.18674 2025-10-22 cs.CR cs.AI 57%

Exploring Membership Inference Vulnerabilities in Clinical Large Language Models

Alexander Nemecek, Zebin Yun, Zahra Rahmani, Yaniv Harel, Vipin Chaudhary, Mahmood Sharif, Erman Ayday

机构 * Case Western Reserve University(凯斯西储大学) Tel Aviv University(特拉维夫大学)

专题命中 隐私与版权 :alignment(abstract);分类 cs.AI

Comments Accepted at the 1st IEEE Workshop on Healthcare and Medical Device Security, Privacy, Resilience, and Trust (IEEE HMD-SPiRiT)

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.18493 2025-10-22 cs.CR cs.AI cs.HC 57%

One Size Fits All? A Modular Adaptive Sanitization Kit (MASK) for Customizable Privacy-Preserving Phone Scam Detection

Kangzhong Wang, Zitong Shen, Youqian Zhang, Michael MK Cheung, Xiapu Luo, Grace Ngai, Eugene Yujun Fu

机构 * The Hong Kong Polytechnic University(香港理工大学) The Education University of Hong Kong(香港教育大学)

专题命中 隐私与版权 :safety(abstract);分类 cs.AI

Comments 9 pages

详情

展开后加载摘要…

URL PDF HTML 收藏

7. 安全评测 15 篇

2503.04150 2025-10-22 cs.CL cs.AI 81%

Temporal Alignment of LLMs through Cycle Encoding for Long-Range Time Representations

Xue Han, Qian Hu, Yitong Wang, Wenchun Gao, Lianlian Zhang, Qing Wang, Lijun Mei, Chao Deng, Junlan Feng

机构 * JIUTIAN Team China Mobile Research Institute(中国移动研究院)

专题命中 安全评测 :alignment(title,abstract);分类 cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2404.12494 2025-10-22 cs.CL 79%

BIRD: A Trustworthy Bayesian Inference Framework for Large Language Models

Yu Feng, Ben Zhou, Weidong Lin, Dan Roth

机构 * University of Pennsylvania(宾夕法尼亚大学) Arizona State University(亚利桑那州立大学)

专题命中 安全评测 :trustworthy(title,abstract);分类 cs.CL

Journal ref ICLR 2025 (Oral)

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.18550 2025-10-22 cs.NI 78%

JAUNT: Joint Alignment of User Intent and Network State for QoE-centric LLM Tool Routing

Enhan Li, Hongyang Du

专题命中 安全评测 :alignment(title,abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.17575 2025-10-22 cs.HC 67%

DeTAILS: Deep Thematic Analysis with Iterative LLM Support

Ansh Sharma, Karen Cochrane, James R. Wallace

专题命中 安全评测 :alignment(abstract);trustworthy(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.17910 2025-10-22 cs.CY cs.AI cs.CL 67%

Interpretability Framework for LLMs in Undergraduate Calculus

Sagnik Dakshit, Sushmita Sinha Roy

机构 * University of Texas at Tyler(德克萨斯理工大学) Florida Gulf Coast University(佛罗里达盖恩斯维尔大学)

专题命中 安全评测 :alignment(abstract);分类 cs.CL、cs.AI、cs.CY

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.18581 2025-10-22 cs.CY cs.AI 62%

The Cost-Benefit of Interdisciplinarity in AI for Mental Health

Katerina Drakos, Eva Paraschou, Simay Toplu, Line Harder Clemmensen, Christoph Lütge, Nicole Nadine Lønfeldt, Sneha Das

机构 * Center for Social Data Science, Faculty of Social Sciences, University of Copenhagen(哥本哈根大学社会科学学院社会数据科学中心) Dept. of Applied Mathematics and Computer Science, Technical University of Denmark(丹麦技术大学应用数学与计算机科学系) Institute for Ethics in Artificial Intelligence, School of Social Sciences and Technology, Technical University of Munich(慕尼黑技术大学社会科学与技术学院人工智能伦理研究所) Dept. of Mathematical Sciences, University of Copenhagen(哥本哈根大学数学科学系) Child and Adolescent Mental Health Center, Copenhagen University Hospital(哥本哈根大学医院青少年与儿童心理健康中心)

专题命中 安全评测 :alignment(abstract);分类 cs.AI、cs.CY

Comments Accepted for poster presentation at the AI in Science Summit 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.18898 2025-10-22 cs.CV cs.AI cs.LG cs.RO 62%

Interpretable Decision-Making for End-to-End Autonomous Driving

Mona Mirzaie, Bodo Rosenhahn

机构 * Institute for Information Processing, Leibniz University Hannover(信息处理研究所,汉诺威莱布尼茨大学)

专题命中 安全评测 :trustworthy(abstract);分类 cs.AI、cs.LG

Comments Accepted to the ICCV 2025 2nd Workshop on the Challenge Of Out-of-Label Hazards in Autonomous Driving (2COOOL)

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.18425 2025-10-22 cs.AI 57%

Automated urban waterlogging assessment and early warning through a mixture of foundation models

Chenxu Zhang, Fuxiang Huang, Lei Zhang

专题命中 安全评测 :safety(abstract);分类 cs.AI

Comments Submitted to Nature

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.18417 2025-10-22 cs.NI cs.AI 57%

On AI Verification in Open RAN

Rahul Soundrarajan, Claudio Fiandrino, Michele Polese, Salvatore D'Oro, Leonardo Bonati, Tommaso Melodia

机构 * Tejas Networks IMDEA Networks Institute Northeastern University

专题命中 安全评测 :trustworthy(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.06239 2025-10-22 cs.AI 57%

Proof2Silicon: Prompt Repair for Verified Code and Hardware Generation via Reinforcement Learning

Manvi Jha, Jiaxin Wan, Deming Chen

机构 * Electrical and Computer Engineering(电气与计算机工程系) University of Illinois Urbana-Champaign(伊利诺伊大学厄巴纳-香槟分校)

专题命中 安全评测 :safety(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏