arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

AI 大模型

大模型对齐与安全

大模型对齐、安全、越狱、红队、提示注入和可信评测。

2025-10-14 至 2025-10-14 共收录 97 信号源:cs.CL, cs.AI, cs.CY, cs.LG

1. 偏好对齐 7 篇

2510.09615 2025-10-14 cs.CR 88%

A Biosecurity Agent for Lifecycle LLM Biosecurity Alignment

Meiyin Meng, Zaixi Zhang

专题命中 偏好对齐 :alignment(title,abstract);DPO(abstract);safety(abstract);jailbreak(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.22578 2025-10-14 cs.LG cs.AI stat.ML 86%

The Hidden Link Between RLHF and Contrastive Learning

Xufei Lv, Kehai Chen, Haoyuan Sun, Xuefeng Bai, Min Zhang, Houde Liu, Kehai Chen

机构 * Tsinghua University(清华大学) Harbin Institute of Technology(哈尔滨工业大学) Soochow University(苏州大学)

专题命中 偏好对齐 :RLHF(title,abstract);alignment(abstract);DPO(abstract);分类 cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.10013 2025-10-14 cs.CL 85%

Path Drift in Large Reasoning Models:How First-Person Commitments Override Safety

Yuyi Huang, Runzhe Zhan, Lidia S. Chao, Ailin Tao, Derek F. Wong

机构 * The Second Affiliated Hospital, Guangdong Provincial Key Laboratory of Allergy and Clinical Immunology, Guangzhou Medical University(广东省过敏与临床免疫学重点实验室,广州医科大学第二附属医院) NLP 2 CT Lab, Department of Computer and Information Science, University of Macau(澳门大学计算机与信息科学系NLP2CT实验室)

专题命中 偏好对齐 :safety(title,abstract);alignment(abstract);RLHF(abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.09895 2025-10-14 cs.CL cs.AI cs.LG 85%

Learning from Reference Answers: Versatile Language Model Alignment without Binary Human Preference Data

Shuai Zhao, Yunqiu Xu, Linchao Zhu, Yi Yang

专题命中 偏好对齐 :alignment(title,abstract);safety(abstract);分类 cs.CL、cs.AI、cs.LG

Comments The code is at https://github.com/mzhaoshuai/RefAlign

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.10963 2025-10-14 cs.LG cs.AI 62%

APLOT: Robust Reward Modeling via Adaptive Preference Learning with Optimal Transport

Zhuo Li, Yuege Feng, Dandan Guo, Jinpeng Hu, Anningzhe Gao, Xiang Wan

机构 * Shenzhen International Center for Industrial and Applied Mathematics(深圳工业与应用数学国际中心) Shenzhen Research Institute of Big Data(深圳大数据研究 institute) The Chinese University of Hong Kong Shenzhen(香港中文大学(深圳)) Birmingham City University(伯明翰城市大学) Jilin University(吉林大学) KAUST(科威特大学) Hefei University of Technology(合肥工业大学)

专题命中 偏好对齐 :RLHF(abstract);分类 cs.AI、cs.LG

Comments EMNLP2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.21776 2025-10-14 cs.CL cs.AI cs.IR 62%

WebThinker: Empowering Large Reasoning Models with Deep Research Capability

Xiaoxi Li, Jiajie Jin, Guanting Dong, Hongjin Qian, Yongkang Wu, Ji-Rong Wen, Yutao Zhu, Zhicheng Dou

机构 * Renmin University of China(中国人民大学) BAAI(北京人工智能研究院) Huawei Poisson Lab(华为Poisson实验室)

专题命中 偏好对齐 :DPO(abstract);分类 cs.CL、cs.AI

Comments Accepted by NeurIPS 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.19223 2025-10-14 cs.LG 57%

LLaDA 1.5: Variance-Reduced Preference Optimization for Large Language Diffusion Models

Fengqi Zhu, Rongzhen Wang, Shen Nie, Xiaolu Zhang, Chunwei Wu, Jun Hu, Jun Zhou, Jianfei Chen, Yankai Lin, Ji-Rong Wen, Chongxuan Li

机构 * Gaoling School of AI, Renmin University of China(中国人民大学人工智能学院) Beijing Key Laboratory of Research on Large Models and Intelligent Governance(北京大型模型与智能治理研究重点实验室) Engineering Research Center of Next-Generation Intelligent Search and Recommendation, MOE(下一代智能搜索与推荐工程技术研究中心) Tsinghua University(清华大学) Ant Group(蚂蚁集团)

专题命中 偏好对齐 :alignment(abstract);分类 cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏

2. 安全训练 8 篇

2506.17368 2025-10-14 cs.LG cs.AI cs.CR 84%

SAFEx: Analyzing Vulnerabilities of MoE-Based LLMs via Stable Safety-critical Expert Identification

Zhenglin Lai, Mengyao Liao, Bingzhe Wu, Dong Xu, Zebin Zhao, Zhihang Yuan, Chao Fan, Jianqiang Li

机构 * School of Artificial Intelligence, Shenzhen University(深圳大学人工智能学院) ByteDance Inc(字节跳动公司)

专题命中 安全训练 :safety(title,abstract);alignment(abstract);分类 cs.AI、cs.LG

Comments 10 pages, 8 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.10452 2025-10-14 cs.CL 83%

Steering Over-refusals Towards Safety in Retrieval Augmented Generation

Utsav Maskey, Mark Dras, Usman Naseem

机构 * Macquarie University(麦觉里大学)

专题命中 安全训练 :safety(title,abstract);alignment(abstract);分类 cs.CL

Comments Preprint

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.10823 2025-10-14 cs.AI cs.NE cs.RO 79%

The Irrational Machine: Neurosis and the Limits of Algorithmic Safety

Daniel Howard

机构 * Howard Science Limited, Malvern, UK(霍华德科学有限公司,英国马尔文) QinetiQ Fellow, UK(QinetiQ Fellow,英国) Member of Senior Common Room, Pembroke College, University of Oxford(奥克斯福德大学彭伯里学院高级共同房间成员)

专题命中 安全训练 :safety(title,abstract);分类 cs.AI

Comments 41 pages, 17 figures, 5 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.10390 2025-10-14 cs.CL cs.AI cs.LG 75%

RefusalBench: Generative Evaluation of Selective Refusal in Grounded Language Models

Aashiq Muhamed, Leonardo F. R. Ribeiro, Markus Dreyer, Virginia Smith, Mona T. Diab

机构 * Carnegie Mellon University(卡内基梅隆大学) Amazon AGI(亚马逊人工智能研究院)

专题命中 安全训练 :alignment(abstract);safety(abstract);分类 cs.CL、cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.09781 2025-10-14 cs.LG cs.AI cs.CL 67%

Building a Foundational Guardrail for General Agentic Systems via Synthetic Data

Yue Huang, Hang Hua, Yujun Zhou, Pengcheng Jing, Manish Nagireddy, Inkit Padhi, Greta Dolcetti, Zhangchen Xu, Subhajit Chaudhury, Ambrish Rawat, Liubov Nedoshivina, Pin-Yu Chen, Prasanna Sattigeri, Xiangliang Zhang

机构 * University of Notre Dame(诺特大学) MIT-IBM Watson AI Lab(MIT-IBM Watson AI实验室) University of Washington(华盛顿大学) Ca’ Foscari University of Venice(威尼斯卡福尔学院) IBM Research(IBM研究院)

专题命中 安全训练 :safety(abstract);分类 cs.CL、cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2110.15907 2025-10-14 cs.AI cs.LG 62%

Learning to Be Cautious

Montaser Mohammedalamen, Dustin Morrill, Alexander Sieusahai, Yash Satsangi, Michael Bowling

机构 * University of Alberta(阿尔伯塔大学) Alberta Machine Intelligence Institute (Amii)(阿尔伯塔人工智能研究所)

专题命中 安全训练 :safety(abstract);分类 cs.AI、cs.LG

Comments Published at the Transactions on Machine Learning Research Journal (TMLR)

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.10633 2025-10-14 cs.AI 57%

Collaborative Text-to-Image Generation via Multi-Agent Reinforcement Learning and Semantic Fusion

Jiabao Shi, Minfeng Qi, Lefeng Zhang, Di Wang, Yingjie Zhao, Ziying Li, Yalong Xing, Ningran Li

机构 * Minzu University of China(民族大学) City University of Macau(澳门城市大学) Key Laboratory of Computing Power Network and Information Security, Ministry of Education, Shandong Computer Science Center (National Supercomputer Center in Jinan), Qilu University of Technology (Shandong Academy of Sciences)(计算能力网络与信息安全重点实验室,教育部,山东计算机科学中心(济南国家超级计算机中心),齐鲁工业大学(山东省科学院)) The University of Adelaide(阿德莱德大学)

专题命中 安全训练 :alignment(abstract);分类 cs.AI

Comments 16 pages, 13 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.11185 2025-10-14 cs.HC 50%

Principles of Safe AI Companions for Youth: Parent and Expert Perspectives

Yaman Yu, Mohi, Aishi Debroy, Xin Cao, Karen Rudolph, Yang Wang

专题命中 安全训练 :safety(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏

3. 越狱攻击 4 篇

2505.13500 2025-10-14 cs.CL cs.AI cs.LG 87%

Noise Injection Systemically Degrades Large Language Model Safety Guardrails

Prithviraj Singh Shahani, Kaveh Eskandari Miandoab, Matthias Scheutz

机构 * Tufts University(塔夫茨大学)

专题命中 越狱攻击 :safety(title,abstract);alignment(abstract);AI safety(abstract);分类 cs.CL、cs.AI、cs.LG

Comments 9 pages,3 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.09699 2025-10-14 cs.CR cs.AI 70%

VisualDAN: Exposing Vulnerabilities in VLMs with Visual-Driven DAN Commands

Aofan Liu, Lulu Tang

机构 * Beijing Academy of Artificial Intelligence(北京人工智能研究院)

专题命中 越狱攻击 :alignment(abstract);jailbreak(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.12575 2025-10-14 cs.CR cs.AI 57%

DemonAgent: Dynamically Encrypted Multi-Backdoor Implantation Attack on LLM-based Agent

Pengyu Zhu, Zhenhong Zhou, Yuanhe Zhang, Shilinlu Yan, Kun Wang, Sen Su

机构 * Beijing University of Posts and Telecommunications(北京邮电大学) Nanyang Technological University(南洋理工大学)

专题命中 越狱攻击 :safety(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.11246 2025-10-14 cs.CR 50%

Collaborative Shadows: Distributed Backdoor Attacks in LLM-Based Multi-Agent Systems

Pengyu Zhu, Lijun Li, Yaxing Lyu, Li Sun, Sen Su, Jing Shao

专题命中 越狱攻击 :safety(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏

4. 提示注入 2 篇

2510.09849 2025-10-14 cs.CL cs.CV 83%

Text Prompt Injection of Vision Language Models

Ruizhe Zhu

机构 * Ruizhe Zhu(独立研究者)

专题命中 提示注入 :prompt injection(title,abstract);safety(abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.11238 2025-10-14 cs.CL cs.AI 62%

Attacks by Content: Automated Fact-checking is an AI Security Issue

Michael Schlichtkrull

机构 * School of Electronic Engineering and Computer Science(电子工程与计算机科学学院) Queen Mary University of London(伦敦女王学院)

专题命中 提示注入 :prompt injection(abstract);分类 cs.CL、cs.AI

Comments Accepted to EMNLP 2025

详情

展开后加载摘要…

URL PDF HTML 收藏

5. 幻觉与事实性 7 篇

2507.09279 2025-10-14 cs.CV cs.AI cs.CL 73%

Prompt4Trust: A Reinforcement Learning Prompt Augmentation Framework for Clinically-Aligned Confidence Calibration in Multimodal Large Language Models

Anita Kriz, Elizabeth Laura Janes, Xing Shen, Tal Arbel

机构 * McGill University(麦吉尔大学) Mila – Quebec AI Institute(魁北克AI研究所)

专题命中 幻觉与事实性 :safety(abstract);trustworthy(abstract);分类 cs.CL、cs.AI

Comments Accepted to ICCV 2025 Workshop CVAMD

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.11457 2025-10-14 cs.AI 57%

From <Answer> to <Think>: Multidimensional Supervision of Reasoning Process for LLM Optimization

Beining Wang, Weihang Su, Hongtao Tian, Tao Yang, Yujia Zhou, Ting Yao, Qingyao Ai, Yiqun Liu

机构 * Tsinghua University(清华大学) Tencent Inc(腾讯公司)

专题命中 幻觉与事实性 :alignment(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.11242 2025-10-14 astro-ph.EP astro-ph.IM cs.LG 57%

Analyzing Data Quality and Decay in Mega-Constellations: A Physics-Informed Machine Learning Approach

Katarina Dyreby, Francisco Caldas, Cláudia Soares

机构 * FCT-UNL, Portugal(葡萄牙FCT-UNL)

专题命中 幻觉与事实性 :safety(abstract);分类 cs.LG

Comments 76th International Astronautical Congress

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.10462 2025-10-14 cs.CV cs.AI 57%

Learning from Disagreement: A Group Decision Simulation Framework for Robust Medical Image Segmentation

Chen Zhong, Yuxuan Yang, Xinyue Zhang, Ruohan Ma, Yong Guo, Gang Li, Jupeng Li

机构 * School of Electronics and Information Engineering, Beijing Jiaotong University, China(电子信息工程学院,北京交通大学) Peking University School and Hospital of Stomatology, China(北京大学口腔医院)

专题命中 幻觉与事实性 :trustworthy(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.17178 2025-10-14 cs.CL 57%

Attention Consistency for LLMs Explanation

Tian Lan, Jinyuan Xu, Xue He, Jenq-Neng Hwang, Lei Li

机构 * Milkuya Studio(Milkuya工作室) ERTIM, INALCO(ERTIM与INALCO) Sorbonne University(索邦大学) IRD(国家农业与食品研究所) University of Washington(华盛顿大学) VitaSight

专题命中 幻觉与事实性 :trustworthy(abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.10290 2025-10-14 cs.SE cs.LG 57%

Grounded AI for Code Review: Resource-Efficient Large-Model Serving in Enterprise Pipelines

Sayan Mandal, Hua Jiang

专题命中 幻觉与事实性 :safety(abstract);分类 cs.LG

Comments Submitted to MLSys 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.09913 2025-10-14 cs.CL 57%

Don't Throw Away Your Pretrained Model

Shangbin Feng, Wenhao Yu, Yike Wang, Hongming Zhang, Yulia Tsvetkov, Dong Yu

机构 * University of Washington(华盛顿大学) Tencent AI Seattle Lab(腾讯AI西雅图实验室)

专题命中 幻觉与事实性 :alignment(abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏

6. 隐私与版权 1 篇

2510.10805 2025-10-14 cs.HC cs.AI 57%

Therapeutic AI and the Hidden Risks of Over-Disclosure: An Embedded AI-Literacy Framework for Mental Health Privacy

Soraya S. Anvari, Rina R. Wehbe

机构 * Dalhousie University(达尔豪西大学)

专题命中 隐私与版权 :safety(abstract);分类 cs.AI

Comments Accepted to SMASH 2025

详情

展开后加载摘要…

URL PDF HTML 收藏

7. 安全评测 36 篇

2510.11235 2025-10-14 cs.AI 89%

AI Alignment Strategies from a Risk Perspective: Independent Safety Mechanisms or Shared Failures?

Leonard Dung, Florian Mai

机构 * Ruhr-Universität Bochum(博尔塔伦大学) Rheinische Friedrich-Wilhelms-Universität Bonn(波恩莱茵-斐迪南-威廉大学) Lamarr Institute for Machine Learning and Artificial Intelligence(拉马尔人工智能与机器学习研究所)

专题命中 安全评测 :alignment(title,abstract);safety(title,abstract);AI safety(abstract);分类 cs.AI

Comments under review

详情

展开后加载摘要…

URL PDF HTML 收藏