arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

大模型对齐与安全

大模型对齐、安全、越狱、红队、提示注入和可信评测。

2026-01-27 至 2026-01-27 共收录 11 信号源:cs.CL, cs.AI, cs.CY, cs.LG

1. 安全训练 11 篇

2601.18730 2026-01-27 cs.CL cs.LG 86%

Reflect: Transparent Principle-Guided Reasoning for Constitutional Alignment at Scale

Reflect: 为大规模宪法对齐的透明原则引导推理

Henry Bell, Caroline Zhang, Mohammed Mobasserul Haque, Dhaval Potdar, Samia Zaman, Brandon Fain

机构 * Duke University(杜克大学) Independent Researcher(独立研究者)

专题命中 安全训练 :alignment(title,abstract);RLHF(abstract);safety(abstract);分类 cs.CL、cs.LG

AI总结 Reflect通过透明推理框架在不需训练数据的情况下提升大语言模型对多样原则的对齐能力,增强安全性和鲁棒性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.17911 2026-01-27 cs.CR 82%

Prompt Injection Evaluations: Refusal Boundary Instability and Artifact-Dependent Compliance in GPT-4-Series Models

提示注入评估:GPT-4系列模型中的拒绝边界不稳定性与依赖于人工制品的合规性

Thomas Heverin

专题命中 安全训练 :prompt injection(title,abstract);safety(abstract)

AI总结 本研究发现GPT-4系列模型的拒绝行为受人工制品类型影响,存在拒绝边界不稳定性,单提示评估高估安全鲁棒性。

Comments 15 pages, 3 figures, 1 table

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.17563 2026-01-27 cs.LG cs.AI 76%

Towards Generalisable Imitation Learning Through Conditioned Transition Estimation and Online Behaviour Alignment

通过条件转移估计和在线行为对齐实现通用模仿学习

Nathan Gavenski, Matteo Leonetti, Odinaldo Rodrigues

机构 * King's College London(伦敦国王学院)

专题命中 安全训练 :alignment(title);分类 cs.AI、cs.LG

AI总结 本文提出UfO方法,通过条件转移估计和在线行为对齐,实现无监督模仿学习,提升模型在未见场景中的泛化能力。

Comments The 25th International Conference on Autonomous Agents and Multi-Agent Systems (AAMAS 2026)

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.17892 2026-01-27 cs.CY cs.AI cs.CL cs.LG 70%

Artificial Intelligence and Intellectual Property Rights: Comparative Transnational Policy Analysis

人工智能与知识产权权利:比较跨国政策分析

Sahibpreet Singh, Manjit Singh

专题命中 安全训练 :alignment(abstract);分类 cs.CL、cs.AI、cs.CY

AI总结 本研究分析了人工智能与知识产权法的结合,揭示印度法律在适应AI生成内容方面的不足,并提出协调法律分类以促进公平创新。

Comments Published in Journal of University Institute of Legal Studies, Vol. 19, Issue 1, pp. 182-208, 2025

Journal ref Journal of University Institute of Legal Studies 19(1), 182-208 (2025)

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.17168 2026-01-27 cs.AI cs.MA 70%

Interpreting Agentic Systems: Beyond Model Explanations to System-Level Accountability

解释代理系统:超越模型解释到系统级问责

Judy Zhu, Dhari Gandhi, Himanshu Joshi, Ahmad Rezaie Mianroodi, Sedef Akinli Kocak, Dhanesh Ramachandran

机构 * Vector Institute for Artificial Intelligence(向量人工智能研究所) University of Texas, Austin(德克萨斯大学奥斯汀分校) Dalhousie University(达尔豪斯大学)

专题命中 安全训练 :safety(abstract);AI safety(abstract);分类 cs.AI

AI总结 本文探讨了代理系统中可解释性技术的必要性,提出需设计专门方法以确保系统在目标形成、环境交互和结果评估等阶段的可追溯性和问责性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.17284 2026-01-27 cs.CL cs.AI 62%

Mind the Ambiguity: Aleatoric Uncertainty Quantification in LLMs for Safe Medical Question Answering

注意模糊性:在LLMs中进行偶然不确定性量化以实现安全的医疗问答

Yaokun Liu, Yifan Liu, Phoebe Mbuvi, Zelin Li, Ruichen Yao, Gawon Lim, Dong Wang

机构 * University of Illinois Urbana-Champaign(伊利诺伊大学厄巴纳-香槟分校)

专题命中 安全训练 :safety(abstract);分类 cs.CL、cs.AI

AI总结 本文提出AU-Probe框架,通过检测输入模糊性提升医疗问答的安全性与准确性。

Comments Accepted at The Web Conference 2026 (WWW 2026)

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.18790 2026-01-27 cs.CL 57%

MortalMATH: Evaluating the Conflict Between Reasoning Objectives and Emergency Contexts

MortalMATH: 评估推理目标与紧急情境之间的冲突

Etienne Lanzeray, Stephane Meilliez, Malo Ruelle, Damien Sileo

机构 * Univ. Lille(里尔大学) Univ. Lille, Inria, CNRS, Centrale Lille, UMR 9189 - CRIStAL(里尔大学)

专题命中 安全训练 :safety(abstract);分类 cs.CL

AI总结 MortalMATH研究了推理模型在紧急情境下的表现差异,发现通用模型能拒绝危险任务,而专门模型则忽视危险,导致安全风险。

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.18619 2026-01-27 cs.AI 57%

Visual Attention Reasoning via Hierarchical Search and Self-Verification

通过分层搜索与自验证的视觉注意力推理

Wei Cai, Jian Zhao, Yuchen Yuan, Tianle Zhang, Ming Zhu, Haichuan Tang, Xuelong Li

专题命中 安全训练 :safety(abstract);分类 cs.AI

AI总结 本文提出通过分层搜索与自验证的视觉注意力推理框架,有效提升多模态大语言模型的视觉定位和推理能力,显著降低幻觉发生率。

Comments The paper is withdrawn by the authors after discovering a flaw in the theoretical derivation presented in the Method section. This incorrect step leads to conclusions that are not supported by the corrected derivation. The authors plan to reconstruct the argument and will release an updated version once the issue is fully resolved

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.16891 2026-01-27 cs.AI 57%

LLMs as Layout Designers: Enhanced Spatial Reasoning for Content-Aware Layout Generation

LLMs作为布局设计师:增强的空间推理用于内容感知布局生成

Sha Li, Stefano Petrangeli, Yu Shen, Xiang Chen, Naren Ramakrishnan

机构 * Virginia Tech(弗吉尼亚理工大学) Adobe Research(Adobe研究)

专题命中 安全训练 :alignment(abstract);分类 cs.AI

AI总结 LaySPA通过强化学习框架增强LLM的空间推理能力,实现内容感知布局生成,优于通用LLM和专用布局模型。

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.03906 2026-01-27 cs.CV 50%

From Filters to VLMs: Benchmarking Defogging Methods through Object Detection and Segmentation Performance

从滤波器到视觉语言模型:通过目标检测和分割性能评估去雾方法

Ardalan Aryashad, Parsa Razmara, Amin Mahjoub, Seyedarmin Azizi, Mahdi Salmani, Arad Firouzkouhi

机构 * University of Southern California(南加州大学)

专题命中 安全训练 :alignment(abstract)

AI总结 本文通过目标检测和分割性能评估,探讨了去雾方法在真实与合成环境中的有效性,揭示了视觉语言模型在恶劣天气下的应用潜力。

Comments Accepted at WACV 2026 Proceedings (Oral), 5th Workshop on Image, Video, and Audio Quality Assessment in Computer Vision, with a focus on VLM and Diffusion Models

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.17107 2026-01-27 cs.CV 50%

StealthMark: Harmless and Stealthy Ownership Verification for Medical Segmentation via Uncertainty-Guided Backdoors

StealthMark: 通过不确定性引导后门实现医疗分割的无害且隐蔽的所有权验证

Qinkai Yu, Chong Zhang, Gaojie Jin, Tianjin Huang, Wei Zhou, Wenhui Li, Xiaobo Jin, Bo Huang, Yitian Zhao, Guang Yang, Gregory Y. H. Lip, Yalin Zheng, Aline Villavicencio, Yanda Meng

机构 * Computer Science Department, University of Exeter(埃克塞特大学计算机科学系) Bioengineering Program, Biological and Environmental Science and Engineering Division (BESE), King Abdullah University of Science and Technology (KAUST)(科廷大学科学与技术学院生物工程项目) School of Advanced Technology, Xi’an Jiaotong-Liverpool University(西安交通大学利物浦大学分校高级技术学院) School of Computer Science and Informatics, Cardiff University(卡迪夫大学计算机科学与信息学院) College of Optoelectronic Engineering, Chongqing University(重庆大学光电工程学院) Ningbo Cixi Institute of Biomedical Engineering, Chinese Academy of Sciences(宁波慈溪生物医学工程研究所,中国科学院) School of Bioengineering, Imperial College London(伦敦帝国理工学院生物工程学院) Liverpool Centre for Cardiovascular Science at University of Liverpool, Liverpool John Moores University and Liverpool Heart & Chest Hospital(利物浦大学心血管科学中心,利物浦约翰摩尔斯大学,利物浦心脏和胸科医院) Eye and Vision Department, University of Liverpool(利物浦大学眼科学与视觉科学系)

专题命中 安全训练 :harmlessness(abstract)

AI总结 StealthMark通过不确定性引导后门实现医疗分割模型的隐蔽无害所有权验证,有效提升模型安全性与实用性。

Comments 15 pages,7 figures. Accepted to IEEE Transactions on Image Processing (TIP) 2026

详情

展开后加载摘要…

URL PDF HTML 收藏