arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

大模型对齐与安全

大模型对齐、安全、越狱、红队、提示注入和可信评测。

2026-01-23 至 2026-01-23 共收录 4 信号源:cs.CL, cs.AI, cs.CY, cs.LG

1. 幻觉与事实性 4 篇

2601.12471 2026-01-23 cs.CL cs.AI 73%

Knowing When to Abstain: Medical LLMs Under Clinical Uncertainty

知何时退避:医疗大语言模型在临床不确定性中的表现

Sravanthi Machcha, Sushrita Yerra, Sahil Gupta, Aishwarya Sahoo, Sharmin Sultana, Hong Yu, Zonghai Yao

机构 * Manning College of Information and Computer Sciences, UMass Amherst, MA, USA(马萨诸塞大学阿姆赫斯特曼宁信息与计算机科学学院) Center for Healthcare Organization and Implementation Research, VA Bedford Health Care(医疗组织与实施研究中心) Miner School of Computer and Information Sciences, UMass Lowell, MA, USA(米纳尔计算机与信息科学学院)

专题命中 幻觉与事实性 :safety(abstract);trustworthy(abstract);分类 cs.CL、cs.AI

AI总结 本文提出MedAbstain基准,探讨医疗LLM在临床不确定性中的退避能力,发现显式退避选项能显著提升安全性,而模型规模和提示方法效果有限。

Comments Equal contribution for the first two authors; To appear in proceedings of the Main Conference of the European Chapter of the Association for Computational Linguistics (EACL) 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.15745 2026-01-23 cs.CL 57%

Hallucination Mitigating for Medical Report Generation

缓解医疗报告生成中的幻觉

Ruoqing Zhao, Runze Xia, Piji Li

机构 * College of Artificial Intelligence, Nanjing University of Aeronautics and Astronautics(南京航空航天大学人工智能学院) MIIT Key Laboratory of Pattern Analysis and Machine Intelligence(信息产业部模式分析与机器智能重点实验室) The Key Laboratory of Brain-Machine Intelligence Technology, Ministry of Education(教育部脑机智能技术重点实验室)

专题命中 幻觉与事实性 :alignment(abstract);分类 cs.CL

AI总结 KERM框架通过知识检索、净化模块和细粒度奖励,有效缓解医疗报告生成中的幻觉问题,提升报告质量。

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.15707 2026-01-23 cs.RO 50%

D-Optimality-Guided Reinforcement Learning for Efficient Open-Loop Calibration of a 3-DOF Ankle Rehabilitation Robot

基于D-最优性的强化学习用于3-自由度踝关节康复机器人高效开环校准

Qifan Hu, Branko Celler, Weidong Mu, Steven W. Su

机构 * Affiliated Provincial Hospital, Shandong First Medical University(山东第一医科大学附属省立医院) Faculty of Engineering, University of New South Wales(新南威尔士大学工程学院) Faculty of Engineering and IT, University of Technology Sydney(悉尼大学技术与信息工程学院)

专题命中 幻觉与事实性 :alignment(abstract)

AI总结 本文提出基于D-最优性的强化学习方法,用于高效校准三自由度踝关节康复机器人,提升校准效率和参数估计的鲁棒性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.25856 2026-01-23 cs.CV 50%

PatchEAD: Unifying Industrial Visual Prompting Frameworks for Patch-Exclusive Anomaly Detection

PatchEAD: 统一工业视觉提示框架用于补丁专属异常检测

Po-Han Huang, Jeng-Lin Li, Po-Hsuan Huang, Ming-Ching Chang, Wei-Chao Chen

机构 * Inventec Corporation(Inventec公司) University at Albany, State University of New York(纽约州立大学阿尔巴尼分校)

专题命中 幻觉与事实性 :alignment(abstract)

AI总结 PatchEAD提出统一的补丁聚焦框架,实现无需训练的工业异常检测,兼容多种基础模型并提升补丁相似性鲁棒性。

Comments 10 pages, 5 figures. WACV 2026 (Accepted)

详情

展开后加载摘要…

URL PDF HTML 收藏