arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

AI 大模型

大模型对齐与安全

大模型对齐、安全、越狱、红队、提示注入和可信评测。

2025-08-29 至 2025-08-29 共收录 7 信号源:cs.CL, cs.AI, cs.CY, cs.LG

1. 安全评测 7 篇

2508.20776 2025-08-29 cs.CV cs.AI 70%

Safer Skin Lesion Classification with Global Class Activation Probability Map Evaluation and SafeML

Kuniko Paxton, Koorosh Aslansefat, Amila Akagić, Dhavalkumar Thakker, Yiannis Papadopoulos

机构 * School of Computer Science, University of Hull(赫尔大学计算机科学学院) Faculty of Electrical Engineering, University of Sarajevo(萨拉热窝大学电气工程学院)

专题命中 安全评测 :safety(abstract);trustworthy(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.20737 2025-08-29 cs.SE cs.AI 70%

Rethinking Testing for LLM Applications: Characteristics, Challenges, and a Lightweight Interaction Protocol

Wei Ma, Yixiao Yang, Qiang Hu, Shi Ying, Zhi Jin, Bo Du, Zhenchang Xing, Tianlin Li, Junjie Shi, Yang Liu, Linxiao Jiang

机构 * Singapore Management University Singapore Capital Normal University Beijing China Tianjin University Tianjin China Wuhan University China CSIRO's Data61 \& Australian National University Australia Nanyang Technological University Singapore Singapore Management University Capital Normal University Tianjin University Wuhan University CSIRO's Data61 \& Australian National University Nanyang Technological University

专题命中 安全评测 :safety(abstract);trustworthy(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.18076 2025-08-29 cs.CL 70%

Neither Valid nor Reliable? Investigating the Use of LLMs as Judges

Khaoula Chehbouni, Mohammed Haddou, Jackie Chi Kit Cheung, Golnoosh Farnadi

机构 * McGill University(麦吉尔大学) Mila - Quebec AI Institute(魁北克AI研究所) Statistics Canada(加拿大统计局)

专题命中 安全评测 :alignment(abstract);safety(abstract);分类 cs.CL

Comments Prepared for conference submission

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.21061 2025-08-29 cs.HC cs.AI cs.LG 62%

OnGoal: Tracking and Visualizing Conversational Goals in Multi-Turn Dialogue with Large Language Models

Adam Coscia, Shunan Guo, Eunyee Koh, Alex Endert

机构 * Georgia Institute of Technology(佐治亚理工学院) Adobe Research(Adobe研究)

专题命中 安全评测 :alignment(abstract);分类 cs.AI、cs.LG

Comments Accepted to UIST 2025. 18 pages, 9 figures, 2 tables. For a demo video, see https://youtu.be/uobhmxo6EIE

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.20416 2025-08-29 cs.CL cs.AI 62%

DentalBench: Benchmarking and Advancing LLMs Capability for Bilingual Dentistry Understanding

Hengchuan Zhu, Yihuan Xu, Yichen Li, Zijie Meng, Zuozhu Liu

机构 * Zhejiang University(浙江大学) ZJU-Angelalign R&D Center for Intelligence Healthcare(浙江大学智能医疗研发中心)

专题命中 安全评测 :trustworthy(abstract);分类 cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.20288 2025-08-29 eess.SY cs.LG cs.SY 57%

Neural Spline Operators for Risk Quantification in Stochastic Systems

Zhuoyuan Wang, Raffaele Romagnoli, Kamyar Azizzadenesheli, Yorie Nakahira

机构 * Department of Electrical and Computering Engineering, Carnegie Mellon University(电气与计算机工程系,卡内基梅隆大学) School of Science and Engineering, Department of Mathematics and Computer Science, Duquesne University(科学与工程学院,数学与计算机科学系,杜克森大学)

专题命中 安全评测 :safety(abstract);分类 cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.20851 2025-08-29 cs.CV 50%

PathMR: Multimodal Visual Reasoning for Interpretable Pathology Diagnosis

Ye Zhang, Yu Zhou, Jingwen Qi, Yongbing Zhang, Simon Puettmann, Finn Wichmann, Larissa Pereira Ferreira, Lara Sichward, Julius Keyl, Sylvia Hartmann, Shuo Zhao, Hongxiao Wang, Xiaowei Xu, Jianxu Chen

机构 * School of Computer Science and Technology, Harbin Institute of Technology(哈尔滨工业大学计算机科学与技术学院) Leibniz-Institut für Analytische Wissenschaften – ISAS – e.V.(莱比锡分析科学研究所(ISAS)) Department of Pathology, The Sixth Affiliated Hospital, Sun Yat-sen University(中山大学第六附属医院病理科部) Institute of Pathology, University Hospital Essen(埃森大学医院病理科研究所) Academy for Multidisciplinary Studies, Capital Normal University(首都师范大学多学科研究学院)

专题命中 安全评测 :alignment(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏