arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

AI 大模型

大模型对齐与安全

大模型对齐、安全、越狱、红队、提示注入和可信评测。

2025-10-17 至 2025-10-17 共收录 11 信号源:cs.CL, cs.AI, cs.CY, cs.LG

1. 安全评测 11 篇

2510.11278 2025-10-17 cs.LG cs.AI cs.CL 82%

ENIGMA: The Geometry of Reasoning and Alignment in Large-Language Models

Gareth Seneque, Lap-Hang Ho, Nafise Erfanian Saeedi, Jeffrey Molendijk, Ariel Kuperman, Tim Elson

机构 * Australian Broadcasting Corporation(澳大利亚广播公司)

专题命中 安全评测 :alignment(title,abstract);分类 cs.CL、cs.AI、cs.LG

Comments 52 pages, 10 figures, author typo corrected, abstract typo corrected

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.03550 2025-10-17 cs.CL 79%

Beyond the Surface: Enhancing LLM-as-a-Judge Alignment with Human via Internal Representations

Peng Lai, Jianjie Zheng, Sijie Cheng, Yun Chen, Peng Li, Yang Liu, Guanhua Chen

机构 * Southern University of Science and Technology(南方科技大学) Tsinghua University(清华大学) Shanghai University of Finance and Economics(上海金融学院)

专题命中 安全评测 :alignment(title,abstract);分类 cs.CL

Comments Accepted to NeurIPS 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.19179 2025-10-17 cs.AI 79%

A Design Framework for operationalizing Trustworthy Artificial Intelligence in Healthcare: Requirements, Tradeoffs and Challenges for its Clinical Adoption

Pedro A. Moreno-Sánchez, Javier Del Ser, Mark van Gils, Jussi Hernesniemi

机构 * organization= Faculty of Medicine Health Technology, Tampere University , city= Tampere , postcode= 33100 , country= Finland organization= TECNALIA, Basque Research \& Technology Alliance (BRTA) , city= Derio , postcode= 48160 , country= Spain organization= Department of Mathematics, University of the Basque Country (UPV/EHU) , city= Leioa , postcode= 48940 , country= Spain organization= Tampere Heart Hospital, Tampere University Hospital , city= Tampere , postcode= 33520 , country= Finland

专题命中 安全评测 :trustworthy(title,abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.04426 2025-10-17 cs.CL cs.AI cs.CY 67%

The simulation of judgment in LLMs

Edoardo Loru, Jacopo Nudo, Niccolò Di Marco, Alessandro Santirocchi, Roberto Atzeni, Matteo Cinelli, Vincenzo Cestari, Clelia Rossi-Arnaud, Walter Quattrociocchi

机构 * Department of Computer, Control and Management Engineering, Sapienza University of Rome(计算机、控制与管理工程系,罗马萨皮恩扎大学) Department of Computer Science, Sapienza University of Rome(计算机科学系,罗马萨皮恩扎大学) Department of Legal, Social, and Educational Sciences, Tuscia University(法律、社会与教育科学系,图斯西亚大学) Department of Psychology, Sapienza University of Rome(心理学系,罗马萨皮恩扎大学)

专题命中 安全评测 :alignment(abstract);分类 cs.CL、cs.AI、cs.CY

Comments Please refer to published version: https://doi.org/10.1073/pnas.2518443122

Journal ref Proc. Natl. Acad. Sci. U.S.A. 122 (42) e2518443122, 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.14242 2025-10-17 cs.CL cs.LG 62%

Flip-Flop Consistency: Unsupervised Training for Robustness to Prompt Perturbations in LLMs

Parsa Hejabi, Elnaz Rahmati, Alireza S. Ziabari, Morteza Dehghani

机构 * University of Southern California(南加州大学)

专题命中 安全评测 :alignment(abstract);分类 cs.CL、cs.LG

Comments 14 pages, 6 figures, 3 tables, and 1 algorithm

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.12210 2025-10-17 eess.AS cs.CL cs.LG 62%

DiSTAR: Diffusion over a Scalable Token Autoregressive Representation for Speech Generation

Yakun Song, Xiaobin Zhuang, Jiawei Chen, Zhikang Niu, Guanrou Yang, Chenpeng Du, Dongya Jia, Zhuo Chen, Yuping Wang, Yuxuan Wang, Xie Chen

机构 * X-LANCE Lab, School of Computer Science, Shanghai Jiao Tong University(X-LANCE实验室,计算机科学学院,上海交通大学) ByteDance Inc.(字节跳动公司)

专题命中 安全评测 :alignment(abstract);分类 cs.CL、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.14525 2025-10-17 cs.CV cs.AI 57%

Real-Time Surgical Instrument Defect Detection via Non-Destructive Testing

Qurrat Ul Ain, Atif Aftab Ahmed Jilani, Zunaira Shafqat, Nigar Azhar Butt

专题命中 安全评测 :safety(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.14503 2025-10-17 cs.LG 57%

Learning to Undo: Rollback-Augmented Reinforcement Learning with Reversibility Signals

Andrejs Sorstkins, Omer Tariq, Muhammad Bilal

机构 * School of Computing and Communications, Lancaster University, Lancaster LA1 4WA, United Kingdom(1 计算与通信学院,兰卡斯特大学,英国兰卡斯特 LA1 4WA) Neubility, 2F 115 (04768) Wangsimni-ro, Seongdong-gu, Seoul, South Korea(2 Neubility,韩国首尔松江区 Wangsimni-ro 2F 115 (04768))

专题命中 安全评测 :safety(abstract);分类 cs.LG

Comments Submitted PLOS ONE

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.14194 2025-10-17 cs.AI 57%

Implementation of AI in Precision Medicine

Göktuğ Bender, Samer Faraj, Anand Bhardwaj

机构 * Desautels Faculty of Management(德萨尔斯管理学院) McGill University(麦吉尔大学)

专题命中 安全评测 :trustworthy(abstract);分类 cs.AI

Comments Accepted to SMASH 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.14058 2025-10-17 physics.optics cs.AI eess.IV 57%

Optical Computation-in-Communication enables low-latency, high-fidelity perception in telesurgery

Rui Yang, Jiaming Hu, Jian-Qing Zheng, Yue-Zhen Lu, Jian-Wei Cui, Qun Ren, Yi-Jie Yu, John Edward Wu, Zhao-Yu Wang, Xiao-Li Lin, Dandan Zhang, Mingchu Tang, Christos Masouros, Huiyun Liu, Chin-Pang Liu

机构 * University College London(伦敦大学学院) University of Oxford(牛津大学) CAMS Oxford Institute(牛津大学癌症医学学院) Nuffield Department of Medicine(医学系) Imperial College London(帝国理工学院) Department of Bioengineering(生物工程系)

专题命中 安全评测 :safety(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.13828 2025-10-17 cs.CL 57%

From Explainability to Action: A Generative Operational Framework for Integrating XAI in Clinical Mental Health Screening

Ratna Kandala, Akshata Kishore Moharir, Divya Arvinda Nayak

专题命中 安全评测 :trustworthy(abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏