arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

大模型对齐与安全

大模型对齐、安全、越狱、红队、提示注入和可信评测。

共收录 9434 信号源:cs.CL, cs.AI, cs.CY, cs.LG

1. 安全评测 9434 篇

2510.14503 2025-10-17 cs.LG 57%

Learning to Undo: Rollback-Augmented Reinforcement Learning with Reversibility Signals

Andrejs Sorstkins, Omer Tariq, Muhammad Bilal

机构 * School of Computing and Communications, Lancaster University, Lancaster LA1 4WA, United Kingdom(1 计算与通信学院,兰卡斯特大学,英国兰卡斯特 LA1 4WA) Neubility, 2F 115 (04768) Wangsimni-ro, Seongdong-gu, Seoul, South Korea(2 Neubility,韩国首尔松江区 Wangsimni-ro 2F 115 (04768))

专题命中 安全评测 :safety(abstract);分类 cs.LG

Comments Submitted PLOS ONE

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.14194 2025-10-17 cs.AI 57%

Implementation of AI in Precision Medicine

Göktuğ Bender, Samer Faraj, Anand Bhardwaj

机构 * Desautels Faculty of Management(德萨尔斯管理学院) McGill University(麦吉尔大学)

专题命中 安全评测 :trustworthy(abstract);分类 cs.AI

Comments Accepted to SMASH 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.14058 2025-10-17 physics.optics cs.AI eess.IV 57%

Optical Computation-in-Communication enables low-latency, high-fidelity perception in telesurgery

Rui Yang, Jiaming Hu, Jian-Qing Zheng, Yue-Zhen Lu, Jian-Wei Cui, Qun Ren, Yi-Jie Yu, John Edward Wu, Zhao-Yu Wang, Xiao-Li Lin, Dandan Zhang, Mingchu Tang, Christos Masouros, Huiyun Liu, Chin-Pang Liu

机构 * University College London(伦敦大学学院) University of Oxford(牛津大学) CAMS Oxford Institute(牛津大学癌症医学学院) Nuffield Department of Medicine(医学系) Imperial College London(帝国理工学院) Department of Bioengineering(生物工程系)

专题命中 安全评测 :safety(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.13828 2025-10-17 cs.CL 57%

From Explainability to Action: A Generative Operational Framework for Integrating XAI in Clinical Mental Health Screening

Ratna Kandala, Akshata Kishore Moharir, Divya Arvinda Nayak

专题命中 安全评测 :trustworthy(abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.12981 2025-10-16 cs.LG 57%

Reference-Specific Unlearning Metrics Can Hide the Truth: A Reality Check

Sungjun Cho, Dasol Hwang, Frederic Sala, Sangheum Hwang, Kyunghyun Cho, Sungmin Cha

机构 * University of Wisconsin-Madison(威斯康星大学麦迪逊分校) LG AI Research(LG人工智能研究) Seoul National University of Science and Technology(首尔科学技术大学) New York University(纽约大学) Genentech(基因泰克)

专题命中 安全评测 :alignment(abstract);分类 cs.LG

Comments 20 pages, 11 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.12931 2025-10-16 cs.CV cs.CL 57%

Unifying Vision-Language Latents for Zero-label Image Caption Enhancement

Sanghyun Byun, Jung Ick Guack, Mohanad Odema, Baisub Lee, Jacob Song, Woo Seong Chung

机构 * LG Electronics USA(LG电子美国公司)

专题命中 安全评测 :alignment(abstract);分类 cs.CL

Comments Accepted to PMLR and NeurIPS 2025 UniReps

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.09558 2025-10-16 cs.CL 57%

AutoPR: Let's Automate Your Academic Promotion!

Qiguang Chen, Zheng Yan, Mingda Yang, Libo Qin, Yixin Yuan, Hanjing Li, Jinhao Liu, Yiyan Ji, Dengyun Peng, Jiannan Guan, Mengkang Hu, Yantao Du, Wanxiang Che

机构 * LARG Research Center for Social Computing and Interactive Robotics(社会计算与交互机器人研究室) Harbin Institute of Technology(哈尔滨工业大学) School of Computer Science and Engineering(计算机科学与工程学院) Central South University(中南大学) The University of Hong Kong(香港大学) ByteDance China (Seed)(字节跳动中国(种子))

专题命中 安全评测 :alignment(abstract);分类 cs.CL

Comments Preprint. Code: https://github.com/LightChen233/AutoPR . Benchmark: https://huggingface.co/datasets/yzweak/PRBench

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.12713 2025-10-15 cs.AI 57%

Towards Robust Artificial Intelligence: Self-Supervised Learning Approach for Out-of-Distribution Detection

Wissam Salhab, Darine Ameyed, Hamid Mcheick, Fehmi Jaafar

机构 * University of Quebec at Chicoutimi(魁北克大学恰普蒂米分校)

专题命中 安全评测 :safety(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.12704 2025-10-15 cs.CV cs.AI 57%

Hybrid Explanation-Guided Learning for Transformer-Based Chest X-Ray Diagnosis

Shelley Zixin Shu, Haozhe Luo, Alexander Poellinger, Mauricio Reyes

机构 * ARTORG Center for Biomedical Engineering Research, University of Bern(ARTORG生物医学工程研究中心,伯恩大学) Inselspital (Bern University Hospital)(Inselspital(伯恩大学医院)) Insel Gruppe Bern Universitätsinstitut für Diagnostische, Interventionelle und Pädiatrische Radiologie(Bern大学诊断、介入和儿科放射学研究所) Department of Radiation Oncology, Inselspital, Bern University Hospital(放射肿瘤科,Inselspital,伯恩大学医院)

专题命中 安全评测 :alignment(abstract);分类 cs.AI

Comments Accepted by iMIMIC at MICCAI 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.12316 2025-10-15 cs.CL 57%

Beating Harmful Stereotypes Through Facts: RAG-based Counter-speech Generation

Greta Damo, Elena Cabrio, Serena Villata

机构 * Université Côte d’Azur, CNRS, Inria, I3S, France(法国大学-科蒂-阿祖尔大学、国家科学研究中心、法国国家信息与自动化研究所、I3S研究所)

专题命中 安全评测 :trustworthy(abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.12287 2025-10-15 cs.CV cs.CL 57%

Vision Language Models Map Logos to Text via Semantic Entanglement in the Visual Projector

Sifan Li, Hongkai Chen, Yujun Cai, Qingwen Ye, Liyang Chen, Junsong Yuan, Yiwei Wang

机构 * University of California, Merced(加州大学梅尔德分校) vivo Mobile Communication Co., Ltd.(vivo移动通信有限公司) University of Queensland(昆士兰大学) UCLA(加州大学洛杉矶分校) University at Buffalo(布法罗大学)

专题命中 安全评测 :trustworthy(abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.12224 2025-10-15 cs.AI 57%

MedKGEval: A Knowledge Graph-Based Multi-Turn Evaluation Framework for Open-Ended Patient Interactions with Clinical LLMs

Yuechun Yu, Han Ying, Haoan Jin, Wenjian Jiang, Dong Xian, Binghao Wang, Zhou Yang, Mengyue Wu

机构 * Shanghai Jiao Tong University(上海交通大学)

专题命中 安全评测 :safety(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.11745 2025-10-15 cs.LG 57%

Think as a Doctor: An Interpretable AI Approach for ICU Mortality Prediction

Qingwen Li, Xiaohang Zhao, Xiao Han, Hailiang Huang, Lanjuan Liu

机构 * School of Information Management & Engineering, Shanghai University of Finance and Economics(上海金融学院信息管理与工程学院) Key Laboratory of Data Intelligence and Management (Beihang University), Ministry of Industry and Information Technology, School of Economics and Management, Beihang University(北京航空航天大学数据智能与管理重点实验室)

专题命中 安全评测 :trustworthy(abstract);分类 cs.LG

Comments 42 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.11688 2025-10-14 cs.CR cs.AI 57%

PACEbench: A Framework for Evaluating Practical AI Cyber-Exploitation Capabilities

Zicheng Liu, Lige Huang, Jie Zhang, Dongrui Liu, Yuan Tian, Jing Shao

机构 * Shanghai Artificial Intelligence Laboratory(上海人工智能实验室) Shanghai Jiao Tong University(上海交通大学) Institute of Information Engineering, Chinese Academy of Sciences(中国科学院信息工程研究所)

专题命中 安全评测 :trustworthy(abstract);分类 cs.AI

Comments Project webpage available at https://pacebench.github.io/

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.11389 2025-10-14 cs.CL 57%

Beyond Survival: Evaluating LLMs in Social Deduction Games with Human-Aligned Strategies

Zirui Song, Yuan Huang, Junchang Liu, Haozhe Luo, Chenxi Wang, Lang Gao, Zixiang Xu, Mingfei Han, Xiaojun Chang, Xiuying Chen

机构 * Mohamed bin Zayed University of Artificial Intelligence (MBZUAI)(穆罕默德·本·扎耶德人工智能大学)

专题命中 安全评测 :alignment(abstract);分类 cs.CL

Comments 34 pages, 32figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.11313 2025-10-14 cs.AI 57%

Automated Skill Decomposition Meets Expert Ontologies: Bridging the Granularity Gap with LLMs

Le Ngoc Luyen, Marie-Hélène Abel

机构 * Université de technologie de Compiègne, CNRS, Heudiasyc(法国图卢兹技术大学、CNRS、Heudiasyc)

专题命中 安全评测 :alignment(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.11202 2025-10-14 cs.LG cs.CR 57%

Evaluating Line-level Localization Ability of Learning-based Code Vulnerability Detection Models

Marco Pintore, Giorgio Piras, Angelo Sotgiu, Maura Pintor, Battista Biggio

机构 * University of Cagliari, Italy(卡利亚里大学) CINI, Italy(CINI)

专题命中 安全评测 :alignment(abstract);分类 cs.LG

Comments Preprint

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.11151 2025-10-14 cs.CL cs.CR 57%

TypePilot: Leveraging the Scala Type System for Secure LLM-generated Code

Alexander Sternfeld, Andrei Kucharavy, Ljiljana Dolamic

机构 * Institute of Entrepreneurship & Management, HES-SO Le Foyer, Techno-Pôle 1 Sierre, Switzerland(创业与管理学院,HES-SO莱福院,技术园区1,瑞士) Institute of Informatics, HES-SO Techno-Pôle 3 Sierre, Switzerland(信息学院,HES-SO技术园区3,瑞士) Cyber-Defence Campus armasuisse, Science and Technology Thun, Switzerland(网络安全校区,armasuisse,科技,瑞士)

专题命中 安全评测 :safety(abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.11131 2025-10-14 cs.SI cs.CY 57%

SocioBench: Modeling Human Behavior in Sociological Surveys with Large Language Models

Jia Wang, Ziyu Zhao, Tingjuntao Ni, Zhongyu Wei

专题命中 安全评测 :alignment(abstract);分类 cs.CY

Comments Accepted by EMNLP 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.10719 2025-10-14 cs.SD cs.AI 57%

SS-DPPN: A self-supervised dual-path foundation model for the generalizable cardiac audio representation

Ummy Maria Muna, Md Mehedi Hasan Shawon, Md Jobayer, Sumaiya Akter, Md Rakibul Hasan, Md. Golam Rabiul Alam

机构 * Department of Computer Science and Engineering(计算机科学与工程系) BRAC University(布拉克大学) Department of Electricial and Electronic Engineering(电气与电子工程系) Department of Biomedical Engineering(生物医学工程系) Linköping University(林肯堡大学) Department of Electrical and Computer Engineering(电气与计算机工程系) University of Maryland(马里兰大学) School of Electrical Engineering, Computing and Mathematical Sciences(电气工程、计算与数学科学学院)

专题命中 安全评测 :trustworthy(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.09116 2025-10-14 cs.CL 57%

DITING: A Multi-Agent Evaluation Framework for Benchmarking Web Novel Translation

Enze Zhang, Jiaying Wang, Mengxi Xiao, Jifei Liu, Ziyan Kuang, Rui Dong, Eric Dong, Sophia Ananiadou, Min Peng, Qianqian Xie

机构 * School of Artificial Intelligence, Wuhan University(武汉大学人工智能学院) Center for Language and Information Research, Wuhan University(武汉大学语言信息研究中心) Jiangxi Normal University(江西师范大学) The University of Manchester(曼彻斯特大学) Yunnan Trrans Technology Co., Ltd.(云南翻译技术有限公司) Malvern College Chengdu(成都马尔伯学院)

专题命中 安全评测 :safety(abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.04852 2025-10-14 cs.SE cs.AI 57%

FreshBrew: A Benchmark for Evaluating AI Agents on Java Code Migration

Victor May, Diganta Misra, Yanqi Luo, Anjali Sridhar, Justine Gehring, Silvio Soares Ribeiro Junior

机构 * Google(谷歌) Max Planck Institut für Intelligente Systeme (MPI-IS)(马克斯·普朗克智能系统研究所) ELLIS Institute, Tübingen(图宾根ELLIS研究所) Salesforce Gologic Inc(Gologic公司)

专题命中 安全评测 :trustworthy(abstract);分类 cs.AI

Comments 18 pages, 12 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.19073 2025-10-14 cs.CL 57%

MFTCXplain: A Multilingual Benchmark Dataset for Evaluating the Moral Reasoning of LLMs through Multi-hop Hate Speech Explanation

Jackson Trager, Francielle Vargas, Diego Alves, Matteo Guida, Mikel K. Ngueajio, Ameeta Agrawal, Yalda Daryani, Farzan Karimi-Malekabadi, Flor Miriam Plaza-del-Arco

机构 * University of Southern California(南加州大学) São Paulo State University(圣保罗州立大学) Saarland University(萨尔兰大学) University of Melbourne(墨尔本大学) Howard University(霍华德大学) Portland State University(波特兰州立大学) Leiden University(莱顿大学)

专题命中 安全评测 :alignment(abstract);分类 cs.CL

Comments Jackson Trager and Francielle Vargas contributed equally

Journal ref Findings of the Association for Computational Linguistics: EMNLP 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.10224 2025-10-14 cs.CL cs.IR 57%

Text2Token: Unsupervised Text Representation Learning with Token Target Prediction

Ruize An, Richong Zhang, Zhijie Nie, Zhanyu Wu, Yanzhao Zhang, Dingkun Long

机构 * CCSE, School of Computer Science and Engineering, Beihang University, Beijing, China(计算机科学与工程学院,北京航空航天大学,北京,中国) Zhongguancun Laboratory, Beijing, China(中关村实验室,北京,中国) Shen Yuan Honors College, Beihang University, Beijing, China(神元荣誉学院,北京航空航天大学,北京,中国)

专题命中 安全评测 :alignment(abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.10199 2025-10-14 cs.HC cs.AI 57%

Revisiting Trust in the Era of Generative AI: Factorial Structure and Latent Profiles

Haocan Sun, Weizi Liu, Di Wu, Guoming Yu, Mike Yao

专题命中 安全评测 :trustworthy(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.09620 2025-10-14 cs.CR cs.AI 57%

Toward a Unified Security Framework for AI Agents: Trust, Risk, and Liability

Jiayun Mo, Xin Kang, Tieyan Li, Zhongding Lei

专题命中 安全评测 :trustworthy(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.06343 2025-10-14 cs.SE cs.AI cs.CR 57%

Leveraging Large Language Models for Cybersecurity Risk Assessment -- A Case from Forestry Cyber-Physical Systems

Fikret Mert Gultekin, Oscar Lilja, Ranim Khojah, Rebekka Wohlrab, Marvin Damschen, Mazen Mohamad

机构 * Chalmers University of Technology(楚德斯技术大学) University of Gothenburg(哥德堡大学) Carnegie Mellon University(卡内基梅隆大学) RISE Research Institutes of Sweden(瑞典RISE研究机构)

专题命中 安全评测 :safety(abstract);分类 cs.AI

Comments Accepted at Autonomous Agents in Software Engineering (AgenticSE) Workshop, co-located with ASE 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.03418 2025-10-14 cs.AI cs.MA 57%

LegalWiz: A Multi-Agent Generation Framework for Contradiction Detection in Legal Documents

Ananya Mantravadi, Shivali Dalmia, Olga Pospelova, Abhishek Mukherji, Nand Dave, Anudha Mittal

机构 * Centific Amazon(亚马逊)

专题命中 安全评测 :trustworthy(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.14670 2025-10-14 cs.HC cs.AI 57%

StreetLens: Enabling Human-Centered AI Agents for Neighborhood Assessment from Street View Imagery

Jina Kim, Leeje Jang, Yao-Yi Chiang, Guanyu Wang, Michelle C. Pasco

机构 * University of Minnesota(明尼苏达大学)

专题命中 安全评测 :safety(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.01339 2025-10-14 cs.LG 57%

Invariance Makes LLM Unlearning Resilient Even to Unanticipated Downstream Fine-Tuning

Changsheng Wang, Yihua Zhang, Jinghan Jia, Parikshit Ram, Dennis Wei, Yuguang Yao, Soumyadeep Pal, Nathalie Baracaldo, Sijia Liu

机构 * Michigan State University(密歇根州立大学) IBM Research(IBM研究院)

专题命中 安全评测 :safety(abstract);分类 cs.LG

Comments Accepted by ICML 2025

详情

展开后加载摘要…

URL PDF HTML 收藏