arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

大模型对齐与安全

大模型对齐、安全、越狱、红队、提示注入和可信评测。

共收录 9434 信号源:cs.CL, cs.AI, cs.CY, cs.LG

1. 安全评测 9434 篇

2506.09109 2025-11-21 cs.CV cs.CL 57%

CAIRe: Cultural Attribution of Images by Retrieval-Augmented Evaluation

CAIRe:通过检索增强评估进行图像文化归因

Arnav Yayavaram, Siddharth Yayavaram, Simran Khanuja, Michael Saxon, Graham Neubig

机构 * BITS Pilani(比斯·皮兰大学) Carnegie Mellon University(卡内基梅隆大学) University of California, Santa Barbara(加州大学圣巴巴拉分校)

专题命中 安全评测 :alignment(abstract);分类 cs.CL

AI总结 CAIRe通过检索增强评估方法,评估图像在不同文化标签下的相关性,有效衡量文化偏见,提升跨文化公平性。

Comments Preprint, under review

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.15203 2025-11-20 cs.CR cs.AI 57%

Taxonomy, Evaluation and Exploitation of IPI-Centric LLM Agent Defense Frameworks

Zimo Ji, Xunguang Wang, Zongjie Li, Pingchuan Ma, Yudong Gao, Daoyuan Wu, Xincheng Yan, Tian Tian, Shuai Wang

机构 * The Hong Kong University of Science and Technology(香港科技大学) Zhejiang University of Technology(浙江工业大学) Lingnan University(岭南大学) School of Cyber Science and Engineering, Southeast University(东南大学计算机科学与工程学院) ZTE Corporation(中兴通讯有限公司)

专题命中 安全评测 :prompt injection(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.14903 2025-11-20 cs.LG cs.SE 57%

It's LIT! Reliability-Optimized LLMs with Inspectable Tools

Ruixin Zhang, Jon Donnelly, Zhicheng Guo, Ghazal Khalighinejad, Haiyang Huang, Alina Jade Barnett, Cynthia Rudin

机构 * Department of Computer Science(计算机科学系) Duke University(杜克大学)

专题命中 安全评测 :trustworthy(abstract);分类 cs.LG

Comments Accepted to the 39th Conference on Neural Information Processing Systems (NeurIPS 2025) Workshop on Multi-Turn Interactions in Large Language Models

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.14439 2025-11-20 cs.CL 57%

MedBench v4: A Robust and Scalable Benchmark for Evaluating Chinese Medical Language Models, Multimodal Models, and Intelligent Agents

Jinru Ding, Lu Lu, Chao Ding, Mouxiao Bian, Jiayuan Chen, Wenrao Pang, Ruiyao Chen, Xinwei Peng, Renjie Lu, Sijie Ren, Guanxu Zhu, Xiaoqin Wu, Zhiqiang Liu, Rongzhao Zhang, Luyi Jiang, Bing Han, Yunqiu Wang, Jie Xu

专题命中 安全评测 :safety(abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.12527 2025-11-20 cs.LG stat.ML 57%

Selective Risk Certification for LLM Outputs via Information-Lift Statistics: PAC-Bayes, Robustness, and Skeleton Design

Sanjeda Akter, Ibne Farabi Shihab, Anuj Sharma

机构 * Department of Computer Science Iowa State University(计算机科学系爱荷华州立大学) Department of Civil, Construction and Environmental Engineering Iowa State University(土木、建设与环境工程系爱荷华州立大学)

专题命中 安全评测 :safety(abstract);分类 cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.14245 2025-11-19 cs.CL 57%

MuCPT: Music-related Natural Language Model Continued Pretraining

Kai Tian, Yirong Mao, Wendong Bi, Hanjie Wang, Que Wenhui

机构 * Tsinghua University(清华大学) Tencent Inc(腾讯公司)

专题命中 安全评测 :alignment(abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.14120 2025-11-19 cs.CV cs.AI 57%

Multi-view Phase-aware Pedestrian-Vehicle Incident Reasoning Framework with Vision-Language Models

Hao Zhen, Yunxiang Yang, Jidong J. Yang

机构 * College of Engineering University of Georgia(工程学院 乔治亚大学)

专题命中 安全评测 :safety(abstract);分类 cs.AI

Comments 23 pages, 4 figures, 3 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.13970 2025-11-19 cs.AI cs.CV 57%

Scene Graph-Guided Generative AI Framework for Synthesizing and Evaluating Industrial Hazard Scenarios

Sanjay Acharjee, Abir Khan Ratul, Diego Patino, Md Nazmus Sakib

机构 * Ph.D. Student, Dept. of Civil Eng., University of Texas at Arlington. E-mail Assistant Professor, Dept. of Computer Sci. \& Eng., University of Texas at Arlington. E-mail Assistant Professor, Dept. of Civil Eng., University of Texas at Arlington. E-mail

专题命中 安全评测 :safety(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.13771 2025-11-19 cs.CR cs.AI 57%

ExplainableGuard: Interpretable Adversarial Defense for Large Language Models Using Chain-of-Thought Reasoning

Shaowei Guan, Yu Zhai, Zhengyu Zhang, Yanze Wang, Hin Chi Kwok

机构 * Centre for Smart Health, School of Nursing, The Hong Kong Polytechnic University(智能健康研究中心、护理学院、香港理工大学) Department of Language Science and Technology, The Hong Kong Polytechnic University(语言科学与技术系、香港理工大学) Department of Electrical and Electronic Engineering, The Hong Kong Polytechnic University(电子与电气工程系、香港理工大学)

专题命中 安全评测 :trustworthy(abstract);分类 cs.AI

Comments 9 pages, 2 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.10819 2025-11-19 cs.CL 57%

LLM-as-a-Grader: Practical Insights from Large Language Model for Short-Answer and Report Evaluation

Grace Byun, Swati Rajwal, Jinho D. Choi

专题命中 安全评测 :alignment(abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.13626 2025-11-18 cs.AI 57%

CreBench: Human-Aligned Creativity Evaluation from Idea to Process to Product

Kaiwen Xue, Chenglong Li, Zhonghong Ou, Guoxin Zhang, Kaoyan Lu, Shuai Lyu, Yifan Zhu, Ping Zong Junpeng Ding, Xinyu Liu, Qunlin Chen, Weiwei Qin, Yiran Shen, Jiayi Cen

专题命中 安全评测 :alignment(abstract);分类 cs.AI

Comments 13 pages, 3 figures,The 40th Annual AAAI Conference on Artificial Intelligence(AAAI 2026),Paper has been accepted for a poster presentation

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.13169 2025-11-18 cs.CL 57%

TCM-5CEval: Extended Deep Evaluation Benchmark for LLM's Comprehensive Clinical Research Competence in Traditional Chinese Medicine

Tianai Huang, Jiayuan Chen, Lu Lu, Pengcheng Chen, Tianbin Li, Bing Han, Wenchao Tang, Jie Xu, Ming Li

机构 * School of Artificial Intelligence in Traditional Chinese Medicine, Shanghai University of Traditional Chinese Medicine, Shanghai, China(上海中医药大学人工智能学院) Shanghai Artificial Intelligence Laboratory, Shanghai, China(上海人工智能实验室) University of Washington, Seattle, Washington, US(华盛顿大学)

专题命中 安全评测 :alignment(abstract);分类 cs.CL

Comments 17 pages, 8 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.12928 2025-11-18 cs.CL 57%

Visual Room 2.0: Seeing is Not Understanding for MLLMs

Haokun Li, Yazhou Zhang, Jizhi Ding, Qiuchi Li, Peng Zhang

机构 * Tianjin University(天津大学) Shandong Institute of Petroleum and Chemical Technology(山东石油化学技术学院) Beijing Institute of Technology(北京理工大学)

专题命中 安全评测 :alignment(abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.12414 2025-11-18 cs.LG cs.CR 57%

The 'Sure' Trap: Multi-Scale Poisoning Analysis of Stealthy Compliance-Only Backdoors in Fine-Tuned Large Language Models

Yuting Tan, Yi Huang, Zhuo Li

专题命中 安全评测 :alignment(abstract);分类 cs.LG

Comments 13 pages, 5 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.12206 2025-11-18 cs.CV cs.AI 57%

A Novel AI-Driven System for Real-Time Detection of Mirror Absence, Helmet Non-Compliance, and License Plates Using YOLOv8 and OCR

Nishant Vasantkumar Hegde, Aditi Agarwal, Minal Moharir

机构 * Computer Science and Engineering(计算机科学与工程) RV College of Engineering(RV工程学院)

专题命中 安全评测 :safety(abstract);分类 cs.AI

Comments 6 pages, 4 figures. Published in: Proceedings of the 12th International Conference on Emerging Trends in Engineering Technology Signal and Information Processing (ICETET SIP 2025) Note: The conference proceedings contain an outdated abstract due to a publisher-side error. This arXiv version includes the correct and updated abstract

Journal ref 2025 IEEE 12th International Conference on Emerging Trends in Engineering Technology Signal & Information Processing (ICETET SIP 2025)

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.14741 2025-11-18 cs.CV cs.AI 57%

DEXTER: Diffusion-Guided EXplanations with TExtual Reasoning for Vision Models

Simone Carnemolla, Matteo Pennisi, Sarinda Samarasinghe, Giovanni Bellitto, Simone Palazzo, Daniela Giordano, Mubarak Shah, Concetto Spampinato

机构 * University of Catania(卡塔尼亚大学) University of Central Florida(中央佛罗里达大学)

专题命中 安全评测 :trustworthy(abstract);分类 cs.AI

Comments Accepted to NeurIPS 2025 (spotlight)

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.15249 2025-11-18 cs.CL cs.CV 57%

Fooling the LVLM Judges: Visual Biases in LVLM-Based Evaluation

Yerin Hwang, Dongryeol Lee, Kyungmin Min, Taegwan Kang, Yong-il Kim, Kyomin Jung

机构 * IPAI, Seoul National University(IPAI,首尔国立大学) Dept. of ECE, Seoul National University(电子工程系,首尔国立大学) LG AI Research(LG人工智能研究)

专题命中 安全评测 :alignment(abstract);分类 cs.CL

Comments EMNLP 2025 Main (21pgs, 12 Tables, 9 Figures)

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.12052 2025-11-18 cs.CR cs.AI 57%

Exploring AI in Steganography and Steganalysis: Trends, Clusters, and Sustainable Development Potential

Aditya Kumar Sahu, Chandan Kumar, Saksham Kumar, Serdar Solak

专题命中 安全评测 :alignment(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.10027 2025-11-18 cs.AI 57%

ChEmREF: Evaluating Language Model Readiness for Chemical Emergency Response

Risha Surana, Qinyuan Ye, Swabha Swayamdipta

机构 * University of Southern California(南加州大学)

专题命中 安全评测 :safety(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.11252 2025-11-17 cs.AI 57%

UAVBench: An Open Benchmark Dataset for Autonomous and Agentic AI UAV Systems via LLM-Generated Flight Scenarios

Mohamed Amine Ferrag, Abderrahmane Lakas, Merouane Debbah

机构 * Department of Computer and Network Engineering, College of Information Technology, United Arab Emirates University(计算机与网络工程系,信息科技学院,阿联酋大学) Khalifa University of Science and Technology(科技大学)

专题命中 安全评测 :safety(abstract);分类 cs.AI

Comments 18 pages, 5 Figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2410.14720 2025-11-17 cs.LG cs.CV 57%

SGLP: A Similarity Guided Fast Layer Partition Pruning for Compressing Large Deep Models

Yuqi Li, Yao Lu, Junhao Dong, Zeyu Dong, Chuanguang Yang, Xin Yin, Yihao Chen, Jianping Gou, Yingli Tian, Tingwen Huang

机构 * Institute of Computing Technology, Chinese Academy of Sciences(中国科学院计算技术研究所) Zhejiang University of Technology(浙江工业大学) Nanyang Technological University(南洋理工大学) Zhejiang University(浙江大学) Southwest University(西南大学) The City College of New York(纽约城市学院) Shenzhen University of Advanced Technology(深圳大学)

专题命中 安全评测 :alignment(abstract);分类 cs.LG

Comments 16 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.10925 2025-11-17 cs.AI 57%

Multi-Agent Legal Verifier Systems for Data Transfer Planning

Ha-Thanh Nguyen, Wachara Fungwacharakorn, Ken Satoh

机构 * Center of Juris-Informatics, Joint Support-Center for Data Science Research, ROIS, Tokyo, Japan(司法信息中心、数据科学研究联合支持中心、ROIS、东京、日本) Research and Development Center for Large Language Models, NII, ROIS, Tokyo, Japan(大语言模型研究与开发中心、日本信息机构、ROIS、东京、日本)

专题命中 安全评测 :trustworthy(abstract);分类 cs.AI

Comments Presented at NeLaMKRR@KR, 2025 (arXiv:2511.09575)

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.08464 2025-11-17 cs.CV cs.AI 57%

Contrastive Integrated Gradients: A Feature Attribution-Based Method for Explaining Whole Slide Image Classification

Anh Mai Vu, Tuan L. Vo, Ngoc Lam Quang Bui, Nam Nguyen Le Binh, Akash Awasthi, Huy Quoc Vo, Thanh-Huy Nguyen, Zhu Han, Chandra Mohan, Hien Van Nguyen

机构 * ECE Department, University of Houston(德克萨斯大学休斯顿分校电子与计算机工程系) Information Technology, HCMC University of Technology and Education(胡志明市技术与教育大学信息技术系) VN-UK Institute for Research and Executive Education, The University of Da Nang(越南-英国研究与管理教育研究所,丹那大学) Ho Chi Minh City University of Science, Vietnam National University(胡志明市科学大学,越南国家大学) Computational Biology Department, Carnegie Mellon University(卡内基梅隆大学计算生物学系)

专题命中 安全评测 :trustworthy(abstract);分类 cs.AI

Comments Accepted to WACV 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.08954 2025-11-17 cs.CY cs.HC 57%

Should you use LLMs to simulate opinions? Quality checks for early-stage deliberation

Terrence Neumann, Maria De-Arteaga, Sina Fazelpour

专题命中 安全评测 :alignment(abstract);分类 cs.CY

Comments Accepted to AAAI AI for Social Impact (AISI), 2026. This version includes Appendices

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.10203 2025-11-14 cs.CV cs.AI cs.RO 57%

VISTA: A Vision and Intent-Aware Social Attention Framework for Multi-Agent Trajectory Prediction

Stephane Da Silva Martins, Emanuel Aldea, Sylvie Le Hégarat-Mascle

机构 * SATIE - CNRS UMR 8029 Paris-Saclay University, France(巴黎-萨克雷大学)

专题命中 安全评测 :safety(abstract);分类 cs.AI

Comments Paper accepted at WACV 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.07127 2025-11-14 cs.LG 57%

REACT-LLM: A Benchmark for Evaluating LLM Integration with Causal Features in Clinical Prognostic Tasks

Linna Wang, Zhixuan You, Qihui Zhang, Jiunan Wen, Ji Shi, Yimin Chen, Yusen Wang, Fanqi Ding, Ziliang Feng, Li Lu

专题命中 安全评测 :trustworthy(abstract);分类 cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.03553 2025-11-14 cs.CL 57%

CCD-Bench: Probing Cultural Conflict in Large Language Model Decision-Making

Hasibur Rahman, Hanan Salam

专题命中 安全评测 :alignment(abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.05796 2025-11-14 cs.CV cs.AI 57%

Dual-Mode Deep Anomaly Detection for Medical Manufacturing: Structural Similarity and Feature Distance

Julio Zanon Diaz, Georgios Siogkas, Peter Corcoran

专题命中 安全评测 :safety(abstract);分类 cs.AI

Comments 12 pages, 3 figures, 3 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.09742 2025-11-14 cs.CV cs.AI 57%

Feature Quality and Adaptability of Medical Foundation Models: A Comparative Evaluation for Radiographic Classification and Segmentation

Frank Li, Theo Dapamede, Mohammadreza Chavoshi, Young Seok Jeon, Bardia Khosravi, Abdulhameed Dere, Beatrice Brown-Mulry, Rohan Satya Isaac, Aawez Mansuri, Chiratidzo Sanyika, Janice Newsome, Saptarshi Purkayastha, Imon Banerjee, Hari Trivedi, Judy Gichoya

专题命中 安全评测 :alignment(abstract);分类 cs.AI

Comments 7 figures, 3 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.09443 2025-11-13 cs.CV cs.AI 57%

BronchOpt : Vision-Based Pose Optimization with Fine-Tuned Foundation Models for Accurate Bronchoscopy Navigation

Hongchao Shu, Roger D. Soberanis-Mukul, Jiru Xu, Hao Ding, Morgan Ringel, Mali Shen, Saif Iftekar Sayed, Hedyeh Rafii-Tari, Mathias Unberath

机构 * Johnson & Johnson MedTech(强生医疗科技)

专题命中 安全评测 :alignment(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏