arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

大模型对齐与安全

大模型对齐、安全、越狱、红队、提示注入和可信评测。

2025-11-18 至 2025-11-18 共收录 31 信号源:cs.CL, cs.AI, cs.CY, cs.LG

1. 安全评测 31 篇

2502.12659 2025-11-18 cs.CY cs.AI 84%

The Hidden Risks of Large Reasoning Models: A Safety Assessment of R1

Kaiwen Zhou, Chengzhi Liu, Xuandong Zhao, Shreedhar Jangam, Jayanth Srinivasa, Gaowen Liu, Dawn Song, Xin Eric Wang

机构 * UC Santa Cruz(加州大学圣克ruz分校) UC Santa Barbara(加州大学圣芭芭拉分校) UC Berkeley(加州大学伯克利分校) Cisco Research(思科研究)

专题命中 安全评测 :safety(title,abstract);prompt injection(abstract);分类 cs.AI、cs.CY

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.03138 2025-11-18 cs.AI 83%

DeepKnown-Guard: A Proprietary Model-Based Safety Response Framework for AI Agents

Qi Li, Jianjun Xu, Pingtao Wei, Jiu Li, Peiqiang Zhao, Jiwei Shi, Xuan Zhang, Yanhui Yang, Xiaodong Hui, Peng Xu, Wenqin Shao

机构 * Beijing Caizhi Tech(北京彩智科技)

专题命中 安全评测 :safety(title,abstract);trustworthy(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.12155 2025-11-18 cs.LG 83%

Rethinking Deep Alignment Through The Lens Of Incomplete Learning

Thong Bach, Dung Nguyen, Thao Minh Le, Truyen Tran

专题命中 安全评测 :alignment(title,abstract);safety(abstract);分类 cs.LG

Comments AAAI'26

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.12668 2025-11-18 cs.CR cs.AI cs.LG 79%

AI Bill of Materials and Beyond: Systematizing Security Assurance through the AI Risk Scanning (AIRS) Framework

Samuel Nathanson, Alexander Lee, Catherine Chen Kieffer, Jared Junkin, Jessica Ye, Amir Saeed, Melanie Lockhart, Russ Fink, Elisha Peterson, Lanier Watkins

机构 * Johns Hopkins University Applied Physics Laboratory (APL)(约翰霍普金斯大学应用物理实验室)

专题命中 安全评测 :alignment(abstract);safety(abstract);trustworthy(abstract);分类 cs.AI、cs.LG

Comments 13 pages, 4 figures, 6 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.01668 2025-11-18 cs.AI 74%

Hybrid Retrieval-Augmented Generation Agent for Trustworthy Legal Question Answering in Judicial Forensics

Yueqing Xi, Yifan Bai, Huasen Luo, Weiliang Wen, Hui Liu, Haoliang Li

机构 * Department of Electronic Engineering, City University of Hong Kong(DongGuan)(香港城市大学(东莞)电子工程系) Department of Electronic Engineering, City University of Hong Kong(香港城市大学电子工程系)

专题命中 安全评测 :trustworthy(title);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.12014 2025-11-18 cs.CL cs.HC 74%

CURE: Cultural Understanding and Reasoning Evaluation - A Framework for "Thick" Culture Alignment Evaluation in LLMs

Truong Vo, Sanmi Koyejo

专题命中 安全评测 :alignment(title);分类 cs.CL

Comments 7 pages, 5 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.11759 2025-11-18 cs.CR cs.AI cs.CY 73%

Can AI Models be Jailbroken to Phish Elderly Victims? An End-to-End Evaluation

Fred Heiding, Simon Lermen

专题命中 安全评测 :safety(abstract);AI safety(abstract);分类 cs.AI、cs.CY

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.24796 2025-11-18 cs.CY cs.AI 73%

Mutual Wanting in Human--AI Interaction: Empirical Evidence from Large-Scale Analysis of GPT Model Transitions

HaoYang Shang, Xuan Liu

机构 * BreathingCORE

专题命中 安全评测 :alignment(abstract);trustworthy(abstract);分类 cs.AI、cs.CY

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.13670 2025-11-18 cs.HC cs.AI 70%

Person-AI Bidirectional Fit - A Proof-Of-Concept Case Study Of Augmented Human-Ai Symbiosis In Management Decision-Making Process

Agnieszka Bieńkowska, Jacek Małecki, Alexander Mathiesen-Ohman, Katarzyna Tworek

机构 * Department of Management Systems and Organizational Development, Faculty of Management, Wrocław University of Science and Technology(管理系统与组织发展系,管理学院,沃林大学科学与技术学院) Department of Mathematics, Faculty of Mathematics, Wrocław University of Science and Technology(数学系,数学学院,沃林大学科学与技术学院) AMOTHO Research Institute, Vallsjön 20, 780 00 Rörbäcksnäs, Sweden(AMOTHO研究所)

专题命中 安全评测 :alignment(abstract);trustworthy(abstract);分类 cs.AI

Comments 30 pages, 2 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.13381 2025-11-18 cs.CL 70%

Can Large Language Models Function as Qualified Pediatricians? A Systematic Evaluation in Real-World Clinical Contexts

Siyu Zhu, Mouxiao Bian, Yue Xie, Yongyu Tang, Zhikang Yu, Tianbin Li, Pengcheng Chen, Bing Han, Jie Xu, Xiaoyan Dong

机构 * Shanghai Children’s Hospital, School of Medicine,Shanghai Jiao Tong University(上海儿童医学中心,上海交通大学医学院) Shanghai Artificial Intelligence Laboratory(上海人工智能实验室) Shanghai University of Traditional Chinese Medicine(上海中医药大学) University of Washington(华盛顿大学)

专题命中 安全评测 :safety(abstract);trustworthy(abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.15621 2025-11-18 cs.SE 67%

DSCodeBench: A Realistic Benchmark for Data Science Code Generation

Shuyin Ouyang, Dong Huang, Jingwen Guo, Zeyu Sun, Qihao Zhu, Jie M. Zhang

专题命中 安全评测 :alignment(abstract);trustworthy(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2406.08824 2025-11-18 cs.RO cs.AI cs.CL cs.CY 67%

LLM-Driven Robots Risk Enacting Discrimination, Violence, and Unlawful Actions

Andrew Hundt, Rumaisa Azeem, Masoumeh Mansouri, Martim Brandão

机构 * Carnegie Mellon University(卡内基梅隆大学) King’s College London(伦敦国王学院) University of Birmingham(伯明翰大学)

专题命中 安全评测 :safety(abstract);分类 cs.CL、cs.AI、cs.CY

Comments Published in International Journal of Social Robotics (2025). 49 pages (65 with references and appendix), 27 Figures, 8 Tables. Andrew Hundt and Rumaisa Azeem are equal contribution co-first authors. The positions of the two co-first authors were swapped from arxiv version 1 with the written consent of all four authors. The Version of Record is available via DOI: 10.1007/s12369-025-01301-x

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.13712 2025-11-18 cs.LG cs.AI 62%

From Black Box to Insight: Explainable AI for Extreme Event Preparedness

Kiana Vu, İsmet Selçuk Özer, Phung Lai, Zheng Wu, Thilanka Munasinghe, Jennifer Wei

机构 * Department of Cybersecurity University at Albany, SUNY Albany, NY, USA Dept. of Atmospheric \& Environmental Sciences University at Albany, SUNY Albany, NY, USA Lally School of Management Rensselaer Polytechnic Institute Albany, NY, USA Goddard Space Flight Center NASA Greenbelt, MD, USA

专题命中 安全评测 :trustworthy(abstract);分类 cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.04688 2025-11-18 cs.CL cs.LG 62%

Evaluating LLMs' Reasoning Over Ordered Procedural Steps

Adrita Anika, Md Messal Monem Miah

机构 * Amazon(亚马逊公司) Texas A&M University(德克萨斯A&M大学)

专题命中 安全评测 :alignment(abstract);分类 cs.CL、cs.LG

Comments Accepted to IJCNLP-AACL 2025 Findings

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.11691 2025-11-18 cs.LG cs.AI cs.SD 62%

Beyond saliency: enhancing explanation of speech emotion recognition with expert-referenced acoustic cues

Seham Nasr, Zhao Ren, David Johnson

机构 * Center for Cognitive Interaction Technology (CITEC), Bielefeld University, Germany(认知交互技术中心(CITEC),比勒菲尔德大学,德国) Cognitive Systems Lab, University of Bremen, Germany(认知系统实验室,不莱梅大学,德国)

专题命中 安全评测 :trustworthy(abstract);分类 cs.AI、cs.LG

Comments 5 pages, 2 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.11597 2025-11-18 cs.AI cs.CL 62%

CLINB: A Climate Intelligence Benchmark for Foundational Models

Michelle Chen Huebscher, Katharine Mach, Aleksandar Stanić, Markus Leippold, Ben Gaiarin, Zeke Hausfather, Elisa Rawat, Erich Fischer, Massimiliano Ciaramita, Joeri Rogelj, Christian Buck, Lierni Sestorain Saralegui, Reto Knutti

机构 * University of Miami(迈阿密大学) University of Zurich(苏黎世大学) Stripe(Stripe公司) ETH Zurich(苏黎世联邦理工学院) Imperial College London(伦敦帝国理工学院)

专题命中 安全评测 :trustworthy(abstract);分类 cs.CL、cs.AI

Comments Questions, system prompt and model judge prompts available here: https://www.kaggle.com/datasets/deepmind/clinb-questions

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.11583 2025-11-18 cs.LG cs.AI cs.IR 62%

Parallel and Multi-Stage Knowledge Graph Retrieval for Behaviorally Aligned Financial Asset Recommendations

Fernando Spadea, Oshani Seneviratne

机构 * Rensselaer Polytechnic Institute(伦斯勒理工学院)

专题命中 安全评测 :alignment(abstract);分类 cs.AI、cs.LG

Comments 10 pages, 3 figures, RAGE-KG 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.13626 2025-11-18 cs.AI 57%

CreBench: Human-Aligned Creativity Evaluation from Idea to Process to Product

Kaiwen Xue, Chenglong Li, Zhonghong Ou, Guoxin Zhang, Kaoyan Lu, Shuai Lyu, Yifan Zhu, Ping Zong Junpeng Ding, Xinyu Liu, Qunlin Chen, Weiwei Qin, Yiran Shen, Jiayi Cen

专题命中 安全评测 :alignment(abstract);分类 cs.AI

Comments 13 pages, 3 figures,The 40th Annual AAAI Conference on Artificial Intelligence(AAAI 2026),Paper has been accepted for a poster presentation

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.13169 2025-11-18 cs.CL 57%

TCM-5CEval: Extended Deep Evaluation Benchmark for LLM's Comprehensive Clinical Research Competence in Traditional Chinese Medicine

Tianai Huang, Jiayuan Chen, Lu Lu, Pengcheng Chen, Tianbin Li, Bing Han, Wenchao Tang, Jie Xu, Ming Li

机构 * School of Artificial Intelligence in Traditional Chinese Medicine, Shanghai University of Traditional Chinese Medicine, Shanghai, China(上海中医药大学人工智能学院) Shanghai Artificial Intelligence Laboratory, Shanghai, China(上海人工智能实验室) University of Washington, Seattle, Washington, US(华盛顿大学)

专题命中 安全评测 :alignment(abstract);分类 cs.CL

Comments 17 pages, 8 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.12928 2025-11-18 cs.CL 57%

Visual Room 2.0: Seeing is Not Understanding for MLLMs

Haokun Li, Yazhou Zhang, Jizhi Ding, Qiuchi Li, Peng Zhang

机构 * Tianjin University(天津大学) Shandong Institute of Petroleum and Chemical Technology(山东石油化学技术学院) Beijing Institute of Technology(北京理工大学)

专题命中 安全评测 :alignment(abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.12414 2025-11-18 cs.LG cs.CR 57%

The 'Sure' Trap: Multi-Scale Poisoning Analysis of Stealthy Compliance-Only Backdoors in Fine-Tuned Large Language Models

Yuting Tan, Yi Huang, Zhuo Li

专题命中 安全评测 :alignment(abstract);分类 cs.LG

Comments 13 pages, 5 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.12206 2025-11-18 cs.CV cs.AI 57%

A Novel AI-Driven System for Real-Time Detection of Mirror Absence, Helmet Non-Compliance, and License Plates Using YOLOv8 and OCR

Nishant Vasantkumar Hegde, Aditi Agarwal, Minal Moharir

机构 * Computer Science and Engineering(计算机科学与工程) RV College of Engineering(RV工程学院)

专题命中 安全评测 :safety(abstract);分类 cs.AI

Comments 6 pages, 4 figures. Published in: Proceedings of the 12th International Conference on Emerging Trends in Engineering Technology Signal and Information Processing (ICETET SIP 2025) Note: The conference proceedings contain an outdated abstract due to a publisher-side error. This arXiv version includes the correct and updated abstract

Journal ref 2025 IEEE 12th International Conference on Emerging Trends in Engineering Technology Signal & Information Processing (ICETET SIP 2025)

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.14741 2025-11-18 cs.CV cs.AI 57%

DEXTER: Diffusion-Guided EXplanations with TExtual Reasoning for Vision Models

Simone Carnemolla, Matteo Pennisi, Sarinda Samarasinghe, Giovanni Bellitto, Simone Palazzo, Daniela Giordano, Mubarak Shah, Concetto Spampinato

机构 * University of Catania(卡塔尼亚大学) University of Central Florida(中央佛罗里达大学)

专题命中 安全评测 :trustworthy(abstract);分类 cs.AI

Comments Accepted to NeurIPS 2025 (spotlight)

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.15249 2025-11-18 cs.CL cs.CV 57%

Fooling the LVLM Judges: Visual Biases in LVLM-Based Evaluation

Yerin Hwang, Dongryeol Lee, Kyungmin Min, Taegwan Kang, Yong-il Kim, Kyomin Jung

机构 * IPAI, Seoul National University(IPAI,首尔国立大学) Dept. of ECE, Seoul National University(电子工程系,首尔国立大学) LG AI Research(LG人工智能研究)

专题命中 安全评测 :alignment(abstract);分类 cs.CL

Comments EMNLP 2025 Main (21pgs, 12 Tables, 9 Figures)

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.12052 2025-11-18 cs.CR cs.AI 57%

Exploring AI in Steganography and Steganalysis: Trends, Clusters, and Sustainable Development Potential

Aditya Kumar Sahu, Chandan Kumar, Saksham Kumar, Serdar Solak

专题命中 安全评测 :alignment(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.10027 2025-11-18 cs.AI 57%

ChEmREF: Evaluating Language Model Readiness for Chemical Emergency Response

Risha Surana, Qinyuan Ye, Swabha Swayamdipta

机构 * University of Southern California(南加州大学)

专题命中 安全评测 :safety(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.13369 2025-11-18 cs.SI physics.soc-ph 50%

Unifying points of interest taxonomies: mapping OpenStreetMap tags to the Foursquare category system

Lilou Soulas, Lorenzo Lucchini, Maurizio Napolitano, Sebastiano Bontorin, Simone Centellegher, Bruno Lepri, Riccardo Gallotti, Eleonora Andreotti

专题命中 安全评测 :alignment(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.13269 2025-11-18 cs.CV 50%

Is your VLM Sky-Ready? A Comprehensive Spatial Intelligence Benchmark for UAV Navigation

Lingfeng Zhang, Yuchen Zhang, Hongsheng Li, Haoxiang Fu, Yingbo Tang, Hangjun Ye, Long Chen, Xiaojun Liang, Xiaoshuai Hao, Wenbo Ding

专题命中 安全评测 :safety(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.12633 2025-11-18 cs.CV 50%

Denoising Vision Transformer Autoencoder with Spectral Self-Regularization

Xunzhi Xiang, Xingye Tian, Guiyu Zhang, Yabo Chen, Shaofeng Zhang, Xuebo Wang, Xin Tao, Qi Fan

机构 * Nanjing University(南京大学) Kling Team, Kuaishou Technology(快手科技 Kling 团队) Chinese University of Hong Kong, Shenzhen(香港中文大学(深圳)) Shanghai Jiao Tong University(上海交通大学) University of Science and Technology of China(中国科学技术大学)

专题命中 安全评测 :alignment(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.12371 2025-11-18 cs.CV 50%

Reasoning Text-to-Video Retrieval via Digital Twin Video Representations and Large Language Models

Yiqing Shen, Chenxiao Fan, Chenjia Li, Mathias Unberath

机构 * Department of Computer Science, Johns Hopkins University(计算机科学系,约翰霍普金斯大学)

专题命中 安全评测 :alignment(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏