arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

大模型对齐与安全

大模型对齐、安全、越狱、红队、提示注入和可信评测。

2025-11-18 至 2025-11-18 共收录 95 信号源:cs.CL, cs.AI, cs.CY, cs.LG

1. 安全评测 31 篇

2511.13626 2025-11-18 cs.AI 57%

CreBench: Human-Aligned Creativity Evaluation from Idea to Process to Product

Kaiwen Xue, Chenglong Li, Zhonghong Ou, Guoxin Zhang, Kaoyan Lu, Shuai Lyu, Yifan Zhu, Ping Zong Junpeng Ding, Xinyu Liu, Qunlin Chen, Weiwei Qin, Yiran Shen, Jiayi Cen

专题命中 安全评测 :alignment(abstract);分类 cs.AI

Comments 13 pages, 3 figures,The 40th Annual AAAI Conference on Artificial Intelligence(AAAI 2026),Paper has been accepted for a poster presentation

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.13169 2025-11-18 cs.CL 57%

TCM-5CEval: Extended Deep Evaluation Benchmark for LLM's Comprehensive Clinical Research Competence in Traditional Chinese Medicine

Tianai Huang, Jiayuan Chen, Lu Lu, Pengcheng Chen, Tianbin Li, Bing Han, Wenchao Tang, Jie Xu, Ming Li

机构 * School of Artificial Intelligence in Traditional Chinese Medicine, Shanghai University of Traditional Chinese Medicine, Shanghai, China(上海中医药大学人工智能学院) Shanghai Artificial Intelligence Laboratory, Shanghai, China(上海人工智能实验室) University of Washington, Seattle, Washington, US(华盛顿大学)

专题命中 安全评测 :alignment(abstract);分类 cs.CL

Comments 17 pages, 8 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.12928 2025-11-18 cs.CL 57%

Visual Room 2.0: Seeing is Not Understanding for MLLMs

Haokun Li, Yazhou Zhang, Jizhi Ding, Qiuchi Li, Peng Zhang

机构 * Tianjin University(天津大学) Shandong Institute of Petroleum and Chemical Technology(山东石油化学技术学院) Beijing Institute of Technology(北京理工大学)

专题命中 安全评测 :alignment(abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.12414 2025-11-18 cs.LG cs.CR 57%

The 'Sure' Trap: Multi-Scale Poisoning Analysis of Stealthy Compliance-Only Backdoors in Fine-Tuned Large Language Models

Yuting Tan, Yi Huang, Zhuo Li

专题命中 安全评测 :alignment(abstract);分类 cs.LG

Comments 13 pages, 5 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.12206 2025-11-18 cs.CV cs.AI 57%

A Novel AI-Driven System for Real-Time Detection of Mirror Absence, Helmet Non-Compliance, and License Plates Using YOLOv8 and OCR

Nishant Vasantkumar Hegde, Aditi Agarwal, Minal Moharir

机构 * Computer Science and Engineering(计算机科学与工程) RV College of Engineering(RV工程学院)

专题命中 安全评测 :safety(abstract);分类 cs.AI

Comments 6 pages, 4 figures. Published in: Proceedings of the 12th International Conference on Emerging Trends in Engineering Technology Signal and Information Processing (ICETET SIP 2025) Note: The conference proceedings contain an outdated abstract due to a publisher-side error. This arXiv version includes the correct and updated abstract

Journal ref 2025 IEEE 12th International Conference on Emerging Trends in Engineering Technology Signal & Information Processing (ICETET SIP 2025)

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.14741 2025-11-18 cs.CV cs.AI 57%

DEXTER: Diffusion-Guided EXplanations with TExtual Reasoning for Vision Models

Simone Carnemolla, Matteo Pennisi, Sarinda Samarasinghe, Giovanni Bellitto, Simone Palazzo, Daniela Giordano, Mubarak Shah, Concetto Spampinato

机构 * University of Catania(卡塔尼亚大学) University of Central Florida(中央佛罗里达大学)

专题命中 安全评测 :trustworthy(abstract);分类 cs.AI

Comments Accepted to NeurIPS 2025 (spotlight)

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.15249 2025-11-18 cs.CL cs.CV 57%

Fooling the LVLM Judges: Visual Biases in LVLM-Based Evaluation

Yerin Hwang, Dongryeol Lee, Kyungmin Min, Taegwan Kang, Yong-il Kim, Kyomin Jung

机构 * IPAI, Seoul National University(IPAI,首尔国立大学) Dept. of ECE, Seoul National University(电子工程系,首尔国立大学) LG AI Research(LG人工智能研究)

专题命中 安全评测 :alignment(abstract);分类 cs.CL

Comments EMNLP 2025 Main (21pgs, 12 Tables, 9 Figures)

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.12052 2025-11-18 cs.CR cs.AI 57%

Exploring AI in Steganography and Steganalysis: Trends, Clusters, and Sustainable Development Potential

Aditya Kumar Sahu, Chandan Kumar, Saksham Kumar, Serdar Solak

专题命中 安全评测 :alignment(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.10027 2025-11-18 cs.AI 57%

ChEmREF: Evaluating Language Model Readiness for Chemical Emergency Response

Risha Surana, Qinyuan Ye, Swabha Swayamdipta

机构 * University of Southern California(南加州大学)

专题命中 安全评测 :safety(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.13369 2025-11-18 cs.SI physics.soc-ph 50%

Unifying points of interest taxonomies: mapping OpenStreetMap tags to the Foursquare category system

Lilou Soulas, Lorenzo Lucchini, Maurizio Napolitano, Sebastiano Bontorin, Simone Centellegher, Bruno Lepri, Riccardo Gallotti, Eleonora Andreotti

专题命中 安全评测 :alignment(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.13269 2025-11-18 cs.CV 50%

Is your VLM Sky-Ready? A Comprehensive Spatial Intelligence Benchmark for UAV Navigation

Lingfeng Zhang, Yuchen Zhang, Hongsheng Li, Haoxiang Fu, Yingbo Tang, Hangjun Ye, Long Chen, Xiaojun Liang, Xiaoshuai Hao, Wenbo Ding

专题命中 安全评测 :safety(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.12633 2025-11-18 cs.CV 50%

Denoising Vision Transformer Autoencoder with Spectral Self-Regularization

Xunzhi Xiang, Xingye Tian, Guiyu Zhang, Yabo Chen, Shaofeng Zhang, Xuebo Wang, Xin Tao, Qi Fan

机构 * Nanjing University(南京大学) Kling Team, Kuaishou Technology(快手科技 Kling 团队) Chinese University of Hong Kong, Shenzhen(香港中文大学(深圳)) Shanghai Jiao Tong University(上海交通大学) University of Science and Technology of China(中国科学技术大学)

专题命中 安全评测 :alignment(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.12371 2025-11-18 cs.CV 50%

Reasoning Text-to-Video Retrieval via Digital Twin Video Representations and Large Language Models

Yiqing Shen, Chenxiao Fan, Chenjia Li, Mathias Unberath

机构 * Department of Computer Science, Johns Hopkins University(计算机科学系,约翰霍普金斯大学)

专题命中 安全评测 :alignment(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.10974 2025-11-18 cs.MM cs.CV 50%

Failures to Surface Harmful Contents in Video Large Language Models

Yuxin Cao, Wei Song, Derui Wang, Jingling Xue, Jin Song Dong

专题命中 安全评测 :safety(abstract)

Comments 12 pages, 8 figures. Accepted to AAAI 2026

详情

展开后加载摘要…

URL PDF HTML 收藏

2. AI治理与伦理 5 篇

2511.12271 2025-11-18 cs.AI 85%

MoralReason: Generalizable Moral Decision Alignment For LLM Agents Using Reasoning-Level Reinforcement Learning

Zhiyu An, Wan Du

专题命中 AI治理与伦理 :alignment(title,abstract);safety(abstract);AI safety(abstract);分类 cs.AI

Comments Accepted for AAAI 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.12689 2025-11-18 cs.CY cs.AI 76%

From Delegates to Trustees: How Optimizing for Long-Term Interests Shapes Bias and Alignment in LLM

Suyash Fulay, Jocelyn Zhu, Michiel Bakker

机构 * MIT(麻省理工学院)

专题命中 AI治理与伦理 :alignment(title);分类 cs.AI、cs.CY

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.11790 2025-11-18 cs.CY cs.AI 62%

Differences in the Moral Foundations of Large Language Models

Peter Kirgis

专题命中 AI治理与伦理 :alignment(abstract);分类 cs.AI、cs.CY

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.10089 2025-11-18 cs.LG cs.AI 62%

T2IBias: Uncovering Societal Bias Encoded in the Latent Space of Text-to-Image Generative Models

Abu Sufian, Cosimo Distante, Marco Leo, Hanan Salam

机构 * National Research Council of Italy - Institute of Applied Sciences

专题命中 AI治理与伦理 :alignment(abstract);分类 cs.AI、cs.LG

Comments This manuscript has been accepted for presentation in the First Interdisciplinary Workshop on Responsible AI for Value Creation. Dec 1, Copenhagen. The final version will be submitted for inclusion in a Springer LNCS Volume. (The paper is 15 pages with 7 figures)

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.11789 2025-11-18 cs.MA cs.AI 57%

From Single to Societal: Analyzing Persona-Induced Bias in Multi-Agent Interactions

Jiayi Li, Xiao Liu, Yansong Feng

专题命中 AI治理与伦理 :trustworthy(abstract);分类 cs.AI

Comments AAAI-2026

详情

展开后加载摘要…

URL PDF HTML 收藏

3. 其他安全 16 篇

2506.09353 2025-11-18 cs.CR cs.CV 88%

DAVSP: Safety Alignment for Large Vision-Language Models via Deep Aligned Visual Safety Prompt

Yitong Zhang, Jia Li, Liyi Cai, Ge Li

专题命中 其他安全 :alignment(title,abstract);safety(title,abstract)

Comments 16 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.09287 2025-11-18 cs.AI cs.CY cs.LG 67%

From Model Training to Model Raising

Roland Aydin, Christian Cyron, Steve Bachelor, Ashton Anderson, Robert West

机构 * Hamburg University of Technology(汉堡理工大学) Helmholtz-Zentrum Hereon(海德堡研究中心) University of Toronto(多伦多大学) EPFL(苏黎世联邦理工学院)

专题命中 其他安全 :alignment(abstract);分类 cs.AI、cs.CY、cs.LG

Comments Accepted for publication in Communications of the ACM (CACM), Opinion section

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.11622 2025-11-18 cs.LG cs.AI cs.CL 67%

Small Vocabularies, Big Gains: Pretraining and Tokenization in Time Series Models

Alexis Roger, Gwen Legate, Kashif Rasul, Yuriy Nevmyvaka, Irina Rish

专题命中 其他安全 :alignment(abstract);分类 cs.CL、cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.12630 2025-11-18 cs.CL cs.AI 62%

Knots: A Large-Scale Multi-Agent Enhanced Expert-Annotated Dataset and LLM Prompt Optimization for NOTAM Semantic Parsing

Maoqi Liu, Quan Fang, Yang Yang, Can Zhao, Kaiquan Cai

机构 * Beijing University of Posts and Telecommunications(北京邮电大学) Beihang University(北航) State Key Laboratory of CNS/ATM(国家空管流量管理技术实验室) Aviation Data Communication Corporation(航空数据通信公司)

专题命中 其他安全 :safety(abstract);分类 cs.CL、cs.AI

Comments Accepted to Advanced Engineering Informatics

详情

展开后加载摘要…

URL PDF HTML 收藏
2408.04998 2025-11-18 cs.CL cs.AI 62%

ProFuser: Progressive Fusion of Large Language Models

Tianyuan Shi, Fanqi Wan, Canbin Huang, Xiaojun Quan, Chenliang Li, Ming Yan, Ji Zhang, Minhua Huang, Wu Kai

专题命中 其他安全 :safety(abstract);分类 cs.CL、cs.AI

Comments Accepted to AAAI 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.12452 2025-11-18 cs.CV cs.CL 57%

DenseAnnotate: Enabling Scalable Dense Caption Collection for Images and 3D Scenes via Spoken Descriptions

Xiaoyu Lin, Aniket Ghorpade, Hansheng Zhu, Justin Qiu, Dea Rrozhani, Monica Lama, Mick Yang, Zixuan Bian, Ruohan Ren, Alan B. Hong, Jiatao Gu, Chris Callison-Burch

机构 * University of Pennsylvania(宾夕法尼亚大学)

专题命中 其他安全 :alignment(abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.18320 2025-11-18 cs.LG 57%

State of Health Estimation of Batteries Using a Time-Informed Dynamic Sequence-Inverted Transformer

Janak M. Patel, Milad Ramezankhani, Anirudh Deodhar, Dagnachew Birru

机构 * Applied Research, Quantiphi(Quantiphi应用研究)

专题命中 其他安全 :safety(abstract);分类 cs.LG

Comments 11 pages, 3 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.19848 2025-11-18 cs.RO cs.AI 57%

Human-Centered AI and Autonomy in Robotics: Insights from a Bibliometric Study

Simona Casini, Pietro Ducange, Francesco Marcelloni, Lorenzo Pollini

机构 * Department of Information Engineering, University of Pisa(信息工程系,比萨大学)

专题命中 其他安全 :safety(abstract);分类 cs.AI

Comments International Joint Conference on Neural Network 2025 - Accepted

Journal ref 10.1109/IJCNN64981.2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.12167 2025-11-18 cond-mat.mtrl-sci cs.LG 57%

Rapid Machine Learning-Driven Detection of Pesticides and Dyes Using Raman Spectroscopy

Quach Thi Thai Binh, Thuan Phuoc, Xuan Hai, Thang Bach Phan, Vu Thi Hanh Thu, Nguyen Tuan Hung

机构 * Faculty of Physics and Physics Engineering, University of Science, Ho Chi Minh City 700000, Viet Nam(物理系和物理工程系,科学大学,胡志明市700000,越南) Center for Innovative Materials and Architectures (INOMAR)(创新材料与架构中心) Department of Materials Science and Engineering, National Taiwan University, Taipei 10617, Taiwan(材料科学与工程系,台湾国立大学,台北10617,台湾)

专题命中 其他安全 :safety(abstract);分类 cs.LG

Comments 25 pages, 9 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.11885 2025-11-18 cs.DC cs.AI cs.DB 57%

Flash-Fusion: Enabling Expressive, Low-Latency Queries on IoT Sensor Streams with LLMs

Kausar Patherya, Ashutosh Dhekne, Francisco Romero

机构 * Georgia Institute of Technology(佐治亚理工学院)

专题命中 其他安全 :safety(abstract);分类 cs.AI

Comments 12 pages, 5 figures. Under review

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.11780 2025-11-18 cs.CV cs.AI 57%

Image-POSER: Reflective RL for Multi-Expert Image Generation and Editing

Hossein Mohebbi, Mohammed Abdulrahman, Yanting Miao, Pascal Poupart, Suraj Kothawade

机构 * University of Waterloo(滑铁卢大学) Vector Institute(向量研究所) Google(谷歌)

专题命中 其他安全 :alignment(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏