arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

大模型对齐与安全

大模型对齐、安全、越狱、红队、提示注入和可信评测。

共收录 9451 信号源:cs.CL, cs.AI, cs.CY, cs.LG

1. 安全评测 9451 篇

2511.20218 2025-11-26 cs.CV 50%

Text-guided Controllable Diffusion for Realistic Camouflage Images Generation

基于文本引导的可控扩散生成逼真伪装图像

Yuhang Qian, Haiyan Chen, Wentong Li, Ningzhong Liu, Jie Qin

专题命中 安全评测 :alignment(abstract)

AI总结 本文提出CT-CIG方法,通过文本引导和可控扩散生成逼真且逻辑合理的伪装图像,利用VLM和FIRM模块提升伪装图像质量。

Comments Accepted by AAAI 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2501.09436 2025-11-26 cs.CV 50%

Scaling up self-supervised learning for improved surgical foundation models

提升自监督学习以改进手术基础模型

Tim J. M. Jaspers, Ronald L. P. D. de Jong, Yiping Li, Carolus H. J. Kusters, Franciscus H. A. Bakker, Romy C. van Jaarsveld, Gino M. Kuiper, Richard van Hillegersberg, Jelle P. Ruurda, Willem M. Brinkman, Josien P. W. Pluim, Peter H. N. de With, Marcel Breeuwer, Yasmina Al Khalil, Fons van der Sommen

机构 * Department of Electrical Engineering, Video Coding \& Architectures, Eindhoven University of Technology, Eindhoven, The Netherlands Department of Biomedical Engineering, Medical Image Analysis, Eindhoven University of Technology, Eindhoven, The Netherlands Department of Surgery, University Medical Center Utrecht, Utrecht, The Netherlands Department of Oncological Urology, University Medical Center Utrecht, Utrecht, The Netherlands Department of Urology, Catharina Hospital, Eindhoven, The Netherlands

专题命中 安全评测 :safety(abstract)

AI总结 本研究提出SurgeNetXL,通过大规模预训练提升手术计算机视觉性能,实现多个任务上的显著改进。

Journal ref Medical Image Analysis, 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.19526 2025-11-26 cs.CV 50%

Perceptual Taxonomy: Evaluating and Guiding Hierarchical Scene Reasoning in Vision-Language Models

感知分类:评估和引导视觉语言模型中的层次场景推理

Jonathan Lee, Xingrui Wang, Jiawei Peng, Luoxin Ye, Zehan Zheng, Tiezheng Zhang, Tao Wang, Wufei Ma, Siyi Chen, Yu-Cheng Chou, Prakhar Kaushik, Alan Yuille

机构 * Johns Hopkins University(约翰霍普金斯大学)

专题命中 安全评测 :alignment(abstract)

AI总结 本文提出感知分类基准测试,旨在评估和引导视觉语言模型在层次场景推理中的能力,揭示模型在属性驱动推理上的不足,并展示通过上下文示例提升性能的效果。

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.19109 2025-11-25 cs.CV 50%

HABIT: Human Action Benchmark for Interactive Traffic in CARLA

HABIT: 交互交通中的人类行为基准

Mohan Ramesh, Mark Azer, Fabian B. Flohr

机构 * Intelligent Vehicles Lab, Munich University of Applied Sciences(智能车辆实验室,慕尼黑应用科学大学)

专题命中 安全评测 :safety(abstract)

AI总结 HABIT是一个高保真度的交互交通人类行为基准,通过整合现实人类运动数据提升自动驾驶模拟的真实性,揭示现有自动驾驶代理在复杂场景下的性能缺陷。

Comments Accepted to WACV 2026. This is the pre-camera-ready version

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.18983 2025-11-25 cs.CV 50%

UMCL: Unimodal-generated Multimodal Contrastive Learning for Cross-compression-rate Deepfake Detection

UMCL: 单模生成多模对比学习用于跨压缩率深度伪造检测

Ching-Yi Lai, Chih-Yu Jian, Pei-Cheng Chuang, Chia-Ming Lee, Chih-Chung Hsu, Chiou-Ting Hsu, Chia-Wen Lin

专题命中 安全评测 :alignment(abstract)

AI总结 UMCL通过单模生成多模对比学习,提升跨压缩率深度伪造检测的鲁棒性和准确性。

Comments 24-page manuscript accepted to IJCV

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.18921 2025-11-25 cs.CV 50%

BackdoorVLM: A Benchmark for Backdoor Attacks on Vision-Language Models

BackdoorVLM:面向视觉-语言模型的后门攻击基准

Juncheng Li, Yige Li, Hanxun Huang, Yunhao Chen, Xin Wang, Yixu Wang, Xingjun Ma, Yu-Gang Jiang

机构 * Fudan University(复旦大学) Singapore Management University(新加坡管理学院) The University of Melbourne(墨尔本大学)

专题命中 安全评测 :jailbreak(abstract)

AI总结 BackdoorVLM提出首个评估视觉-语言模型后门攻击的基准,揭示VLMs对文本指令的高敏感性及多模后门的有效性

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.18547 2025-11-25 physics.med-ph 50%

Towards Integrated Clinical-Computational Nuclear Medicine

迈向整合的临床-计算核医学

Faraz Farhadi, Shadi A. Esfahani, Fereshteh Yousefirizi, Monica Luo, Pedro Esquinas Fernandez, Arkadiusz Sitek, Hamid Sabet, Babak Saboury, Arman Rahmim, Pedram Heidari

专题命中 安全评测 :safety(abstract)

AI总结 本文探讨了整合临床与计算核医学的重要性,通过AI和PBPK建模提升影像质量和治疗个性化,强调临床监督在确保安全性和准确性中的关键作用。

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.10353 2025-11-25 cs.CV 50%

Motion-R1: Enhancing Motion Generation with Decomposed Chain-of-Thought and RL Binding

Motion-R1: 通过分解的链式推理与RL绑定增强运动生成

Runqi Ouyang, Haoyun Li, Zhenyuan Zhang, Xiaofeng Wang, Zeyu Zhang, Zheng Zhu, Guan Huang, Sirui Han, Xingang Wang

机构 * GigaAI CASIA HKUST(香港科技大学)

专题命中 安全评测 :alignment(abstract)

AI总结 Motion-R1通过结合分解的链式推理与强化学习,提升运动生成的质量和可解释性,实现多项指标的性能提升。

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.02298 2025-11-25 cs.IR 50%

A Zero-shot Explainable Doctor Ranking Framework with Large Language Models

基于大语言模型的零样本可解释医生排名框架

Ziyang Zeng, Dongyuan Li, Yuqing Yang

专题命中 安全评测 :trustworthy(abstract)

AI总结 本文提出基于大语言模型的零样本可解释医生排名框架,通过动态生成疾病特定的排名标准和逐步推理理由,提升医生排名的透明度和可解释性,并在DrRank数据集上取得显著性能提升。

Comments Accepted by Big Data Mining and Analytics (JCR Q1)

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.18486 2025-11-25 cs.RO cs.SY eess.SY 50%

Expanding the Workspace of Electromagnetic Navigation Systems Using Dynamic Feedback for Single- and Multi-agent Control

通过动态反馈扩展电磁导航系统的作业空间以实现单体和多体控制

Jasan Zughaibi, Denis von Arx, Maurus Derungs, Florian Heemeyer, Luca A. Antonelli, Quentin Boehler, Michael Muehlebach, Bradley J. Nelson

机构 * Multi-Scale Robotics Lab, ETH Zurich(多尺度机器人实验室,苏黎世联邦理工学院) Learning and Dynamical Systems Group, Max Planck Institute for Intelligent Systems(学习与动态系统组,马克斯·普朗克智能系统研究所)

专题命中 安全评测 :alignment(abstract)

AI总结 通过动态反馈和系统级控制设计扩展电磁导航系统的作业空间,实现单体和多体控制的稳定与高效

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.18352 2025-11-25 cs.CV 50%

MagicWand: A Universal Agent for Generation and Evaluation Aligned with User Preference

MagicWand: 一个通用代理用于生成与评估,与用户偏好对齐

Zitong Xu, Dake Shen, Yaosong Du, Kexiang Hao, Jinghan Huang, Xiande Huang

机构 * Shanghai Jiao Tong University(上海交通大学) De Artificial Intelligence Lab(德人工智能实验室) The Chinese University of Hong Kong, Shenzhen China(香港中文大学(深圳))

专题命中 安全评测 :alignment(abstract)

AI总结 MagicWand通过结合用户偏好数据集和先进生成模型,实现高质量内容生成与偏好对齐评估。

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.18343 2025-11-25 cs.SE 50%

A Needle in a Haystack: Intent-driven Reusable Artifacts Recommendation with LLMs

在 haystack 中寻找针:基于意图的可重用 artifacts 推荐 with LLMs

Dongming Jin, Zhi Jin, Xiaohong Chen, Zheng Fang, Linyu Li, Yuanpeng He, Jia Li, Yirang Zhang, Yingtao Fang

专题命中 安全评测 :alignment(abstract)

AI总结 本文提出 TreeRec 框架,通过语义抽象组织 artifacts 成层次结构树,提升 LLMs 在意图驱动推荐中的性能和效率。

Comments 15 pages, 7 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.17755 2025-11-25 cs.CV 50%

CORA: Consistency-Guided Semi-Supervised Framework for Reasoning Segmentation

CORA: 一致性引导的半监督框架用于推理分割

Prantik Howlader, Hoang Nguyen-Canh, Srijan Das, Jingyi Xu, Hieu Le, Dimitris Samaras

机构 * Stony Brook University(石溪大学) UNC-Charlotte(北卡罗来纳大学夏洛特分校)

专题命中 安全评测 :alignment(abstract)

AI总结 CORA通过一致性引导的半监督方法实现高效推理分割,在Cityscapes和PanNuke数据集上分别提升2.3%和2.4%

Comments WACV 2026 accepted

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.17735 2025-11-25 cs.CV 50%

Towards Open-Ended Visual Scientific Discovery with Sparse Autoencoders

基于稀疏自编码器的开放性视觉科学发现

Samuel Stevens, Jacob Beattie, Tanya Berger-Wolf, Yu Su

机构 * The Ohio State University(俄亥俄州立大学)

专题命中 安全评测 :alignment(abstract)

AI总结 本研究探讨稀疏自编码器在基础模型表示中实现开放性特征发现的潜力,通过生态图像案例展示其在无标签数据下的细粒度结构识别能力,为科学发现提供了新工具。

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.02188 2025-11-25 cs.CR 50%

Whispering Agents: An Event-driven Covert Communication Protocol For the Internet of Agents

低语代理:一种基于事件的隐秘通信协议用于代理互联网

Kaibo Huang, Yukun Wei, Wansheng Wu, Tianhua Zhang, Zhongliang Yang, Linna Zhou

专题命中 安全评测 :trustworthy(abstract)

AI总结 本文提出了一种基于事件的隐秘通信协议ΠCCAP,利用代理对话的特性实现安全隐秘通信,为代理互联网的安全监控和防御提供基础支持。

Comments Accepted to AAAI-26 (Main, Oral)

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.17448 2025-11-24 cs.CV 50%

MMT-ARD: Multimodal Multi-Teacher Adversarial Distillation for Robust Vision-Language Models

MMT-ARD: 多模态多教师对抗蒸馏用于鲁棒视觉-语言模型

Yuqi Li, Junhao Dong, Chuanguang Yang, Shiping Wen, Piotr Koniusz, Tingwen Huang, Yingli Tian, Yew-Soon Ong

机构 * The City University of New York, CUNY(纽约城市大学) Nanyang Technological University(南洋理工大学) Institute of Computing Technology, Chinese Academy of Sciences(中国科学院计算技术研究所) University of Technology Sydney(悉尼技术大学) Data61, CSIRO(CSIRO数据61研究所) Shenzhen University of Advanced Technology(深圳先进技术大学)

专题命中 安全评测 :safety(abstract)

AI总结 MMT-ARD通过多教师对抗蒸馏提升视觉-语言模型的对抗鲁棒性,实验显示鲁棒精度提升4.32%,训练效率提高2.3倍。

Comments 10 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.18267 2025-11-24 cs.HC 50%

From Checking to Sensemaking: A Caregiver-in-the-Loop Framework for AI-Assisted Task Verification in Dementia Care

从验证到理解:一种护理人员在循环框架中的AI辅助任务验证方法用于痴呆症护理

Joy Lai, Kelly Beaton, David Black, Bing Ye, Alex Mihailidis

专题命中 安全评测 :trustworthy(abstract)

AI总结 本文提出了一种护理人员在循环框架中的AI辅助任务验证方法,通过生成式AI支持照护者主导的任务验证,提升痴呆症护理中的信任与协作。

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.16369 2025-11-21 eess.SP cs.NI 50%

Reasoning Meets Representation: Envisioning Neuro-Symbolic Wireless Foundation Models

推理与表征的结合:展望神经符号无线基础模型

Jaron Fontaine, Mohammad Cheraghinia, John Strassner, Adnan Shahid, Eli De Poorter

专题命中 安全评测 :trustworthy(abstract)

AI总结 本文提出神经符号无线基础模型,结合神经网络与符号推理,以解决无线通信中的可解释性、鲁棒性和合规性问题,推动6G网络的智能化发展。

Comments Accepted at the 39th Conference on Neural Information Processing Systems (NeurIPS 2025) Workshop: AI and ML for Next-Generation Wireless Communications and Networking (AI4NextG)

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.12197 2025-11-21 cs.CV 50%

Beyond Patches: Mining Interpretable Part-Prototypes for Explainable AI

超越补丁:挖掘可解释的部分原型用于可解释AI

Mahdi Alehdaghi, Rajarshi Bhattacharya, Pourya Shamsolmoali, Rafael M. O. Cruz, Maguelonne Heritier, Eric Granger

专题命中 安全评测 :alignment(abstract)

AI总结 PCMNet通过挖掘可解释的部分原型提升AI系统的可解释性、稳定性和鲁棒性,为构建可靠且一致的AI系统提供新方法。

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.16150 2025-11-21 cs.CV 50%

Reasoning Guided Embeddings: Leveraging MLLM Reasoning for Improved Multimodal Retrieval

基于推理的嵌入:利用多模态大语言模型推理提升多模态检索

Chunxu Liu, Jiyuan Yang, Ruopeng Gao, Yuhan Zhu, Feng Zhu, Rui Zhao, Limin Wang

机构 * State Key Laboratory for Novel Software Technology, Nanjing University(南京大学新型软件技术国家重点实验室) Sensetime Research(商汤科技研究院) Beijing Institute of Technology(北京理工大学) Shanghai AI Lab(上海人工智能实验室)

专题命中 安全评测 :alignment(abstract)

AI总结 本文提出基于推理的嵌入方法,利用多模态大语言模型的推理能力提升多模态检索性能。

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.15308 2025-11-20 cs.CV 50%

Text2Loc++: Generalizing 3D Point Cloud Localization from Natural Language

Yan Xia, Letian Shi, Yilin Di, Joao F. Henriques, Daniel Cremers

机构 * School of Artificial Intelligence and Data Science, University of Science and Technology of China(人工智能与数据科学学院,中国科学技术大学) Technical University of Munich(慕尼黑技术大学) Visual Geometry Group, University of Oxford(牛津大学视觉几何组)

专题命中 安全评测 :alignment(abstract)

Comments This paper builds upon and extends our earlier conference paper Text2Loc presented at CVPR 2024

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.14391 2025-11-19 cs.CV 50%

Enhancing LLM-based Autonomous Driving with Modular Traffic Light and Sign Recognition

Fabian Schmidt, Noushiq Mohammed Kayilan Abdul Nazar, Markus Enzweiler, Abhinav Valada

机构 * Institute for Intelligent Systems(智能系统研究所) Esslingen University of Applied Sciences(应用科学大学埃斯林根) Department of Computer Science(计算机科学系) University of Freiburg(弗赖堡大学)

专题命中 安全评测 :safety(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.09611 2025-11-19 cs.CV 50%

MMaDA-Parallel: Multimodal Large Diffusion Language Models for Thinking-Aware Editing and Generation

Ye Tian, Ling Yang, Jiongfan Yang, Anran Wang, Yu Tian, Jiani Zheng, Haochen Wang, Zhiyang Teng, Zhuochen Wang, Yinjie Wang, Yunhai Tong, Mengdi Wang, Xiangtai Li

机构 * Peking University(北京大学) ByteDance(字节跳动) Princeton University(普林斯顿大学) CASIA(中国科学院自动化研究所) The University of Chicago(芝加哥大学)

专题命中 安全评测 :alignment(abstract)

Comments Project Page: https://tyfeld.github.io/mmadaparellel.github.io/

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.13369 2025-11-18 cs.SI physics.soc-ph 50%

Unifying points of interest taxonomies: mapping OpenStreetMap tags to the Foursquare category system

Lilou Soulas, Lorenzo Lucchini, Maurizio Napolitano, Sebastiano Bontorin, Simone Centellegher, Bruno Lepri, Riccardo Gallotti, Eleonora Andreotti

专题命中 安全评测 :alignment(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.13269 2025-11-18 cs.CV 50%

Is your VLM Sky-Ready? A Comprehensive Spatial Intelligence Benchmark for UAV Navigation

Lingfeng Zhang, Yuchen Zhang, Hongsheng Li, Haoxiang Fu, Yingbo Tang, Hangjun Ye, Long Chen, Xiaojun Liang, Xiaoshuai Hao, Wenbo Ding

专题命中 安全评测 :safety(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.12633 2025-11-18 cs.CV 50%

Denoising Vision Transformer Autoencoder with Spectral Self-Regularization

Xunzhi Xiang, Xingye Tian, Guiyu Zhang, Yabo Chen, Shaofeng Zhang, Xuebo Wang, Xin Tao, Qi Fan

机构 * Nanjing University(南京大学) Kling Team, Kuaishou Technology(快手科技 Kling 团队) Chinese University of Hong Kong, Shenzhen(香港中文大学(深圳)) Shanghai Jiao Tong University(上海交通大学) University of Science and Technology of China(中国科学技术大学)

专题命中 安全评测 :alignment(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.12371 2025-11-18 cs.CV 50%

Reasoning Text-to-Video Retrieval via Digital Twin Video Representations and Large Language Models

Yiqing Shen, Chenxiao Fan, Chenjia Li, Mathias Unberath

机构 * Department of Computer Science, Johns Hopkins University(计算机科学系,约翰霍普金斯大学)

专题命中 安全评测 :alignment(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.10974 2025-11-18 cs.MM cs.CV 50%

Failures to Surface Harmful Contents in Video Large Language Models

Yuxin Cao, Wei Song, Derui Wang, Jingling Xue, Jin Song Dong

专题命中 安全评测 :safety(abstract)

Comments 12 pages, 8 figures. Accepted to AAAI 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.11356 2025-11-17 cs.CR 50%

SEAL: Subspace-Anchored Watermarks for LLM Ownership

Yanbo Dai, Zongjie Li, Zhenlan Ji, Shuai Wang

专题命中 安全评测 :alignment(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.10668 2025-11-17 cs.CV 50%

A Mathematical Framework for AI Singularity: Conditions, Bounds, and Control of Recursive Improvement

Akbar Anbar Jafari, Cagri Ozcinar, Gholamreza Anbarjafari

机构 * University of Tartu(塔尔图大学) Loughborough University London(洛桑大学伦敦分校) Estonian Business School(爱沙尼亚商业学院)

专题命中 安全评测 :safety(abstract)

Comments 41 pages

详情

展开后加载摘要…

URL PDF HTML 收藏