arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

视觉大模型 / VLM

视觉语言模型、视觉推理、视觉问答、图文理解和视觉 grounding。

2026-01-13 至 2026-01-13 共收录 9 信号源:cs.CV, cs.AI, cs.LG

1. 幻觉与鲁棒性 9 篇

2503.11742 2026-01-13 cs.CV cs.AI 81%

Safe Vision-Language Models via Unsafe Weights Manipulation

通过不安全权重操作实现安全的视觉-语言模型

Moreno D'Incà, Elia Peruzzo, Xingqian Xu, Humphrey Shi, Nicu Sebe, Massimiliano Mancini

机构 * University of Trento(特伦托大学) NVIDIA(NVIDIA公司) Georgia Tech(佐治亚理工学院)

专题命中 幻觉与鲁棒性 :vision-language model(title,abstract);分类 cs.CV、cs.AI

AI总结 本文提出UWM方法,通过不训练的方式提升视觉-语言模型在不安全查询上的安全性,同时在安全输入上表现更优。

Comments WACV 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.07812 2026-01-13 cs.CV 79%

More Images, More Problems? A Controlled Analysis of VLM Failure Modes

更多图像,更多问题?对VLV失败模式的受控分析

Anurag Das, Adrian Bulat, Alberto Baldrati, Ioannis Maniadis Metaxas, Bernt Schiele, Georgios Tzimiropoulos, Brais Martinez

机构 * MPI for Informatics Saarland Informatics Campus(马克斯·普朗克研究所(信息学)萨尔兰信息学校区) Samsung AI Cambridge(三星人工智能(剑桥)) Technical University of Iași Romania(罗文扎大学(罗马尼亚)) Queen Mary University of London UK(伦敦女王大学(英国))

专题命中 幻觉与鲁棒性 :VLM(title);vision language model(abstract);分类 cs.CV

AI总结 MIMIC基准通过诊断实验揭示LVLM在多图像任务中的失败模式,并提出数据生成和注意力遮罩方案以提升跨图像信息聚合能力。

Comments 19 pages, 16 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.06757 2026-01-13 cs.CL cs.AI 79%

MTMCS-Bench: Evaluating Contextual Safety of Multimodal Large Language Models in Multi-Turn Dialogues

MTMCS-Bench: 多轮对话中多模态大语言模型上下文安全性的评估

Zheyuan Liu, Dongwhi Kim, Yixin Wan, Xiangchi Yuan, Zhaoxuan Tan, Fengran Mo, Meng Jiang

机构 * University of Notre Dame(诺丁汉大学) University of California, Los Angeles(加州大学洛杉矶分校) Georgia Institute of Technology(佐治亚理工学院) University of Montreal(蒙特利尔大学)

专题命中 幻觉与鲁棒性 :multimodal large language model(title,abstract);分类 cs.AI

AI总结 MTMCS-Bench评估多模态大语言模型在多轮对话中的上下文安全性,揭示了安全与效用之间的权衡及现有防护措施的不足。

Comments A benchmark of realistic images and multi-turn conversations that evaluates contextual safety in MLLMs under two complementary settings

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.06464 2026-01-13 cs.CV 79%

On the Adversarial Robustness of 3D Large Vision-Language Models

关于3D大视觉-语言模型的对抗鲁棒性

Chao Liu, Ngai-Man Cheung

机构 * Singapore University of Technology and Design (SUTD)(新加坡科技设计大学)

专题命中 幻觉与鲁棒性 :vision-language model(title,abstract);分类 cs.CV

AI总结 本文研究了3D大视觉-语言模型的对抗鲁棒性,提出两种攻击策略评估其鲁棒性,发现其在无目标攻击下较脆弱,但在有目标攻击中更稳健。

Comments Under Review

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.06460 2026-01-13 cs.CV cs.AI cs.CL 73%

Tone Matters: The Impact of Linguistic Tone on Hallucination in VLMs

语气至关重要:语言语气对VLMs幻觉影响的研究

Weihao Hong, Zhiyuan Jiang, Bingyu Shen, Xinlei Guan, Yangyi Feng, Meng Xu, Boyang Li

机构 * Department of Computer Science and Technology, Kean University(计算机科学与技术系,凯恩大学) Department of Computer Science and Engineering, University of Notre Dame(计算机科学与工程系,圣母大学)

专题命中 幻觉与鲁棒性 :vision-language model(abstract);grounding(abstract);分类 cs.CV、cs.AI

AI总结 本文研究了提示语气对VLMs幻觉的影响,通过Ghost-100数据集发现幻觉率与提示强度非线性相关,揭示模型在处理结构性强制时的局限性。

Comments 10 pages, 6 figures, WACV Workshop

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.01592 2026-01-13 cs.CR cs.CV 70%

OpenRT: An Open-Source Red Teaming Framework for Multimodal LLMs

OpenRT: 一种用于多模态大语言模型的开源红队框架

Xin Wang, Yunhao Chen, Juncheng Li, Yixu Wang, Yang Yao, Tianle Gu, Jie Li, Yan Teng, Yingchun Wang, Xia Hu

机构 * Shanghai Artificial Intelligence Laboratory(上海人工智能实验室)

专题命中 幻觉与鲁棒性 :multimodal large language model(abstract);MLLM(abstract);分类 cs.CV

AI总结 OpenRT框架通过模块化设计和高吞吐量运行时,系统性评估多模态大语言模型的安全性,揭示了前沿模型在复杂攻击下的脆弱性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.07779 2026-01-13 cs.MA cs.AI cs.CL cs.CV cs.HC 62%

OS-Symphony: A Holistic Framework for Robust and Generalist Computer-Using Agent

OS-Symphony: 一个用于鲁棒且通用计算机使用代理的综合框架

Bowen Yang, Kaiming Jin, Zhenyu Wu, Zhaoyang Liu, Qiushi Sun, Zehao Li, JingJing Xie, Zhoumianze Liu, Fangzhi Xu, Kanzhi Cheng, Qingyun Li, Yian Wang, Yu Qiao, Zun Wang, Zichen Ding

机构 * University of Science and Technology of China(中国科学技术大学) Shanghai AI Laboratory(上海人工智能实验室) National University of Singapore(新加坡国立大学) The Hong Kong University of Science and Technology(香港科学与技术大学) The University of Hong Kong(香港大学) CUHK MMLab(香港中文大学MMLab) Xi’an Jiaotong University(西安交通大学) Nanjing University(南京大学) Harbin Institute of Technology(哈尔滨工业大学)

专题命中 幻觉与鲁棒性 :vision-language model(abstract);分类 cs.CV、cs.AI

AI总结 OS-Symphony通过反思记忆代理和多功能工具代理,提升计算机使用代理在长周期任务和新领域中的鲁棒性和泛化能力。

Comments 31 pages, 11 figures, 12 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2501.13772 2026-01-13 cs.SD cs.AI cs.LG cs.MM eess.AS 62%

Jailbreak-AudioBench: In-Depth Evaluation and Analysis of Jailbreak Threats for Large Audio Language Models

Jailbreak-AudioBench: 对大型音频语言模型中 jailbreak 威胁的深入评估与分析

Hao Cheng, Erjia Xiao, Jing Shao, Yichi Wang, Le Yang, Chao Shen, Philip Torr, Jindong Gu, Renjing Xu

机构 * Hong Kong University of Science and Technology (Guangzhou)(香港科技大学(广州)) University of Oxford(牛津大学) Xi’an Jiaotong University(西安交通大学) Hong Kong University of Science and Technology(香港科技大学) Northeastern University(东北大学) Beijing University of Technology(北京理工大学)

专题命中 幻觉与鲁棒性 :multimodal large language model(abstract);分类 cs.AI、cs.LG

AI总结 Jailbreak-AudioBench 通过构建工具箱、数据集和基准,深入评估大型音频语言模型中 jailbreak 威胁,并促进安全防护机制的发展。

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.07558 2026-01-13 cs.RO 50%

FlyCo: Foundation Model-Empowered Drones for Autonomous 3D Structure Scanning in Open-World Environments

FlyCo:基于基础模型的无人机自主3D结构扫描系统

Chen Feng, Guiyong Zheng, Tengkai Zhuang, Yongqian Wu, Fangzhan He, Haojia Li, Juepeng Zheng, Shaojie Shen, Boyu Zhou

机构 * Department of Electronic and Computer Engineering, The Hong Kong University of Science and Technology(香港科技大学电子与计算机工程系) School of Artificial Intelligence, Sun Yat-sen University(中山大学人工智能学院) Department of Mechanical and Energy Engineering, Southern University of Science and Technology(南方科技大学机械与能源工程系) Differential Robotics, Hangzhou, China(杭州差分机器人)

专题命中 幻觉与鲁棒性 :grounding(abstract)

AI总结 FlyCo通过整合基础模型实现无人机自主3D扫描,提升开放世界环境下的目标定位与预测效率。

Comments 34 pages, 24 figures, 9 tables. Video: https://www.youtube.com/playlist?list=PLqjZjnqsCyl40rw3y15Yzc7Mdo-z1y2j8

详情

展开后加载摘要…

URL PDF HTML 收藏