arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

视觉大模型 / VLM

视觉语言模型、视觉推理、视觉问答、图文理解和视觉 grounding。

2025-12-17 至 2025-12-17 共收录 40 信号源:cs.CV, cs.AI, cs.LG

1. 视觉问答 5 篇

2503.20047 2025-12-17 cs.CV eess.IV 85%

Med3DVLM: An Efficient Vision-Language Model for 3D Medical Image Analysis

Med3DVLM: 一种高效的视觉-语言模型用于3D医学图像分析

Yu Xin, Gorkem Can Ates, Kuang Gong, Wei Shao

机构 * University of Florida(佛罗里达大学)

专题命中 视觉问答 :vision-language model(title,abstract);VLM(abstract);visual question answering(abstract);分类 cs.CV

AI总结 Med3DVLM通过三个创新提出,实现了高效的3D医学图像分析,显著提升了图像-文本检索、报告生成和视觉问答的性能。

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.14017 2025-12-17 cs.CV cs.AI 62%

KFS-Bench: Comprehensive Evaluation of Key Frame Sampling in Long Video Understanding

KFS-Bench: 长视频理解中关键帧采样的全面评估

Zongyao Li, Kengo Ishida, Satoshi Yamazaki, Xiaotong Ji, Jianquan Liu

机构 * Visual Intelligence Research Laboratories, NEC Corporation(NEC公司视觉智能研究实验室)

专题命中 视觉问答 :multimodal large language model(abstract);分类 cs.CV、cs.AI

AI总结 KFS-Bench通过多场景标注评估关键帧采样策略,提出新度量标准和方法提升问答性能。

Comments WACV2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.13164 2025-12-17 cs.CV cs.AI 62%

A Semantically Enhanced Generative Foundation Model Improves Pathological Image Synthesis

语义增强的生成基础模型提升病理图像合成

Xianchao Guan, Zhiyuan Fan, Yifeng Wang, Fuqiang Chen, Yanjiang Zhou, Zengyang Che, Hongxue Meng, Xin Li, Yaowei Wang, Hongpeng Wang, Min Zhang, Heng Tao Shen, Zheng Zhang, Yongbing Zhang

专题命中 视觉问答 :visual question answering(abstract);分类 cs.CV、cs.AI

AI总结 CRAFTS通过语义增强的生成基础模型提升病理图像合成质量,生成多样化病理图像并增强多种临床任务性能。

Comments 68 pages, 9 figures, 16 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.11872 2025-12-17 cs.RO cs.AI cs.CV 62%

WAM-Diff: A Masked Diffusion VLA Framework with MoE and Online Reinforcement Learning for Autonomous Driving

WAM-Diff: 一种结合MoE和在线强化学习的掩码扩散VLA框架用于自动驾驶

Mingwang Xu, Jiahao Cui, Feipeng Cai, Hanlin Shang, Zhihao Zhu, Shan Luan, Yifang Xu, Neng Zhang, Yaoyi Li, Jia Cai, Siyu Zhu

机构 * Fudan University(复旦大学) Yinwang Intelligent Technology Co., Ltd(云网智能科技有限公司)

专题命中 视觉问答 :visual question answering(abstract);分类 cs.CV、cs.AI

AI总结 WAM-Diff通过结合MoE和在线强化学习的掩码扩散框架,提升自动驾驶轨迹生成的性能与灵活性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.13177 2025-12-17 cs.CV cs.RO 57%

MMDrive: Interactive Scene Understanding Beyond Vision with Multi-representational Fusion

MMDrive: 通过多表示融合超越视觉的交互场景理解

Minghui Hou, Wei-Hsing Huang, Shaofeng Liang, Daizong Liu, Tai-Hao Wen, Gang Wang, Runwei Guan, Weiping Ding

机构 * organization= College of Computer Science Technology, Jilin University , city= Changchun , country= China organization= Georgia Institute of Technology , city= Atlanta , country= USA organization= Qingdao Institute of Software, College of Computer Science Technology, China University of Petroleum (East China) , city= Qingdao , country= China organization= Institute for Math \& AI, Wuhan University , city= Wuhan , country= China organization= University of Michigan, Ann Arbor , country= USA organization= Thrust of Artificial Intelligence, Hong Kong University of Science organization= School of Artificial Intelligence Computer Science, Nantong University , city= Nantong , country= China

专题命中 视觉问答 :vision-language model(abstract);分类 cs.CV

AI总结 MMDrive通过融合占用图、LiDAR点云和文本描述,实现超越视觉的三维场景理解,提升自动驾驶的多模态推理能力。

详情

展开后加载摘要…

URL PDF HTML 收藏

2. 视觉推理 10 篇

2506.11162 2025-12-17 cs.CV cs.LG 84%

VIBE: Can a VLM Read the Room?

VIBE:视觉语言模型能否读懂房间?

Tania Chakraborty, Eylon Caplan, Dan Goldwasser

机构 * Purdue University(普渡大学)

专题命中 视觉推理 :VLM(title,abstract);vision language model(abstract);分类 cs.CV、cs.LG

AI总结 本文提出视觉社交-语用推理任务,揭示VLM在社交推理中的局限性,并通过高质量数据集评估多种VLM的性能。

Comments Findings of EMNLP, 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.14257 2025-12-17 cs.CV 79%

Enhancing Visual Programming for Visual Reasoning via Probabilistic Graphs

通过概率图增强视觉编程用于视觉推理

Wentao Wan, Kaiyu Wu, Qingyang Ma, Nan Kang, Yunjie Chen, Liang Lin, Keze Wang

机构 * School of Computer Science and Engineering, Sun Yat-sen University(计算机科学与工程学院,中山大学)

专题命中 视觉推理 :visual reasoning(title,abstract);分类 cs.CV

AI总结 通过构建概率图解决视觉编程非可微问题,提升视觉推理任务的性能。

Comments 13 Pages, 12 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.14052 2025-12-17 cs.CV cs.CL 79%

HyperVL: An Efficient and Dynamic Multimodal Large Language Model for Edge Devices

HyperVL: 一种高效的多模态大语言模型用于边缘设备

HyperAI Team, Yuchen Liu, Kaiyang Han, Zhiqiang Xia, Yuhang Dong, Chen Song, Kangyu Tang, Jiaming Xu, Xiushi Feng, WenXuan Yu, Li Peng, Mingyang Wang, Kai Wang, Changpeng Yang, Yang Li, Haoyu Lu, Hao Wang, Bingna Xu, Guangyao Liu, Long Huang, Kaibin Guo, Jinyang Wu, Dan Wu, Hongzhen Wang, Peng Zhou, Shuai Nie, Shande Wang, Runyu Shi, Ying Huang

机构 * HyperAI Team(HyperAI团队) Xiaomi Corporation(小米公司)

专题命中 视觉推理 :multimodal large language model(title,abstract);分类 cs.CV

AI总结 HyperVL是一种为边缘设备优化的高效多模态大语言模型,通过图像分块、视觉分辨率压缩和双一致性学习技术,实现低延迟、低功耗的多模态推理。

Comments Technical report of Xiaomi HyperAI Team

详情

展开后加载摘要…

URL PDF HTML 收藏
2412.07148 2025-12-17 cs.CV cs.AI cs.CL cs.LG 75%

MM-PoE: Multiple Choice Reasoning via. Process of Elimination using Multi-Modal Models

MM-PoE:通过多模态模型的排除过程进行多选推理

Sayak Chakrabarty, Souradip Pal

专题命中 视觉推理 :vision-language model(abstract);visual reasoning(abstract);分类 cs.CV、cs.AI、cs.LG

AI总结 MM-PoE通过多模态模型的排除过程提升视觉语言模型在多选推理任务中的性能。

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.14442 2025-12-17 cs.CV cs.RO 70%

A4-Agent: An Agentic Framework for Zero-Shot Affordance Reasoning

A4-Agent: 一种用于零样本 affordance 推理的代理框架

Zixin Zhang, Kanghao Chen, Hanqing Wang, Hongfei Zhang, Harold Haodong Chen, Chenfei Liao, Litao Guo, Ying-Cong Chen

机构 * HKUST(GZ)(香港科技大学(广州)) HKUST(香港科技大学) SJTU(上海交通大学) Knowin

专题命中 视觉推理 :vision-language model(abstract);grounding(abstract);分类 cs.CV

AI总结 A4-Agent 提出一种无需训练的代理框架,通过三个阶段的流水线实现零样本 affordance 推理,优于现有监督方法。

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.13752 2025-12-17 cs.CV cs.AI 62%

STAR: STacked AutoRegressive Scheme for Unified Multimodal Learning

STAR:用于统一多模态学习的堆叠自回归方案

Jie Qin, Jiancheng Huang, Limeng Qiao, Lin Ma

机构 * Meituan Inc(美团公司)

专题命中 视觉推理 :multimodal large language model(abstract);分类 cs.CV、cs.AI

AI总结 STAR通过堆叠自回归方案提升多模态生成性能,同时保持理解能力,实验验证其在统一多模态学习中的有效性。

Comments 18 pages, 7 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.13747 2025-12-17 cs.CV cs.AI 62%

Why Text Prevails: Vision May Undermine Multimodal Medical Decision Making

为何文本占上风:视觉可能损害多模态医疗决策制定

Siyuan Dai, Lunxiao Li, Kun Zhao, Eardi Lila, Paul K. Crane, Heng Huang, Dongkuan Xu, Haoteng Tang, Liang Zhan

机构 * University of Texas Rio Grande Valley(德克萨斯大学里奥格兰德谷大学) University of Pittsburgh(匹兹堡大学) NC State University(北卡罗来纳州立大学) University of Washington(华盛顿大学) University of Maryland(马里兰大学)

专题命中 视觉推理 :multimodal large language model(abstract);分类 cs.CV、cs.AI

AI总结 本研究发现文本推理在医疗多模态决策中优于多模态输入,提出三种策略以提升多模态医疗决策能力。

Comments Accepted by ICDM 2025 the Workshop on Synergy of AI and Multimodal Biomedical Data Mining

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.11827 2025-12-17 cs.CY cs.AI cs.CV 62%

Assessing Greenspace Attractiveness with ChatGPT, Claude, and Gemini: Do AI Models Reflect Human Perceptions?

利用ChatGPT、Claude和Gemini评估绿地吸引力:AI模型能反映人类感知吗?

Milad Malekzadeh, Magdalena Biernacka, Elias Willberg, Jussi Torkko, Edyta Łaszkiewicz, Tuuli Toivonen

专题命中 视觉推理 :multimodal large language model(abstract);分类 cs.CV、cs.AI

AI总结 本文研究了AI模型在评估绿地吸引力方面的表现,发现其在正式绿地和非正式空间的判断一致性较高,但存在对安全性和本地嵌入质量的低估问题。

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.16313 2025-12-17 cs.LG cs.AI cs.CL 62%

Retrieval Enhanced Feedback via In-context Neural Error-book

通过上下文神经错误本增强的检索反馈

Jongyeop Hyun, Bumsoo Kim

机构 * School of CSE Chung-Ang University(计算机科学与工程学院 Chung-Ang 大学)

专题命中 视觉推理 :multimodal large language model(abstract);分类 cs.AI、cs.LG

AI总结 REFINE通过结构化反馈框架,系统分析并缓解多模态推理中的错误,提升推理效率和可扩展性。

Comments Accepted at EMNLP 2025 main

详情

展开后加载摘要…

URL PDF HTML 收藏
2410.07553 2025-12-17 cs.AI 57%

COMMA: A Communicative Multimodal Multi-Agent Benchmark

COMMA:一种基于通信的多模态多智能体基准

Timothy Ossowski, Danyal Maqbool, Jixuan Chen, Zefan Cai, Tyler Bradshaw, Junjie Hu

机构 * Department of Computer Sciences University of Wisconsin-Madison(计算机科学系威斯康星大学麦迪逊分校) Department of Computer Sciences UC San Diego(计算机科学系加州大学圣地亚哥分校) Department of Radiology University of Wisconsin-Madison(放射学系威斯康星大学麦迪逊分校) Department of Computer Sciences Department of Biostatistics and Medical Informatics University of Wisconsin-Madison(计算机科学系生物统计学与医学信息学系威斯康星大学麦迪逊分校)

专题命中 视觉推理 :LLaVA(abstract);分类 cs.AI

AI总结 COMMA基准通过语言通信评估多模态多智能体系统的协作性能,揭示现有模型在智能体协作中的不足。

Journal ref Transactions on Machine Learning Research, 2025

详情

展开后加载摘要…

URL PDF HTML 收藏

3. 视觉定位与Grounding 9 篇

2512.12012 2025-12-17 cs.CV cs.AI cs.CL cs.RO 88%

Semantic-Drive: Democratizing Long-Tail Data Curation via Open-Vocabulary Grounding and Neuro-Symbolic VLM Consensus

语义驱动:通过开放词汇锚定和神经符号视觉语言共识民主化长尾数据整理

Antonio Guillen-Perez

机构 * Independent Researcher(独立研究者)

专题命中 视觉定位与Grounding :VLM(title,abstract);grounding(title,abstract);分类 cs.CV、cs.AI

AI总结 Semantic-Drive通过开放词汇锚定和神经符号视觉语言共识,提升自动驾驶中长尾数据整理的效率与隐私保护。

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.13771 2025-12-17 cs.AI 79%

Semantic Grounding Index: Geometric Bounds on Context Engagement in RAG Systems

语义 grounding 指数:RAG 系统中上下文参与的几何界限

Javier Marín

机构 * CERT

专题命中 视觉定位与Grounding :grounding(title,abstract);分类 cs.AI

AI总结 本文提出语义 grounding 指数(SGI),通过几何角度分析 RAG 系统中响应与问题、上下文之间的关系,揭示幻觉响应在角度上接近问题而非上下文,验证了 SGI 在评估响应真实性中的有效性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.14312 2025-12-17 cs.CV cs.AI 73%

From YOLO to VLMs: Advancing Zero-Shot and Few-Shot Detection of Wastewater Treatment Plants Using Satellite Imagery in MENA Region

从YOLO到VLMs:利用卫星图像在中东和北非地区推进零样本和少样本废水处理厂检测

Akila Premarathna, Kanishka Hewageegana, Garcia Andarcia Mariangel

专题命中 视觉定位与Grounding :vision-language model(abstract);VLM(abstract);分类 cs.CV、cs.AI

AI总结 本研究利用VLMs替代YOLOv8,通过零样本和少样本方法高效识别中东和北非地区废水处理厂,提升遥感应用的可扩展性。

Comments 9 pages, 9 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.12492 2025-12-17 cs.CV cs.CL 70%

Adaptive Detector-Verifier Framework for Zero-Shot Polyp Detection in Open-World Settings

面向开放世界设置的零样本息肉检测自适应检测-验证框架

Shengkai Xu, Hsiang Lun Kao, Tianxiang Xu, Honghui Zhang, Junqiao Wang, Runmeng Ding, Guanyu Liu, Tianyu Shi, Zhenyu Yu, Guofeng Pan, Ziqian Bi, Yuqi Ouyang

机构 * College of Computer Science, Sichuan University(四川大学计算机学院) Columbia University(哥伦比亚大学) School of Software and Microelectronics, Peking University(北京大学软件与微电子学院) Apon AI and Brain-Computer Engineering Research Institute(Apon人工智能与脑机工程研究院) Faculty of Science and Technology, University of Macau(澳门大学科学与技术学院) Faculty of Applied Science and Engineering, University of Toronto(多伦多大学应用科学与工程学院) Faculty of Computer Science and Information Technology, University of Malaya(马来亚大学计算机科学与信息技术学院) Zhaolong Technology(智龙科技) Purdue University(普渡大学)

专题命中 视觉定位与Grounding :vision-language model(abstract);VLM(abstract);分类 cs.CV

AI总结 本文提出AdaptiveDetector,通过自适应阈值调整和成本敏感强化学习,实现开放世界中零样本息肉检测,提升召回率并减少假阴性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.14273 2025-12-17 cs.CV 57%

Zoom-Zero: Reinforced Coarse-to-Fine Video Understanding via Temporal Zoom-in

Zoom-Zero: 通过时间放大实现强化的粗到细视频理解

Xiaoqian Shen, Min-Hung Chen, Yu-Chiang Frank Wang, Mohamed Elhoseiny, Ryo Hachiuma

机构 * NVIDIA(英伟达) KAUST(卡塔尔人工智能研究所在线大学)

专题命中 视觉定位与Grounding :grounding(abstract);分类 cs.CV

AI总结 Zoom-Zero通过时间放大和令牌选择性信用分配提升视频问题回答的时空定位精度和答案准确性。

Comments Project page: https://xiaoqian-shen.github.io/Zoom-Zero/

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.02503 2025-12-17 cs.CV 57%

Adapting General-Purpose Foundation Models for X-ray Ptychography in Low-Data Regimes

为低数据情形下的X射线衍射成像适应通用基础模型

Robinson Umeike, Neil Getty, Yin Xiangyu, Yi Jiang

机构 * The University of Alabama(阿拉巴马大学) Argonne National Laboratory(阿贡国家实验室)

专题命中 视觉定位与Grounding :vision-language model(abstract);分类 cs.CV

AI总结 本文提出PtychoBench基准,通过比较SFT和ICL策略,在低数据环境下优化X射线衍射成像任务的模型适应性,发现任务模态决定最佳专门化路径。

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.04687 2025-12-17 cs.CV 57%

Guideline-Consistent Segmentation via Multi-Agent Refinement

通过多智能体细化实现指南一致的分割

Vanshika Vats, Ashwani Rathee, James Davis

专题命中 视觉定位与Grounding :vision-language model(abstract);分类 cs.CV

AI总结 本文提出一种多智能体无训练框架,通过Worker-Supervisor迭代细化架构实现指南一致的分割,有效应对复杂文本指南。

Comments To be published in The Fortieth AAAI Conference on Artificial Intelligence (AAAI 2026)

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.11252 2025-12-17 cs.CV eess.IV 57%

MFGDiffusion: Mask-Guided Smoke Synthesis for Enhanced Forest Fire Detection

MFGDiffusion:基于掩码的烟雾合成以提升森林火灾检测

Guanghao Wu, Yunqing Shang, Chen Xu, Hai Song, Chong Wang, Qixing Zhang

机构 * University of Science and Technology of China(中国科学技术大学)

专题命中 视觉定位与Grounding :multimodal large language model(abstract);分类 cs.CV

AI总结 MFGDiffusion通过掩码引导的烟雾合成提升森林火灾检测性能

Comments 14 pages, 11 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.05310 2025-12-17 cs.CL cs.SI 50%

Listening Between the Lines: Decoding Podcast Narratives with Language Modeling

在字里行间倾听:利用语言模型解码播客叙事

Shreya Gupta, Ojasva Saxena, Arghodeep Nandi, Sarah Masud, Kiran Garimella, Tanmoy Chakraborty

机构 * Indian Institute Of Technology Delhi(印度理工学院德里分校) University of Copenhagen(哥本哈根大学) Rutgers University(罗格斯大学)

专题命中 视觉定位与Grounding :grounding(abstract)

AI总结 本文提出一种基于 BERT 的方法,通过标注叙事框架与对话实体的关系,揭示播客中主题与框架的系统关联,提升对数字媒体影响的分析能力。

Comments 10 pages, 6 Figures, 5 Tables. Under review

详情

展开后加载摘要…

URL PDF HTML 收藏

4. 文档图表理解 2 篇

2512.14040 2025-12-17 cs.CV cs.LG 62%

ChartAgent: A Chart Understanding Framework with Tool Integrated Reasoning

ChartAgent: 一种集成工具推理的图表理解框架

Boran Wang, Xinming Wang, Yi Chen, Xiang Li, Jian Xu, Jing Yuan, Chenglin Liu

机构 * College of Artificial Intelligence, Nankai University(人工智能学院,南开大学) University of Chinese Academy of Sciences(中国科学院大学) State Key Laboratory of Multimodal Artificial Intelligence Systems (MAIS), Institution of Automation, Chinese Academy of Sciences(多模态人工智能系统国家重点实验室(MAIS),自动化研究所,中国科学院) Zhongguancun Academy(中关村学院) State Key Laboratory for Novel Software Technology, Nanjing University(新型软件技术国家重点实验室,南京大学)

专题命中 文档图表理解 :multimodal large language model(abstract);分类 cs.CV、cs.LG

AI总结 ChartAgent通过集成工具推理框架提升图表理解的鲁棒性与可追溯性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.09554 2025-12-17 cs.IR 50%

Mixture-of-RAG: Integrating Text and Tables with Large Language Models

混合RAG:利用大语言模型整合文本和表格

Chi Zhang, Qiyang Chen, Mengqi Zhang

专题命中 文档图表理解 :grounding(abstract)

AI总结 MixRAG通过三阶段框架整合文本和表格,提升异构文档检索性能,实现混合模态文档接地的最新成果。

Comments Accepted to SIGKDD 2026

详情

展开后加载摘要…

URL PDF HTML 收藏

5. GUI与屏幕智能体 3 篇

2512.13974 2025-12-17 cs.RO 82%

Autonomous Construction-Site Safety Inspection Using Mobile Robots: A Multilayer VLM-LLM Pipeline

基于移动机器人的自主施工工地安全检查:一种多层VLM-LLM流水线

Hossein Naderi, Alireza Shojaei, Philip Agee, Kereshmeh Afsari, Abiola Akanmu

专题命中 GUI与屏幕智能体 :VLM(title,abstract);vision language model(abstract)

AI总结 本文提出一种多层VLM-LLM流水线,通过机器人自主导航与人工智能结合,实现施工工地安全检查的自动化报告生成。

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.14014 2025-12-17 cs.AI 70%

MobileWorldBench: Towards Semantic World Modeling For Mobile Agents

MobileWorldBench: 向移动智能体的语义世界建模迈进

Shufan Li, Konstantinos Kallidromitis, Akash Gokul, Yusuke Kato, Kazuki Kozuka, Aditya Grover

机构 * UCLA(加州大学洛杉矶分校) Panasonic AI Research(松下人工智能研究) Salesforce AI Research(Salesforce人工智能研究)

专题命中 GUI与屏幕智能体 :vision-language model(abstract);VLM(abstract);分类 cs.AI

AI总结 MobileWorldBench通过语义世界模型提升移动智能体任务成功率

Comments 21 pages, 13 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.14130 2025-12-17 cs.CR cs.AI 57%

UIXPOSE: Mobile Malware Detection via Intention-Behaviour Discrepancy Analysis

基于意图-行为不一致分析的移动恶意软件检测:UIXPOSE

Amirmohammad Pasdar, Toby Murray, Van-Thuan Pham

机构 * School of Computing and Information Systems(计算与信息系统学院) The University of Melbourne(墨尔本大学)

专题命中 GUI与屏幕智能体 :vision-language model(abstract);分类 cs.AI

AI总结 UIXPOSE通过意图-行为不一致分析提升移动恶意软件的动态检测能力

Comments 15 pages

详情

展开后加载摘要…

URL PDF HTML 收藏

6. 幻觉与鲁棒性 3 篇

2512.14320 2025-12-17 cs.CV cs.AI cs.CY cs.LG 67%

Semantic Mismatch and Perceptual Degradation: A New Perspective on Image Editing Immunity

语义不匹配与感知退化:图像编辑免疫的新视角

Shuai Dong, Jie Zhang, Guoying Zhao, Shiguang Shan, Xilin Chen

机构 * State Key Laboratory of AI Safety, Institute of Computing Technology, Chinese Academy of Sciences (CAS)(人工智能安全国家重点实验室,计算技术研究所,中国科学院) University of China Academy of Sciences(中国科学院大学) Center for Machine Vision and Signal Analysis, University of Oulu(信号分析中心,奥卢大学) School of Computer Science, China University of Geosciences(计算机科学学院,中国地质大学)

专题命中 幻觉与鲁棒性 :multimodal large language model(abstract);分类 cs.CV、cs.AI、cs.LG

AI总结 本文提出SIFM方法和ISR度量标准,通过语义不匹配和感知退化来评估图像编辑免疫效果,提升对抗恶意扩散式操纵的能力。

Comments 11 pages, 4 figures

详情

展开后加载摘要…

URL PDF HTML 收藏