arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

视觉大模型 / VLM

视觉语言模型、视觉推理、视觉问答、图文理解和视觉 grounding。

共收录 7409 信号源:cs.CV, cs.AI, cs.LG

1. 视觉定位与Grounding 7409 篇

2408.01942 2024-08-06 cs.AI cs.CV 86%

Visual Grounding for Object-Level Generalization in Reinforcement Learning

Haobin Jiang, Zongqing Lu

专题命中 视觉定位与Grounding :grounding(title,abstract);vision-language model(abstract);VLM(abstract);分类 cs.CV、cs.AI

Comments 35 pages, 14 figures, 17 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2311.17095 2024-06-18 cs.CV cs.AI 86%

Emergent Open-Vocabulary Semantic Segmentation from Off-the-shelf Vision-Language Models

Jiayun Luo, Siddhesh Khandelwal, Leonid Sigal, Boyang Li

专题命中 视觉定位与Grounding :vision-language model(title,abstract);VLM(abstract);visual question answering(abstract);分类 cs.CV、cs.AI

Comments Accepted to CVPR 2024; Earlier version of this paper contained an unintentional error stemming from a bug in the code. This version corrects this error, which had to do with filtering of class names. In consultation with CVPR Program Chairs it was suggested errata be submitted as the updated (fixed) code reinforced original findings (albeit with slightly different final numbers)

详情

展开后加载摘要…

URL PDF HTML 收藏
2404.02726 2024-04-04 cs.CV cs.CR cs.LG 86%

Harnessing the Power of Large Vision Language Models for Synthetic Image Detection

Mamadou Keita, Wassim Hamidouche, Hassen Bougueffa, Abdenour Hadid, Abdelmalik Taleb-Ahmed

专题命中 视觉定位与Grounding :vision language model(title);vision-language model(abstract);VLM(abstract);visual language model(abstract)

Comments arXiv admin note: substantial text overlap with arXiv:2404.01959

详情

展开后加载摘要…

URL PDF HTML 收藏
2403.03170 2024-03-12 cs.MM cs.AI cs.CL cs.CV cs.CY 86%

SNIFFER: Multimodal Large Language Model for Explainable Out-of-Context Misinformation Detection

Peng Qi, Zehong Yan, Wynne Hsu, Mong Li Lee

专题命中 视觉定位与Grounding :multimodal large language model(title,abstract);visual reasoning(abstract);MLLM(abstract);分类 cs.CV、cs.AI

Comments To appear in CVPR 2024

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.16658 2026-06-16 cs.CV 新提交 86%

Vision-Language Models as Zero-Annotation Oracles in Histopathology

视觉-语言模型作为组织病理学中的零标注预言机

Vishal Jain, Giorgio Buzzanca, Sarah Cechnicka, Maarten Naesens, Priyanka Koshy, Tri Nguyen, Jesper Kers, Candice Roufosse, Bernhard Kainz

机构 * Imperial College London(帝国理工学院) Leiden University Medical Center(莱顿大学医学中心) KU Leuven(鲁汶大学) University Hospitals Leuven(鲁汶大学医院) University Medical Center Utrecht(乌得勒支大学医学中心) Friedrich-Alexander University Erlangen-Nürnberg(埃尔朗根-纽伦堡大学)

专题命中 视觉定位与Grounding :vision-language model(title,abstract);VLM(abstract,abstract_cn);分类 cs.CV

AI总结 提出一种粗到细方法,利用通用视觉-语言模型作为零标注预言机进行前景分割,在特殊染色上优于监督基线,并通过伪标签蒸馏轻量学生模型。

Comments 11 pages, 1 figure, 6 tables. Code available at https://github.com/VishalJ99/vlm-wsi-auto-context

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.31227 2026-06-01 cs.CV 86%

HiERO-StepG @ Ego4D Step Grounding Challenge: hierarchical activity understanding enables zero-shot step grounding

HiERO-StepG @ Ego4D Step Grounding Challenge: 层次化活动理解实现零样本步骤定位

Andrea Zenotto, Simone Alberto Peirone, Francesca Pistilli, Giuseppe Averta

机构 * Politecnico di Torino(托里诺理工大学)

专题命中 视觉定位与Grounding :grounding(title,title_cn);分类 cs.CV

AI总结 提出HiERO-StepG方法,利用弱监督层次化表示学习和聚类,无需任务特定微调即可实现零样本步骤定位,在Ego4D挑战中达到56.27% R@1 (IoU=0.3)。

Comments Technical report for the Ego4D Goal Step - Step Grounding challenge at CVPR 2026, derived from arXiv:2505.12911

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.10024 2026-07-20 cs.CV cs.AI cs.LG 版本更新 86%

LVSum: A Benchmark for Timestamp-Aware Long Video Summarization

LVSum:一个用于时间感知长视频摘要的基准测试

Alkesh Patel, Melis Ozyildirim, Ying-Chang Cheng, Ganesh Nagarajan

机构 * Apple(苹果公司)

专题命中 视觉定位与Grounding :MLLM(summary_cn,abstract_cn);grounding(abstract);multimodal large language model(abstract);分类 cs.CV、cs.AI、cs.LG

AI总结 本文提出LVSum基准测试,用于评估长视频摘要中时间对齐的性能,通过引入新的评估指标揭示现有MLLM在时间理解上的系统性差距。

Comments 25 pages, 5 tables, 3 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2607.01079 2026-07-02 cs.RO 新提交 86%

Where Am I? Semantic Map Grounding via Vision-Language Models for Multi-Modal Localization

我在哪里?基于视觉-语言模型的语义地图定位用于多模态定位

Suraj Borate, Aarav Shah, Madhu Vadali

机构 * IIT Gandhinagar(印度理工学院甘地讷格尔分校)

专题命中 视觉定位与Grounding :vision-language model(title,abstract);grounding(title)

AI总结 将机器人定位重构为语义推理任务,使用微调的Qwen2.5-VL-7B模型融合摄像头、LiDAR和语义地图预测位姿,在室内数据集上达到98.23%位置精度,并展现出跨模态互补性和对新场景的泛化能力。

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.11906 2026-06-26 cs.RO 版本更新 86%

Visual-Language-Guided Task Planning for Horticultural Robots

面向园艺机器人的视觉-语言引导任务规划

Jose Cuaran, Kendall Koe, Aditya Potnis, Naveen Kumar Uppalapati, Girish Chowdhary

机构 * the Siebel School of Computing and Data Science(塞比尔计算与数据科学学院) the Department of Agricultural and Biological Engineering(农业与生物工程系) National Center for Supercomputing Applications(国家超级计算中心)

专题命中 视觉定位与Grounding :VLM(summary_cn,abstract);vision language model(abstract);grounding(abstract)

AI总结 提出一种模块化框架,利用视觉语言模型(VLM)主动查询异构数据源来引导机器人任务规划,在单作和混作环境中进行短期和长期作物监测任务基准测试,发现零样本VLM在短期任务中表现稳健(成功率87%),但长期多目标任务成功率低于10%,且依赖噪声语义地图时性能下降。

Comments 18 pages, 6 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.19902 2026-04-23 cs.CV cs.AI cs.LG 86%

MMCORE: MultiModal COnnection with Representation Aligned Latent Embeddings

MMCORE:多模态连接与表征对齐的潜在嵌入

Zijie Li, Yichun Shi, Jingxiang Sun, Ye Wang, Yixuan Huang, Zhiyao Guo, Xiaochen Lian, Peihao Zhu, Yu Tian, Zhonghua Zhai, Peng Wang

机构 * ByteDance Seed(字节跳动种子)

专题命中 视觉定位与Grounding :VLM(abstract,abstract_cn);grounding(abstract,abstract_cn);vision-language model(abstract);分类 cs.CV、cs.AI、cs.LG

AI总结 MMCORE通过预训练视觉-语言模型生成语义视觉嵌入,结合扩散模型实现多模态图像生成与编辑,提升生成质量并降低计算开销。

详情

展开后加载摘要…

URL PDF HTML 收藏
2407.10385 2024-10-01 cs.CL cs.AI cs.CV cs.LG 86%

By My Eyes: Grounding Multimodal Large Language Models with Sensor Data via Visual Prompting

Hyungjun Yoon, Biniyam Aschalew Tolera, Taesik Gong, Kimin Lee, Sung-Ju Lee

专题命中 视觉定位与Grounding :grounding(title);multimodal large language model(title);分类 cs.CV、cs.AI、cs.LG

Comments Accepted to the 2024 Conference on Empirical Methods in Natural Language Processing (EMNLP 2024) Main

详情

展开后加载摘要…

URL PDF HTML 收藏
2304.08587 2023-06-21 cs.RO 86%

Grounding Classical Task Planners via Vision-Language Models

Xiaohan Zhang, Yan Ding, Saeid Amiri, Hao Yang, Andy Kaminski, Chad Esselink, Shiqi Zhang

专题命中 视觉定位与Grounding :vision-language model(title,abstract);grounding(title)

Comments ICRA Workshop on Robot Execution Failures and Failure Management Strategies, 2023

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.11201 2026-08-12 cs.CV 新提交 85%

VidForensics-M1: Meta-Detection Reinforcement Learning with Verifiable Temporal Grounding for AI-Generated Video Forensics

VidForensics-M1:用于AI生成视频取证的、具备可验证时间定位的元检测强化学习

Bowei Liu, Zheng Lu, Yuhan Bian, Xinchen Zhang, Xingming Shui, Yuesheng Huang, Xuhuan Li, Zihao Liu, Yifan Yang, Jun Zhou, Xiu Li

机构 * Tsinghua University(清华大学) Peking University(北京大学) Renmin University of China(中国人民大学) Microsoft(微软公司)

专题命中 视觉定位与Grounding :grounding(title,abstract);MLLM(abstract,abstract_cn);分类 cs.CV

AI总结 该研究针对AI生成视频检测泛化性不足的问题,首次将元检测引入该领域,提出结合可验证时间证据的VidForensics-M1模型及相关机制,实现了鲁棒可泛化的检测。

Comments 27 pages, 15 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.06869 2026-08-12 cs.CV cs.CL 版本更新 85%

DAEP: Difficulty-Aware Evidence Planning for Medical Video Corpus Temporal Answer Grounding

DAEP:面向医学视频语料时序答案 grounding 的难度感知证据规划

Tianjian He, Yujie Liu, Zhiping Huang, Changbo Xu

机构 * TikTok, ByteDance(字节跳动TikTok) Beijing Institute of Graphic Communication(北京印刷学院) Lingnan University(岭南大学)

专题命中 视觉定位与Grounding :grounding(title,title_cn);分类 cs.CV

AI总结 BIGC团队提出DAEP方案,通过多模态证据排名与难度感知规划,在NLPCC 2026共享任务中以0.2728的平均得分获第一名,提升了医学视频时序答案定位的性能。

Comments 12 pages, 2 figures, 5 tables, accepted by NLPCC 2026 Shared Task Track 3

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.08867 2026-08-11 cs.CV 新提交 85%

Zero-Shot Traffic Accident Detection via a Coarse-to-Fine VLM-Tracking Pipeline

基于粗到细VLM跟踪流水线的零样本交通事故检测

Dipit Saha, Shah Mohammad Abdul Mannan, Mohammad Raihan Rashid, Ruwad Naswan, Ahnaf Tahmid

机构 * Bangladesh University of Engineering and Technology(孟加拉工程技术大学)

专题命中 视觉定位与Grounding :VLM(title,title_cn);vision-language model(abstract);分类 cs.CV

AI总结 本文提出一种无需训练的双通粗到细流水线,结合Qwen3-VL-32B-Instruct、YOLO11x与BoT-SORT,在零样本约束下于ACCIDENT @ CVPR基准测试中实现22%相对优势,达成0.504的三方调和均值得分。

Comments Accepted at the AUTOPILOT Workshop, CVPR 2026, Denver, CO

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.01008 2026-08-04 cs.AI 新提交 85%

Toward Fine-Grained Forgetting:Attribute Unlearning for Multimodal Large Language Models

面向细粒度遗忘:多模态大语言模型的属性遗忘

Junkai Lin, Junkai Chen, Siqi Hou, Yuhao He, Ruiqi Liu, Chenhan Jin, Shengze Xu, Tieyong Zeng

专题命中 视觉定位与Grounding :multimodal large language model(title,abstract);MLLM(abstract);分类 cs.AI

AI总结 针对多模态大语言模型(MLLMs)属性遗忘的细粒度需求,提出轻量级训练无关框架CLRP,通过激活修补定位因果层并应用保留感知投影,实现目标属性遗忘同时保留同一身份信息,在多种MLLMs上验证了有效性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.27176 2026-07-29 cs.CV 版本更新 85%

MEDIC-AD: Towards Medical Vision-Language Model's Clinical Intelligence

MEDIC-AD:迈向医疗视觉-语言模型的临床智能

Woohyeon Park, Jaeik Kim, Sunghwan Steve Cho, Pa Hong, Wookyoung Jeong, Yoojin Nam, Namjoon Kim, Ginny Y. Wong, Ka Chun Cheung, Jaeyoung Do

机构 * AIDAS Laboratory, Seoul National University(首尔大学AIDAS实验室) Samsung Changwon Hospital(三星昌原医院) Samsung Medical Center(三星医疗中心) NVIDIA, Santa Clara, USA(英伟达(美国圣克拉拉))

专题命中 视觉定位与Grounding :vision-language model(title,abstract);VLM(abstract,abstract_cn);分类 cs.CV

AI总结 MEDIC-AD通过分阶段框架提升医疗视觉-语言模型在病变检测、症状跟踪和视觉可解释性方面的性能,实现临床应用中的最优表现。

Journal ref CVPR 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2607.23962 2026-07-28 cs.CV cs.CR 新提交 85%

Development of Vision-Language Model-based GNSS Spoofing Detection for Autonomous Vehicle Navigation

基于视觉语言模型的自动驾驶车辆导航GNSS欺骗检测技术发展

Mohammed Aldeen, Muhammad Sami Irfan, Sagar Dasgupta, Long Cheng, Mizanur Rahman, Mashrur Chowdhury

专题命中 视觉定位与Grounding :vision-language model(title,abstract);VLM(abstract,abstract_cn);分类 cs.CV

AI总结 研究自动驾驶车辆GNSS欺骗检测问题,提出基于视觉语言模型的框架,融合多源数据,经三阶段微调及自适应推理策略,大幅提升检测准确率与效率,并生成数据集验证跨区域泛化能力,为自动驾驶提供实用道路防御层。

详情

展开后加载摘要…

URL PDF HTML 收藏
2607.23951 2026-07-28 cs.CV 新提交 85%

TimePLE: Rethinking Temporal Representation for Video Temporal Grounding

TimePLE:重新思考视频时间定位的时间表示

Yuhui Zeng, Xinyu Mao, Xiaokun Liu, Xin Tao, Jinfa Huang, Jiayi Ji, Xiawu Zheng

专题命中 视觉定位与Grounding :grounding(title,abstract);VLM(abstract,abstract_cn);分类 cs.CV

AI总结 研究视频时间定位问题,提出TimePLE方法,通过预测有效时间区间的联合分布将VTG从端点预测转为区间原生定位,经实验验证该方法优于端点预测基线,在短和中等持续时间事件上表现出色。

Comments 25 pages, 13 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2607.23132 2026-07-28 cs.CV 新提交 85%

DispatchRAG: Grounding Emergency Dispatch Decisions in Real-World Protocols from Traffic Accident Video

DispatchRAG:基于交通事故视频中的现实世界协议进行应急调度决策

Muhammad Sulthan Adhipradhana, Ehsan Javanmardi, Naren Bao, Manabu Tsukada

机构 * Graduate School of Information Science and Technology, University of Tokyo(东京大学信息科学与技术研究生院)

专题命中 视觉定位与Grounding :grounding(title);VLM(abstract,abstract_cn);vision-language model(abstract);分类 cs.CV

AI总结 研究旨在通过DispatchRAG框架,利用基于RAG的检索机制和大型语言模型驱动的推理器,根据日本现实交通事故响应协议进行事故评估和调度,引入事故调度数据集验证框架,为自动驾驶车辆事故报告提供支持。

详情

展开后加载摘要…

URL PDF HTML 收藏
2607.21065 2026-07-24 cs.CV 新提交 85%

Do Pathology Vision-Language Models Truly See Pathology?

病理学视觉语言模型真的能‘看到’病理学吗?

Chengyang Zhang, Wenchuan Zhang, Bo Li, Xinyu Liu, Jiaming Yang, Mengran Li, Chenxun Deng, Jie Chen, Yang Zhang, Wei Ju, Yuhao Yi, Hong Bu, Jiancheng Lv

专题命中 视觉定位与Grounding :vision-language model(title,abstract);grounding(abstract,abstract_cn);分类 cs.CV

AI总结 研究病理学视觉语言模型评估中被忽视的问题,提出含多类型样本的PathBind基准,通过对多个VLMs评估发现,当前病理学VLMs在答案性能与视觉语义绑定间有显著差距。

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.21751 2026-07-17 cs.LG 版本更新 85%

Models Can Model, But Can't Bind: Structured Grounding in Text-to-Optimization

模型可以建模,但无法绑定:文本到优化中的结构化 grounding

Zhiqi Gao, Albert Ge, Alexander Berenbeim, Nathaniel D. Bastian, Frederic Sala

机构 * University of Wisconsin-Madison(威斯康星大学麦迪逊分校) United States Military Academy(美国军事学院)

专题命中 视觉定位与Grounding :grounding(title,title_cn);分类 cs.LG

AI总结 本文研究了文本到优化任务中建模与绑定两个关键能力的分离性,发现随着实例数据增长,模型准确性下降,提出BIND方法通过结构化文件外部化数据来提升绑定性能,验证了绑定专精模型在不同优化类别中的优势。

Comments Accepted to COLM 2026. Code and data: https://github.com/SprocketLab/Text2Opt-Bench

详情

展开后加载摘要…

URL PDF HTML 收藏
2607.11598 2026-07-14 cs.AI 新提交 85%

Interaction Scaling: Grounding the Third Axis of Test-Time Compute

交互缩放:奠定测试时计算的第三轴基础

Bojie Li, Noah Shi

机构 * Pine AI(松树人工智能公司) University of Washington(华盛顿大学)

专题命中 视觉定位与Grounding :grounding(title,abstract);vision-language model(abstract);VLM(abstract);分类 cs.AI

AI总结 研究测试时增加计算量的新方法,提出交互缩放概念,通过模型与外部仪器交互突破传统方法局限,在硬编码任务和视觉工件处理上展现优势,表明交互缩放真实且区别于推理和采样,需反馈和度量基于实际。

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.28862 2026-07-07 cs.CV 版本更新 85%

HKVLM: Faithful Query--Region Binding for Frozen-Detector Visual Grounding

HKVLM:通过将语言查询绑定到冻结检测器实现忠实推理定位

Bo Ma

机构 * Bo Ma

专题命中 视觉定位与Grounding :grounding(title,abstract);VLM(abstract_cn);分类 cs.CV

AI总结 提出HKVLM框架,通过将定位从语言路径中分离,利用冻结检测器和语言模型,仅训练轻量对齐钩子实现推理定位,在少样本冷启动场景下显著提升定位精度并减少幻觉。

详情

展开后加载摘要…

URL PDF HTML 收藏
2501.16947 2026-06-29 cs.CV cs.RO 版本更新 85%

Image-based Geo-localization for Robotics: Are Black-box Vision-Language Models there yet?

基于图像的机器人地理定位:黑盒视觉语言模型是否已经足够?

Sania Waheed, Bruno Ferrarini, Michael Milford, Sarvapali D. Ramchurn, Shoaib Ehsan

机构 * University of Southampton(南安普顿大学) MyWay srl Queensland University of Technology(昆士兰科技大学) University of Essex(埃塞克斯大学)

专题命中 视觉定位与Grounding :vision-language model(title,abstract);VLM(abstract,abstract_cn);分类 cs.CV

AI总结 本文首次系统研究黑盒生成式视觉语言模型作为独立零样本地理定位系统的潜力,发现其在粗粒度定位上表现良好,但在细粒度定位上因现实变化而显著退化。

Comments Accepted to the ICRA 2026 Workshop on Multi-Modal Spatial AI for Robust Navigation and Open-World Understanding (MM-SpatialAI)

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.25160 2026-06-25 cs.RO cs.CV 新提交 85%

Toward Low-Latency Vision-Language Models with Doubly-Correct Predictions in Egocentric Visual Understanding

面向低延迟视觉语言模型:在自我中心视觉理解中实现双重正确预测

Qitong Wang, Fan Du, Pranav Maneriker, Jihui Jin, Christopher Rasmussen

机构 * Dolby Laboratories, Inc.(杜比实验室公司) University of Delaware(特拉华大学)

专题命中 视觉定位与Grounding :vision-language model(title,abstract);VLM(abstract,abstract_cn);分类 cs.CV

AI总结 针对人机协作中低延迟需求,提出基于双重正确预测的剪枝策略,在保持证据定位的同时提升预测准确性,在自我中心视频数据集上实现最高精度与双重正确性。

Comments International Conference on Intelligent Robots and Systems (IROS) 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.01482 2026-06-24 cs.AI 版本更新 85%

Grounding Multi-Hop Reasoning in Structural Causal Models via Group Relative Policy Optimization

通过群体相对策略优化在结构因果模型中 grounding 多跳推理

Yunhan Bu, Quan Zhang, Huaping Zhang, Guotong Geng, Chunxiao Gao, Askar Hamdulla, Juan Wang, Qiuchi Li, Baohua Zhang, Shuai Lei, Yunbo Cao, Zhunchen Luo

机构 * School of Computer Science and Technology, Xinjiang University, Urumqi, China(新疆大学计算机科学与技术学院,乌鲁木齐,中国) Beijing Institute of Technology, Beijing, China(北京理工大学,北京,中国) Military Science Information Research Center, Academy of Military Science, Beijing, China(军事科学院军事科学信息研究中心,北京,中国)

专题命中 视觉定位与Grounding :grounding(title,title_cn);分类 cs.AI

AI总结 本文提出基于结构因果模型的多跳事实验证框架,通过群体相对策略优化解决推理链长度与准确率的平衡问题,实验表明其在HoVer和EX-FEVER数据集上优于现有方法。

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.19053 2026-06-18 cs.CV 新提交 85%

Benchmarking Large Vision-Language Models on Fine-Grained Image Tasks: From Evaluation to Diagnosis

大规模视觉-语言模型在细粒度图像任务上的基准测试:从评估到诊断

Hong-Tao Yu, Chen-Wei Xie, Yuxin Peng, Serge Belongie, Xiu-Shen Wei

机构 * School of Computer Science and Engineering, Southeast University, China(东南大学计算机科学与工程学院,中国) Alibaba Group(阿里巴巴集团) School of Computer Science and Engineering, School of Intelligence Science and Engineering, and Key Laboratory of New Generation Artificial Intelligence Technology and Its Interdisciplinary Applications, Southeast University, China(东南大学计算机科学与工程学院、智能科学与工程学院以及新一代人工智能技术及其交叉应用关键实验室,中国) Wangxuan Institute of Computer Technology, National Key Laboratory for Multimedia Information Processing, Peking University, China(北京大学王轩计算机技术研究所、多媒体信息处理国家重点实验室,中国) University of Copenhagen, Denmark(丹麦哥本哈根大学)

专题命中 视觉定位与Grounding :vision-language model(title,abstract);VLM(abstract_cn);grounding(abstract);分类 cs.CV

AI总结 提出FG-BMK基准,含101万问题和28万图像,通过人机双范式评估LVLM的细粒度语义识别与视觉判别能力,诊断失败原因,发现视觉表示、语义对齐等瓶颈。

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.15320 2026-06-16 cs.CV 新提交 85%

Conditional Multi-Event Temporal Grounding in Long-Form Video

长视频中的条件多事件时间定位

Yuanhao Zou, Arthad Kulkarni, Lucas Tonanez, Lincoln Spencer, Guangyu Sun, Tianxingjian Ding, Andong Deng, Yi Li, Shuangjun Liu, Yuan Li, Dashan Gao, Ning Bi, Taotao Jing, Shuai Zhang, Chen Chen

机构 * University of Central Florida(中佛罗里达大学) Qualcomm AI Research(高通人工智能研究院)

专题命中 视觉定位与Grounding :grounding(title,abstract);multimodal large language model(abstract);MLLM(abstract_cn);分类 cs.CV

AI总结 提出CoMET-Bench基准和CoMET-Agent框架,解决长视频中基于组合时空条件定位所有事件的任务,F1@0.5提升6.1%。

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.09132 2026-06-09 cs.AI 新提交 85%

Vision Language Model Helps Private Information De-Identification in Vision Data

视觉语言模型助力视觉数据中的隐私信息去标识化

Tiejin Chen, Pingzhi Li, Kaixiong Zhou, Tianlong Chen, Hua Wei

机构 * Arizona State University(亚利桑那州立大学) University of North Carolina at Chapel Hill(北卡罗来纳大学教堂山分校) North Carolina State University(北卡罗来纳州立大学)

专题命中 视觉定位与Grounding :vision language model(title);vision-language model(abstract);VLM(abstract_cn);visual language model(abstract)

AI总结 提出VisShield框架,通过专用指令微调数据集OPTIC和训练策略,使视觉语言模型精准定位并掩码敏感文本,有效保护医学图像等视觉数据中的隐私信息。

详情

展开后加载摘要…

URL PDF HTML 收藏