arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

视觉大模型 / VLM

视觉语言模型、视觉推理、视觉问答、图文理解和视觉 grounding。

2026-01-15 至 2026-01-15 共收录 11 信号源:cs.CV, cs.AI, cs.LG

1. 视觉定位与Grounding 11 篇

2601.08183 2026-01-15 cs.CV cs.AI 84%

GI-Bench: A Panoramic Benchmark Revealing the Knowledge-Experience Dissociation of Multimodal Large Language Models in Gastrointestinal Endoscopy Against Clinical Standards

GI-Bench: 一个全景基准测试,揭示多模态大语言模型在内窥镜检查中与临床标准相比的知识-经验脱节

Yan Zhu, Te Luo, Pei-Yao Fu, Zhen Zhang, Zi-Long Wang, Yi-Fan Qu, Zi-Han Geng, Jia-Qi Xu, Lu Yao, Li-Yun Ma, Wei Su, Wei-Feng Chen, Quan-Lin Li, Shuo Wang, Ping-Hong Zhou

专题命中 视觉定位与Grounding :multimodal large language model(title,abstract);grounding(abstract);分类 cs.CV、cs.AI

AI总结 GI-Bench通过评估多模态大语言模型在胃肠内窥镜中的性能,揭示其在诊断推理和空间定位方面的不足,发现模型在语言流畅度上优于人类,但事实准确性较低。

Comments 45 pages, 17 figures, 6 tables. Leaderboard available at: https://roterdl.github.io/GIBench/ . Includes supplementary material

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.09449 2026-01-15 cs.CV 83%

PrivLEX: Detecting legal concepts in images through Vision-Language Models

PrivLEX:通过视觉-语言模型检测图像中的法律概念

Darya Baranouskaya, Andrea Cavallaro

机构 * EPFL(苏黎世联邦理工学院) Idiap Research Institute(Idiap研究机构)

专题命中 视觉定位与Grounding :vision-language model(title,abstract);VLM(abstract);分类 cs.CV

AI总结 PrivLEX通过视觉-语言模型实现图像中法律概念的检测,无需显式标签即可进行可解释的隐私分类。

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.09212 2026-01-15 cs.CV cs.AI cs.LG 67%

Annealed Relaxation of Speculative Decoding for Faster Autoregressive Image Generation

退火放松的投机解码用于更快的自回归图像生成

Xingyao Li, Fengzhuo Zhang, Cunxiao Du, Hui Ji

专题命中 视觉定位与Grounding :grounding(abstract);分类 cs.CV、cs.AI、cs.LG

AI总结 本文提出COOL-SD,通过退火放松的投机解码方法提升自回归图像生成的速度与质量。

Comments Accepted to AAAI 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.09663 2026-01-15 cs.CV 57%

Self-Supervised Animal Identification for Long Videos

长时间视频中的自监督动物识别

Xuyang Fang, Sion Hannuna, Edwin Simpson, Neill Campbell

机构 * University of Bristol(布里斯托大学)

专题命中 视觉定位与Grounding :vision-language model(abstract);分类 cs.CV

AI总结 本研究提出了一种高效的自监督方法,通过全局聚类任务实现长时间视频中动物个体的高精度识别,准确率超过97%,且内存消耗低,适用于资源受限环境。

Comments 11 pages, 1 figure

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.09661 2026-01-15 cs.CV 57%

LiteEmbed: Adapting CLIP to Rare Classes

LiteEmbed: 适应CLIP的稀有类别

Aishwarya Agarwal, Srikrishna Karanam, Vineet Gandhi

机构 * CVIT, Kohli Centre for Intelligent Systems, IIIT Hyderabad, India(印度海得拉巴IIIT大学计算机视觉研究所、Kohli智能系统中心) Adobe Research, Bengaluru, India(印度班加罗尔Adobe研究)

专题命中 视觉定位与Grounding :vision-language model(abstract);分类 cs.CV

AI总结 LiteEmbed通过子空间优化和双目标方法,实现CLIP在稀有类别上的少样本个性化,提升分类和检索任务的性能。

Comments 14 pages, 12 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.09575 2026-01-15 cs.CV 57%

OpenVoxel: Training-Free Grouping and Captioning Voxels for Open-Vocabulary 3D Scene Understanding

OpenVoxel: 一种无需训练的稀疏体素分组与标注算法用于开放词汇3D场景理解

Sheng-Yu Huang, Jaesung Choe, Yu-Chiang Frank Wang, Cheng Sun

机构 * NVIDIA National Taiwan University(国立台湾大学)

专题命中 视觉定位与Grounding :vision language model(abstract);分类 cs.CV

AI总结 OpenVoxel通过无需训练的多模态大语言模型实现稀疏体素的分组与标注,提升开放词汇3D场景理解的性能。

Comments project page: https://peterjohnsonhuang.github.io/openvoxel-pages/

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.04678 2026-01-15 cs.CV 57%

Tracking and Understanding Object Transformations

跟踪和理解物体变换

Yihong Sun, Xinyu Yang, Jennifer J. Sun, Bharath Hariharan

机构 * Cornell University(康奈尔大学)

专题命中 视觉定位与Grounding :grounding(abstract);分类 cs.CV

AI总结 本文提出TubeletGraph系统,用于跟踪和理解物体在变换中的状态变化,并引入新的基准数据集VOST-TAS,以提升复杂物体变换的跟踪与理解能力。

Comments NeurIPS 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2501.04184 2026-01-15 cs.CV 57%

MedicalNarratives: Connecting Medical Vision and Language with Localized Narratives

MedicalNarratives: 通过局部化叙述连接医学视觉与语言

Wisdom O. Ikezogwo, Kevin Zhang, Mehmet Saygin Seyfioglu, Fatemeh Ghezloo, Linda Shapiro, Ranjay Krishna

机构 * University of Washington(华盛顿大学) Allen Institute for Artificial Intelligence(人工智能研究院) Amazon(亚马逊)

专题命中 视觉定位与Grounding :grounding(abstract);分类 cs.CV

AI总结 MedicalNarratives通过局部化鼠标轨迹连接医学视觉与语言,训练出的GenMedClip在12个医学领域均优于现有最佳模型。

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.09036 2026-01-15 cs.CL cs.IR 50%

SpectraQuery: A Hybrid Retrieval-Augmented Conversational Assistant for Battery Science

SpectraQuery: 一种用于电池科学的混合检索增强型对话助手

Sreya Vangara, Jagjit Nanda, Yan-Kai Tzeng, Eric Darve

机构 * Mechanical Engineering, Stanford University(斯坦福大学机械工程系) Applied Energy Division, SLAC National Accelerator Laboratory(SLAC国家加速器实验室应用能源部门) Institute of Computational and Mathematical Engineering, Stanford University(斯坦福大学计算与数学工程研究所)

专题命中 视觉定位与Grounding :grounding(abstract)

AI总结 SpectraQuery通过结合结构化数据库和文献语料库,为电池科学提供检索增强型对话助手,有效整合数值证据与机理解释,提升科学推理效率。

Comments 11 pages, 8 figures, appendix included

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.08852 2026-01-15 cs.CL 50%

NewsScope: Schema-Grounded Cross-Domain News Claim Extraction with Open Models

NewsScope:基于模式的跨领域新闻声明提取与开放模型

Nidhi Pandya

机构 * Pace University(帕克大学)

专题命中 视觉定位与Grounding :grounding(abstract)

AI总结 NewsScope通过开放模型实现基于模式的跨领域新闻声明提取,展现优于GPT-4o-mini的准确性和泛化能力。

Comments 5 pages, 3 tables. Code, model, and benchmark publicly released

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.06230 2026-01-15 physics.ed-ph 50%

Introducing the Physics of Complex Systems through Videogames

通过视频游戏引入复杂系统的物理

Alessio Focardi, Franco Bagnoli, Andrea Guazzini, Giorgio Gronchi

专题命中 视觉定位与Grounding :grounding(abstract)

AI总结 本研究通过视频游戏教学,帮助高中生理解复杂系统的物理概念,如相变、敏感性和同步,并评估教学效果。

Comments Same as version 1.0, the only changes is Focardi's email address

详情

展开后加载摘要…

URL PDF HTML 收藏