arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

视觉大模型 / VLM

视觉语言模型、视觉推理、视觉问答、图文理解和视觉 grounding。

2025-12-17 至 2025-12-17 共收录 9 信号源:cs.CV, cs.AI, cs.LG

1. 视觉定位与Grounding 9 篇

2512.12012 2025-12-17 cs.CV cs.AI cs.CL cs.RO 88%

Semantic-Drive: Democratizing Long-Tail Data Curation via Open-Vocabulary Grounding and Neuro-Symbolic VLM Consensus

语义驱动:通过开放词汇锚定和神经符号视觉语言共识民主化长尾数据整理

Antonio Guillen-Perez

机构 * Independent Researcher(独立研究者)

专题命中 视觉定位与Grounding :VLM(title,abstract);grounding(title,abstract);分类 cs.CV、cs.AI

AI总结 Semantic-Drive通过开放词汇锚定和神经符号视觉语言共识,提升自动驾驶中长尾数据整理的效率与隐私保护。

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.13771 2025-12-17 cs.AI 79%

Semantic Grounding Index: Geometric Bounds on Context Engagement in RAG Systems

语义 grounding 指数:RAG 系统中上下文参与的几何界限

Javier Marín

机构 * CERT

专题命中 视觉定位与Grounding :grounding(title,abstract);分类 cs.AI

AI总结 本文提出语义 grounding 指数(SGI),通过几何角度分析 RAG 系统中响应与问题、上下文之间的关系,揭示幻觉响应在角度上接近问题而非上下文,验证了 SGI 在评估响应真实性中的有效性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.14312 2025-12-17 cs.CV cs.AI 73%

From YOLO to VLMs: Advancing Zero-Shot and Few-Shot Detection of Wastewater Treatment Plants Using Satellite Imagery in MENA Region

从YOLO到VLMs:利用卫星图像在中东和北非地区推进零样本和少样本废水处理厂检测

Akila Premarathna, Kanishka Hewageegana, Garcia Andarcia Mariangel

专题命中 视觉定位与Grounding :vision-language model(abstract);VLM(abstract);分类 cs.CV、cs.AI

AI总结 本研究利用VLMs替代YOLOv8,通过零样本和少样本方法高效识别中东和北非地区废水处理厂,提升遥感应用的可扩展性。

Comments 9 pages, 9 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.12492 2025-12-17 cs.CV cs.CL 70%

Adaptive Detector-Verifier Framework for Zero-Shot Polyp Detection in Open-World Settings

面向开放世界设置的零样本息肉检测自适应检测-验证框架

Shengkai Xu, Hsiang Lun Kao, Tianxiang Xu, Honghui Zhang, Junqiao Wang, Runmeng Ding, Guanyu Liu, Tianyu Shi, Zhenyu Yu, Guofeng Pan, Ziqian Bi, Yuqi Ouyang

机构 * College of Computer Science, Sichuan University(四川大学计算机学院) Columbia University(哥伦比亚大学) School of Software and Microelectronics, Peking University(北京大学软件与微电子学院) Apon AI and Brain-Computer Engineering Research Institute(Apon人工智能与脑机工程研究院) Faculty of Science and Technology, University of Macau(澳门大学科学与技术学院) Faculty of Applied Science and Engineering, University of Toronto(多伦多大学应用科学与工程学院) Faculty of Computer Science and Information Technology, University of Malaya(马来亚大学计算机科学与信息技术学院) Zhaolong Technology(智龙科技) Purdue University(普渡大学)

专题命中 视觉定位与Grounding :vision-language model(abstract);VLM(abstract);分类 cs.CV

AI总结 本文提出AdaptiveDetector,通过自适应阈值调整和成本敏感强化学习,实现开放世界中零样本息肉检测,提升召回率并减少假阴性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.14273 2025-12-17 cs.CV 57%

Zoom-Zero: Reinforced Coarse-to-Fine Video Understanding via Temporal Zoom-in

Zoom-Zero: 通过时间放大实现强化的粗到细视频理解

Xiaoqian Shen, Min-Hung Chen, Yu-Chiang Frank Wang, Mohamed Elhoseiny, Ryo Hachiuma

机构 * NVIDIA(英伟达) KAUST(卡塔尔人工智能研究所在线大学)

专题命中 视觉定位与Grounding :grounding(abstract);分类 cs.CV

AI总结 Zoom-Zero通过时间放大和令牌选择性信用分配提升视频问题回答的时空定位精度和答案准确性。

Comments Project page: https://xiaoqian-shen.github.io/Zoom-Zero/

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.02503 2025-12-17 cs.CV 57%

Adapting General-Purpose Foundation Models for X-ray Ptychography in Low-Data Regimes

为低数据情形下的X射线衍射成像适应通用基础模型

Robinson Umeike, Neil Getty, Yin Xiangyu, Yi Jiang

机构 * The University of Alabama(阿拉巴马大学) Argonne National Laboratory(阿贡国家实验室)

专题命中 视觉定位与Grounding :vision-language model(abstract);分类 cs.CV

AI总结 本文提出PtychoBench基准,通过比较SFT和ICL策略,在低数据环境下优化X射线衍射成像任务的模型适应性,发现任务模态决定最佳专门化路径。

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.04687 2025-12-17 cs.CV 57%

Guideline-Consistent Segmentation via Multi-Agent Refinement

通过多智能体细化实现指南一致的分割

Vanshika Vats, Ashwani Rathee, James Davis

专题命中 视觉定位与Grounding :vision-language model(abstract);分类 cs.CV

AI总结 本文提出一种多智能体无训练框架,通过Worker-Supervisor迭代细化架构实现指南一致的分割,有效应对复杂文本指南。

Comments To be published in The Fortieth AAAI Conference on Artificial Intelligence (AAAI 2026)

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.11252 2025-12-17 cs.CV eess.IV 57%

MFGDiffusion: Mask-Guided Smoke Synthesis for Enhanced Forest Fire Detection

MFGDiffusion:基于掩码的烟雾合成以提升森林火灾检测

Guanghao Wu, Yunqing Shang, Chen Xu, Hai Song, Chong Wang, Qixing Zhang

机构 * University of Science and Technology of China(中国科学技术大学)

专题命中 视觉定位与Grounding :multimodal large language model(abstract);分类 cs.CV

AI总结 MFGDiffusion通过掩码引导的烟雾合成提升森林火灾检测性能

Comments 14 pages, 11 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.05310 2025-12-17 cs.CL cs.SI 50%

Listening Between the Lines: Decoding Podcast Narratives with Language Modeling

在字里行间倾听:利用语言模型解码播客叙事

Shreya Gupta, Ojasva Saxena, Arghodeep Nandi, Sarah Masud, Kiran Garimella, Tanmoy Chakraborty

机构 * Indian Institute Of Technology Delhi(印度理工学院德里分校) University of Copenhagen(哥本哈根大学) Rutgers University(罗格斯大学)

专题命中 视觉定位与Grounding :grounding(abstract)

AI总结 本文提出一种基于 BERT 的方法,通过标注叙事框架与对话实体的关系,揭示播客中主题与框架的系统关联,提升对数字媒体影响的分析能力。

Comments 10 pages, 6 Figures, 5 Tables. Under review

详情

展开后加载摘要…

URL PDF HTML 收藏