arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

视觉大模型 / VLM

视觉语言模型、视觉推理、视觉问答、图文理解和视觉 grounding。

2025-08-12 至 2025-08-12 共收录 11 信号源:cs.CV, cs.AI, cs.LG

1. 视觉定位与Grounding 11 篇

2508.07925 2025-08-12 cs.CV 83%

TAG: A Simple Yet Effective Temporal-Aware Approach for Zero-Shot Video Temporal Grounding

Jin-Seop Lee, SungJoon Lee, Jaehan Ahn, YunSeok Choi, Jee-Hyong Lee

机构 * Sungkyunkwan University(顺天大学)

专题命中 视觉定位与Grounding :grounding(title,abstract);vision-language model(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.07877 2025-08-12 cs.CV cs.AI 81%

Selective Contrastive Learning for Weakly Supervised Affordance Grounding

WonJun Moon, Hyun Seok Seong, Jae-Pil Heo

机构 * Sungkyunkwan University(全北大学)

专题命中 视觉定位与Grounding :grounding(title,abstract);分类 cs.CV、cs.AI

Comments Accepted to ICCV 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2411.14594 2025-08-12 cs.CV 79%

Solving Zero-Shot 3D Visual Grounding as Constraint Satisfaction Problems

Qihao Yuan, Kailai Li, Jiaming Zhang

机构 * Bernoulli Institute for Mathematics, Computer Science and Artificial Intelligence, University of Groningen(格罗宁根大学伯努利学院) Computer Vision for Human-Computer Interaction Lab (cv:hci), Karlsruhe Institute of Technology(卡尔斯鲁厄理工学院人机交互计算机视觉实验室)

专题命中 视觉定位与Grounding :grounding(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.07466 2025-08-12 cs.AI 74%

Grounding Natural Language for Multi-agent Decision-Making with Multi-agentic LLMs

Dom Huh, Prasant Mohapatra

机构 * UC Davis(加州大学戴维斯分校) University of South Florida(佛罗里达州立大学)

专题命中 视觉定位与Grounding :grounding(title);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2403.09333 2025-08-12 cs.CV cs.AI 73%

Griffon v2: Advancing Multimodal Perception with High-Resolution Scaling and Visual-Language Co-Referring

Yufei Zhan, Shurong Zheng, Yousong Zhu, Hongyin Zhao, Fan Yang, Ming Tang, Jinqiao Wang

机构 * Foundation Model Research Center, Institute of Automation, Chinese Academy of Sciences(中国科学院自动化研究所基础模型研究中心) School of Artificial Intelligence, University of Chinese Academy of Sciences(中国科学院大学人工智能学院) Peng Cheng Laboratory, Shenzhen, China(鹏城实验室) Wuhan AI Research, Wuhan, China(武汉人工智能研究所)

专题命中 视觉定位与Grounding :vision language model(abstract);grounding(abstract);分类 cs.CV、cs.AI

Comments Accepted by ICCV 2025. Codes and datasets are released at https://github.com/jefferyZhan/Griffon

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.19694 2025-08-12 cs.CV 70%

UltraAD: Fine-Grained Ultrasound Anomaly Classification via Few-Shot CLIP Adaptation

Yue Zhou, Yuan Bi, Wenjuan Tong, Wei Wang, Nassir Navab, Zhongliang Jiang

机构 * Computer Aided Medical Procedures (CAMP)(计算机辅助医疗程序) TU Munich, Germany(慕尼黑工业大学) Munich Center for Machine Learning (MCML)(慕尼黑机器学习中心) The First Affiliated Hospital of Sun Yat-Sen University(中山大学附属第一医院)

专题命中 视觉定位与Grounding :vision-language model(abstract);VLM(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.01663 2025-08-12 cs.CV 70%

Zoom-Refine: Boosting High-Resolution Multimodal Understanding via Localized Zoom and Self-Refinement

Xuan Yu, Dayan Guan, Yanfeng Gu

机构 * Harbin Institute of Technology(哈尔滨工业大学)

专题命中 视觉定位与Grounding :multimodal large language model(abstract);MLLM(abstract);分类 cs.CV

Comments Code is available at https://github.com/xavier-yu114/Zoom-Refine

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.07803 2025-08-12 cs.CV 57%

MambaTrans: Multimodal Fusion Image Translation via Large Language Model Priors for Downstream Visual Tasks

Yushen Xu, Xiaosong Li, Zhenyu Kuang, Xiaoqi Cheng, Haishu Tan, Huafeng Li

机构 * School of Physics and Optoelectronic Engineering(物理与光电工程学院) Guangdong-HongKong-Macao Joint Laboratory for Intelligent Micro-Nano Optoelectronic Technology(粤港澳联合智能微纳光电技术实验室) School of Information Engineering and Automation(信息工程与自动化学院)

专题命中 视觉定位与Grounding :multimodal large language model(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.06539 2025-08-12 cs.LG math.OC 57%

Self-Organizing Survival Manifolds: A Theory for Unsupervised Discovery of Prognostic Structures in Biological Systems

Atahan Karagoz

机构 * Department of Computer Science University of Basel Basel, Switzerland(计算机科学系 巴塞尔大学 巴塞尔瑞士)

专题命中 视觉定位与Grounding :grounding(abstract);分类 cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.08084 2025-08-12 cs.CY 50%

$100,000 or the Robot Gets it! Tech Workers' Resistance Guide: Tech Worker Actions, History, Risks, Impacts, and the Case for a Radical Flank

Mohamed Abdalla

专题命中 视觉定位与Grounding :grounding(abstract)

Comments Accepted to AAAI/ACM AIES 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.07786 2025-08-12 math.LO cs.LO 50%

Proof-theoretic Semantics for Second-order Logic

Alexander V. Gheorghiu, David J. Pym

专题命中 视觉定位与Grounding :grounding(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏