arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

视觉大模型 / VLM

视觉语言模型、视觉推理、视觉问答、图文理解和视觉 grounding。

2025-08-19 至 2025-08-19 共收录 14 信号源:cs.CV, cs.AI, cs.LG

1. 视觉定位与Grounding 14 篇

2508.05123 2025-08-19 cs.CV cs.AI 81%

Latent Expression Generation for Referring Image Segmentation and Grounding

Seonghoon Yu, Junbeom Hong, Joonseok Lee, Jeany Son

机构 * GIST(韩国科学技术院) Seoul National University(首尔国立大学) POSTECH

专题命中 视觉定位与Grounding :grounding(title,abstract);分类 cs.CV、cs.AI

Comments Accepted to ICCV 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.11903 2025-08-19 cs.CV 79%

OVG-HQ: Online Video Grounding with Hybrid-modal Queries

Runhao Zeng, Jiaqi Mao, Minghao Lai, Minh Hieu Phan, Yanjie Dong, Wei Wang, Qi Chen, Xiping Hu

机构 * Artificial Intelligence Research Institute, Shenzhen MSU-BIT University(人工智能研究院,深圳MSU-BIT大学) University of Adelaide(阿德莱德大学)

专题命中 视觉定位与Grounding :grounding(title,abstract);分类 cs.CV

Comments Accepted to ICCV 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.12605 2025-08-19 cs.CV 70%

ViDA-UGC: Detailed Image Quality Analysis via Visual Distortion Assessment for UGC Images

Wenjie Liao, Jieyu Yuan, Yifang Xu, Chunle Guo, Zilong Zhang, Jihong Li, Jiachen Fu, Haotian Fan, Tao Li, Junhui Cui, Chongyi Li

机构 * Media Evaluation Lab, ByteDance Inc.(字节跳动公司媒体评估实验室)

专题命中 视觉定位与Grounding :grounding(abstract);multimodal large language model(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.12349 2025-08-19 cs.CV 70%

EgoLoc: A Generalizable Solution for Temporal Interaction Localization in Egocentric Videos

Junyi Ma, Erhang Zhang, Yin-Dong Zheng, Yuchen Xie, Yixuan Zhou, Hesheng Wang

机构 * IRMV Lab, the Department of Automation, Shanghai Jiao Tong University(IRMV实验室,自动化系,上海交通大学)

专题命中 视觉定位与Grounding :vision-language model(abstract);grounding(abstract);分类 cs.CV

Comments Extended journal version of arXiv:2506.03662

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.11886 2025-08-19 cs.CV cs.AI cs.CL cs.LG eess.IV 67%

EVTP-IVS: Effective Visual Token Pruning For Unifying Instruction Visual Segmentation In Multi-Modal Large Language Models

Wenhui Zhu, Xiwen Chen, Zhipeng Wang, Shao Tang, Sayan Ghosh, Xuanzhao Dong, Rajat Koner, Yalin Wang

机构 * Arizona State University(亚利桑那州立大学) Clemson University(克莱姆森大学) LinkedIn Corporation(领英公司) Ludwig Maximilian University of Munich(慕尼黑路德维希-马克西米利安大学)

专题命中 视觉定位与Grounding :multimodal large language model(abstract);分类 cs.CV、cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.10935 2025-08-19 cs.CV cs.LG cs.RO 62%

HQ-OV3D: A High Box Quality Open-World 3D Detection Framework based on Diffision Model

Qi Liu, Yabei Li, Hongsong Wang, Lei He

专题命中 视觉定位与Grounding :vision-language model(abstract);分类 cs.CV、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.13043 2025-08-19 cs.CV 57%

IntelliCap: Intelligent Guidance for Consistent View Sampling

Ayaka Yasunaga, Hideo Saito, Dieter Schmalstieg, Shohei Mori

机构 * Keio University(庆应大学) University of Stuttgart(斯图加特大学)

专题命中 视觉定位与Grounding :vision-language model(abstract);分类 cs.CV

Comments This work is a pre-print version of a paper that has been accepted to the IEEE International Symposium on Mixed and Augmented Reality for future publication. Project Page: https://mediated-reality.github.io/projects/yasunaga_ismar25/

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.07402 2025-08-19 cs.CV 57%

ForensicsSAM: Toward Robust and Unified Image Forgery Detection and Localization Resisting to Adversarial Attack

Rongxuan Peng, Shunquan Tan, Chenqi Kong, Anwei Luo, Alex C. Kot, Jiwu Huang

机构 * Shenzhen Key Laboratory of Media Security, Faculty of Electronic and Information Engineering, Shenzhen University, China(深圳媒体安全重点实验室,电子与信息工程学院,深圳大学,中国) Rapid-Rich Object Search (ROSE) Lab, School of Electrical and Electronic Engineering, Nanyang Technological University, Singapore(快速富对象搜索(ROSE)实验室,电气电子工程学院,南洋理工大学,新加坡) Guangdong Laboratory of Machine Perception and Intelligent Computing, Faculty of Engineering, Shenzhen MSU-BIT University, China(广东机器感知与智能计算实验室,工程学院,深圳MSU-BIT大学,中国)

专题命中 视觉定位与Grounding :LLaVA(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.24039 2025-08-19 cs.CV cs.HC 57%

Foundation Models for Zero-Shot Segmentation of Scientific Images without AI-Ready Data

Shubhabrata Mukherjee, Jack Lang, Obeen Kwon, Iryna Zenyuk, Valerie Brogden, Adam Weber, Daniela Ushizima

机构 * Lawrence Berkeley National Laboratory(伯克利国家实验室) University of California, Irvine(加州大学尔湾分校) University of California, Berkeley(加州大学伯克利分校) Covalent Metrology(协力计量)

专题命中 视觉定位与Grounding :visual reasoning(abstract);分类 cs.CV

Comments This paper has been accepted for presentation at the 59th International Conference on Parallel Processing (ICPP 2025), DRAI workshop

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.02356 2025-08-19 cs.CV 57%

InterRVOS: Interaction-aware Referring Video Object Segmentation

Woojeong Jin, Seongchan Kim, Jaeho Lee, Seungryong Kim

专题命中 视觉定位与Grounding :MLLM(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.11873 2025-08-19 cs.CY cs.AI cs.HC cs.MM 57%

SimInterview: Transforming Business Education through Large Language Model-Based Simulated Multilingual Interview Training System

Truong Thanh Hung Nguyen, Tran Diem Quynh Nguyen, Hoang Loc Cao, Thi Cam Thanh Tran, Thi Cam Mai Truong, Hung Cao

机构 * Analytics Everywhere Lab, University of New Brunswick, Canada(新不伦瑞克大学分析 everywhere 实验室) University of Foreign Language Studies, University of Danang, Vietnam(越南丹绒大学外语学院) Faculty of Information Technology, University of Science, VNU-HCM, Vietnam(越南胡志明市大学信息科技学院) Faculty of Economics and Accounting, Quy Nhon University, Vietnam(越南奎隆大学经济与会计学院) Faculty of Natural Sciences, Quy Nhon University, Vietnam(越南奎隆大学自然科学学院)

专题命中 视觉定位与Grounding :grounding(abstract);分类 cs.AI

Comments Published as a conference paper at ICEFM 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.11690 2025-08-19 cs.CY cs.AI 57%

Real Time Child Abduction And Detection System

Tadisetty Sai Yashwanth, Yangalasetty Sruthi Royal, Vankayala Rajeshwari Shreya, Mayank Kashyap, Divyaprabha K N

机构 * Dept. of CSE PES University Bangalore, India(计算机科学与工程系,PES大学,印度班加罗尔)

专题命中 视觉定位与Grounding :vision-language model(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.12388 2025-08-19 cs.HC 50%

When motivation can be more than a message: designing agents to boost physical activity

Alessandro Silacci, Maurizio Caon, Mauro Cherubini

专题命中 视觉定位与Grounding :grounding(abstract)

Comments This is a pre-peer-review version of a paper with the same title accepted at 20th IFIP TC13 International Conference on Human-Computer Interaction (INTERACT 2025)

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.09095 2025-08-19 cond-mat.stat-mech 50%

Benchmarking Energy Calculations Using Formal Proofs

Ejike D. Ugwuanyi, Colin T. Jones, John Velkey, Tyler R. Josephson

专题命中 视觉定位与Grounding :grounding(abstract)

Comments Molecular Physics (2025)

详情

展开后加载摘要…

URL PDF HTML 收藏