arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

视觉大模型 / VLM

视觉语言模型、视觉推理、视觉问答、图文理解和视觉 grounding。

共收录 7441 信号源:cs.CV, cs.AI, cs.LG

1. 视觉定位与Grounding 7441 篇

2507.17209 2025-07-24 cs.HC cs.LG 57%

HypoChainer: A Collaborative System Combining LLMs and Knowledge Graphs for Hypothesis-Driven Scientific Discovery

Haoran Jiang, Shaohan Shi, Yunjie Yao, Chang Jiang, Quan Li

专题命中 视觉定位与Grounding :grounding(abstract);分类 cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.16208 2025-07-23 cs.SE cs.AI 57%

LOCOFY Large Design Models -- Design to code conversion solution

Sohaib Muhammad, Ashwati Vipin, Karan Shetti, Honey Mittal

专题命中 视觉定位与Grounding :multimodal large language model(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.21980 2025-07-23 cs.CV 57%

R1-Track: Direct Application of MLLMs to Visual Object Tracking via Reinforcement Learning

Biao Wang, Wenwen Li, Jiawei Ge

机构 * Beihang University(北航大学) Southeast University(东南大学)

专题命中 视觉定位与Grounding :grounding(abstract);分类 cs.CV

Comments 7 pages, 2 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.00721 2025-07-22 cs.CV 57%

UPRE: Zero-Shot Domain Adaptation for Object Detection via Unified Prompt and Representation Enhancement

Xiao Zhang, Fei Wei, Yong Wang, Wenda Zhao, Feiyi Li, Xiangxiang Chu

机构 * Dalian University of Technology(大连理工大学) AMAP, Alibaba Group(阿里集团AMAP)

专题命中 视觉定位与Grounding :vision-language model(abstract);分类 cs.CV

Comments ICCV2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.14705 2025-07-22 cs.AI 57%

Configurable multi-agent framework for scalable and realistic testing of llm-based agents

Sai Wang, Senthilnathan Subramanian, Mudit Sahni, Praneeth Gone, Lingjie Meng, Xiaochen Wang, Nicolas Ferradas Bertoli, Tingxian Cheng, Jun Xu

机构 * eBay Inc.(eBay公司)

专题命中 视觉定位与Grounding :grounding(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.11092 2025-07-22 cs.CL cs.AI cs.HC 57%

Dynamic Context Tuning for Retrieval-Augmented Generation: Enhancing Multi-Turn Planning and Tool Adaptation

Jubin Abhishek Soni, Amit Anand, Rajesh Kumar Pandey, Aniket Abhishek Soni

专题命中 视觉定位与Grounding :grounding(abstract);分类 cs.AI

Comments We are withdrawing the submission in order to thoroughly revise the work

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.01431 2025-07-22 cs.CV 57%

ZS-VCOS: Zero-Shot Video Camouflaged Object Segmentation By Optical Flow and Open Vocabulary Object Detection

Wenqi Guo, Mohamed Shehata, Shan Du

机构 * Department of CMPS, University of British Columbia(计算机科学与软件工程系,不列颠哥伦比亚大学) Group of Methane Emission Observation & Warning (MEOW) , Weathon Software(甲烷排放观测与预警小组,Weathon软件)

专题命中 视觉定位与Grounding :vision-language model(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.13113 2025-07-18 cs.CV 57%

Leveraging Language Prior for Infrared Small Target Detection

Pranav Singh, Pravendra Singh

机构 * Department of Computer Science and Engineering, Indian Institute of Technology Roorkee(计算机科学与工程系,印度理工学院Roorkee)

专题命中 视觉定位与Grounding :vision-language model(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.09092 2025-07-17 cs.CV cs.RO 57%

OpenLKA: An Open Dataset of Lane Keeping Assist from Recent Car Models under Real-world Driving Conditions

Yuhang Wang, Abdulaziz Alhuraish, Shengming Yuan, Hao Zhou

机构 * Department of Civil and Environmental Engineering, University of South Florida(土木与环境工程系,佛罗里达州立大学)

专题命中 视觉定位与Grounding :vision language model(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.13026 2025-07-17 cs.CV 57%

HiMTok: Learning Hierarchical Mask Tokens for Image Segmentation with Large Multimodal Model

Tao Wang, Changxu Cheng, Lingfeng Wang, Senda Chen, Wuyue Zhao

机构 * Uni-Ubi Zhejiang University(浙江大学) Tongji University(同济大学)

专题命中 视觉定位与Grounding :grounding(abstract);分类 cs.CV

Comments Accepted by ICCV 2025; the code is at https://github.com/yayafengzi/LMM-HiMTok

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.13497 2025-07-17 cs.CL cs.AI 57%

Towards Geo-Culturally Grounded LLM Generations

Piyawat Lertvittayakumjorn, David Kinney, Vinodkumar Prabhakaran, Donald Martin, Sunipa Dev

机构 * Google(谷歌) Washington University in St. Louis(华盛顿大学圣路易斯分校)

专题命中 视觉定位与Grounding :grounding(abstract);分类 cs.AI

Comments ACL 2025 (main conference)

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.11003 2025-07-16 cs.CV 57%

Bridge Feature Matching and Cross-Modal Alignment with Mutual-filtering for Zero-shot Anomaly Detection

Yuhu Bai, Jiangning Zhang, Yunkang Cao, Guangyuan Lu, Qingdong He, Xiangtai Li, Guanzhong Tian

机构 * Zhejiang University(浙江大学) YouTu Lab, Tencent(腾讯YouTu实验室) Huazhong University of Science and Technology(华中科技大学) Peking University(北京大学)

专题命中 视觉定位与Grounding :vision-language model(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.10053 2025-07-15 cs.CV 57%

CoSMo: A Multimodal Transformer for Page Stream Segmentation in Comic Books

Marc Serra Ortega, Emanuele Vivoli, Artemis Llabrés, Dimosthenis Karatzas

机构 * Computer Vision Center and Universitat Autònoma de Barcelona(计算机视觉中心和巴塞罗那自治大学) MICC, University of Florence, Italy(佛罗伦萨大学MICC)

专题命中 视觉定位与Grounding :vision-language model(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.05285 2025-07-15 cs.CL cs.AI cs.CY cs.IR 57%

Beyond classical and contemporary models: a transformative AI framework for student dropout prediction in distance learning using RAG, Prompt engineering, and Cross-modal fusion

Miloud Mihoubi, Meriem Zerkouk, Belkacem Chikhaoui

专题命中 视觉定位与Grounding :grounding(abstract);分类 cs.AI

Comments 13 pages, 8 figures, 1 Algorithms, 17th International Conference on Education and New Learning Technologies,: 30 June-2 July, 2025 Location: Palma, Spain

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.09008 2025-07-15 cs.CV 57%

VISTA: A Visual Analytics Framework to Enhance Foundation Model-Generated Data Labels

Xiwei Xuan, Xiaoqi Wang, Wenbin He, Jorge Piazentin Ono, Liang Gou, Kwan-Liu Ma, Liu Ren

机构 * Department of Computer Science, University of California, Davis, CA, USA(加州大学戴维斯分校计算机科学系) Bosch Center for Artificial Intelligence (BCAI), Bosch Research North America(博世人工智能中心(BCAI)、博世北美研究部) Splunk Technology, San Jose, CA, USA(Splunk技术公司)

专题命中 视觉定位与Grounding :LLaVA(abstract);分类 cs.CV

Comments IEEE Transactions on Visualization and Computer Graphics (2025)

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.08443 2025-07-14 cs.LG 57%

KGRAG-Ex: Explainable Retrieval-Augmented Generation with Knowledge Graph-based Perturbations

Georgios Balanos, Evangelos Chasanis, Konstantinos Skianis, Evaggelia Pitoura

机构 * University of Ioannina, Greece(希腊伊奥安妮亚大学) Archimedes, Athena Research Center, Greece(阿提卡研究中心)

专题命中 视觉定位与Grounding :grounding(abstract);分类 cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.06667 2025-07-11 cs.IR cs.AI 57%

Toward Holistic Evaluation of Recommender Systems Powered by Generative Models

Yashar Deldjoo, Nikhil Mehta, Maheswaran Sathiamoorthy, Shuai Zhang, Pablo Castells, Julian McAuley

机构 * Polytechnic University of Bari(巴里理工大学) Duke University(杜克大学) Bespoke Labs(Bespoke实验室) ETH Zurich(苏黎世联邦理工学院) Autónoma University of Madrid(马德里自治大学) University of California, San Diego(加州大学圣地亚哥分校)

专题命中 视觉定位与Grounding :grounding(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.20815 2025-07-09 cs.AI 57%

Dynamic Context-Aware Prompt Recommendation for Domain-Specific AI Applications

Xinye Tang, Haijun Zhai, Chaitanya Belwal, Vineeth Thayanithi, Philip Baumann, Yogesh K Roy

机构 * Microsoft(微软)

专题命中 视觉定位与Grounding :grounding(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.04132 2025-07-08 cs.DL cs.CL cs.CV 57%

An HTR-LLM Workflow for High-Accuracy Transcription and Analysis of Abbreviated Latin Court Hand

Joshua D. Isom

机构 * Lehigh University(莱文大学)

专题命中 视觉定位与Grounding :grounding(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.03756 2025-07-08 stat.ML cs.LG math.ST stat.TH 57%

Implicit Regularisation in Diffusion Models: An Algorithm-Dependent Generalisation Analysis

Tyler Farghly, Patrick Rebeschini, George Deligiannidis, Arnaud Doucet

机构 * Department of Statistics, University of Oxford(统计系,牛津大学) Google(谷歌)

专题命中 视觉定位与Grounding :grounding(abstract);分类 cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.02639 2025-07-04 cs.LG 57%

On Efficient Bayesian Exploration in Model-Based Reinforcement Learning

Alberto Caron, Chris Hicks, Vasilios Mavroudis

机构 * The Alan Turing Institute(艾伦·图灵研究所)

专题命中 视觉定位与Grounding :grounding(abstract);分类 cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.16025 2025-07-04 cs.CV 57%

FeatSharp: Your Vision Model Features, Sharper

Mike Ranzinger, Greg Heinrich, Pavlo Molchanov, Jan Kautz, Bryan Catanzaro, Andrew Tao

机构 * NVIDIA

专题命中 视觉定位与Grounding :vision-language model(abstract);分类 cs.CV

Comments ICML 2025 Version

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.19288 2025-07-02 cs.CV cs.RO 57%

Da Yu: Towards USV-Based Image Captioning for Waterway Surveillance and Scene Understanding

Runwei Guan, Ningwei Ouyang, Tianhao Xu, Shaofeng Liang, Wei Dai, Yafeng Sun, Shang Gao, Songning Lai, Shanliang Yao, Xuming Hu, Ryan Wen Liu, Yutao Yue, Hui Xiong

机构 * Hong Kong University of Science and Technology (Guangzhou)(香港科技大学(广州)) Hong Kong University of Science and Technology, Hong Kong SAR, China(香港科技大学,香港特别行政区,中国) Xi’an Jiaotong-Liverpool University(西安交通大学利物浦大学) China University of Petroleum (East China)(中国石油大学(华东)) University of Science and Technology of China(中国科学技术大学) Yancheng Institute of Technology(盐城职业技术学院) Wuhan University of Technology(武汉理工大学)

专题命中 视觉定位与Grounding :vision-language model(abstract);分类 cs.CV

Comments 14 pages, 13 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.23785 2025-07-01 cs.CV 57%

Visual Textualization for Image Prompted Object Detection

Yongjian Wu, Yang Zhou, Jiya Saiyin, Bingzheng Wei, Yan Xu

机构 * School of Biological Science and Medical Engineering, Beihang University(北京航空航天大学生物科学与医学工程学院) ByteDance Inc.(字节跳动公司)

专题命中 视觉定位与Grounding :vision-language model(abstract);分类 cs.CV

Comments Accepted by ICCV 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.23751 2025-07-01 cs.CV 57%

Can We Challenge Open-Vocabulary Object Detectors with Generated Content in Street Scenes?

Annika Mütze, Sadia Ilyas, Christian Dörpelkus, Matthias Rottmann

机构 * University of Wuppertal(乌珀塔尔大学) Aptiv Services Deutschland GmbH(Aptiv Services 德国分公司)

专题命中 视觉定位与Grounding :grounding(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.23607 2025-07-01 cs.CV 57%

PGOV3D: Open-Vocabulary 3D Semantic Segmentation with Partial-to-Global Curriculum

Shiqi Zhang, Sha Zhang, Jiajun Deng, Yedong Shen, Mingxiao MA, Yanyong Zhang

机构 * University of Science and Technology of China(中国科学技术大学) The University of Adelaide(阿德莱德大学)

专题命中 视觉定位与Grounding :MLLM(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.23263 2025-07-01 cs.CV 57%

Causal-Entity Reflected Egocentric Traffic Accident Video Synthesis

Lei-lei Li, Jianwu Fang, Junbin Xiao, Shanmin Pang, Hongkai Yu, Chen Lv, Jianru Xue, Tat-Seng Chua

机构 * Xi’an Jiaotong University(西安交通大学) National University of Singapore(国立新加坡大学) Nanyang Technological University(南洋理工大学) Cleveland State University(克利夫兰州立大学)

专题命中 视觉定位与Grounding :grounding(abstract);分类 cs.CV

Comments Accepted by ICCV2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.22703 2025-07-01 cs.SE cs.AI 57%

P4OMP: Retrieval-Augmented Prompting for OpenMP Parallelism in Serial Code

Wali Mohammad Abdullah, Azmain Kabir

机构 * Mathematics \& Information Technology Concordia University of Edmonton Edmonton, Alberta, Canada Computer Science University of Manitoba Winnipeg, Manitoba, Canada

专题命中 视觉定位与Grounding :grounding(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.22032 2025-06-30 cs.CV 57%

Partial CLIP is Enough: Chimera-Seg for Zero-shot Semantic Segmentation

Jialei Chen, Xu Zheng, Danda Pani Paudel, Luc Van Gool, Hiroshi Murase, Daisuke Deguchi

机构 * Graduate School of Informatics, Nagoya University(名古屋大学信息学研究科) AI Thrust, The Hong Kong University of Science and Technology(香港科学与技术大学人工智能研究组) INSAIT, Sofia University(索菲亚大学INSAIT)

专题命中 视觉定位与Grounding :vision-language model(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.21860 2025-06-30 cs.RO cs.CV 57%

Embodied Domain Adaptation for Object Detection

Xiangyu Shi, Yanyuan Qiao, Lingqiao Liu, Feras Dayoub

机构 * School of Computer Science and the Australian Institute for Machine Learning at the University of Adelaide(澳大利亚阿德莱德大学计算机科学学院和澳大利亚机器学习研究所)

专题命中 视觉定位与Grounding :vision language model(abstract);分类 cs.CV

Comments Accepted by IROS 2025

详情

展开后加载摘要…

URL PDF HTML 收藏