arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

AI 大模型

视觉大模型 / VLM

视觉语言模型、视觉推理、视觉问答、图文理解和视觉 grounding。

2025-09-25 至 2025-09-25 共收录 13 信号源:cs.CV, cs.AI, cs.LG

1. 视觉定位与Grounding 13 篇

2504.04974 2025-09-25 cs.CV cs.AI cs.CL cs.LG 89%

Towards Visual Text Grounding of Multimodal Large Language Model

Ming Li, Ruiyi Zhang, Jian Chen, Chenguang Wang, Jiuxiang Gu, Yufan Zhou, Franck Dernoncourt, Wanrong Zhu, Tianyi Zhou, Tong Sun

机构 * Adobe Research(Adobe研究院) University of Maryland(马里兰大学) University at Buffalo(布法罗大学)

专题命中 视觉定位与Grounding :grounding(title,abstract);multimodal large language model(title,abstract);分类 cs.CV、cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.23745 2025-09-25 cs.CV cs.AI cs.LG 85%

To Trust Or Not To Trust Your Vision-Language Model's Prediction

Hao Dong, Moru Liu, Jian Liang, Eleni Chatzi, Olga Fink

专题命中 视觉定位与Grounding :vision-language model(title,abstract);VLM(abstract);分类 cs.CV、cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.02876 2025-09-25 cs.CV cs.LG 84%

Multimodal Reference Visual Grounding

Yangxiao Lu, Ruosen Li, Liqiang Jing, Jikai Wang, Xinya Du, Yunhui Guo, Nicholas Ruozzi, Yu Xiang

专题命中 视觉定位与Grounding :grounding(title,abstract);vision-language model(abstract);分类 cs.CV、cs.LG

Comments Project page with our code and dataset: https://irvlutd.github.io/MultiGrounding

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.06899 2025-09-25 cs.CL cs.AI 83%

VisualTrap: A Stealthy Backdoor Attack on GUI Agents via Visual Grounding Manipulation

Ziang Ye, Yang Zhang, Wentao Shi, Xiaoyu You, Fuli Feng, Tat-Seng Chua

机构 * University of Science and Technology of China(中国科学技术大学) National University of Singapore(新加坡国立大学) East China University of Science and Technology(东华大学)

专题命中 视觉定位与Grounding :grounding(title,abstract);vision-language model(abstract);分类 cs.AI

Comments Accepted in COLM2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.19090 2025-09-25 cs.CV cs.AI cs.CL 81%

Citrus-V: Advancing Medical Foundation Models with Unified Medical Image Grounding for Clinical Reasoning

Guoxin Wang, Jun Zhao, Xinyi Liu, Yanbo Liu, Xuyang Cao, Chao Li, Zhuoyun Liu, Qintian Sun, Fangru Zhou, Haoqiang Xing, Zhenhong Yang

机构 * JDH Algo(京东健康算法)

专题命中 视觉定位与Grounding :grounding(title,abstract);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.19875 2025-09-25 cs.CV cs.AI 73%

Adaptive Guidance Semantically Enhanced via Multimodal LLM for Edge-Cloud Object Detection

Yunqing Hu, Zheming Yang, Chang Zhao, Wen Ji

机构 * Institute of Computing Technology, Chinese Academy of Sciences(中国科学院计算技术研究所) Institute of AI for Industries(工业人工智能研究所) University of Chinese Academy of Sciences(中国科学院大学)

专题命中 视觉定位与Grounding :multimodal large language model(abstract);MLLM(abstract);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.19843 2025-09-25 cs.CV cs.RO 57%

PersONAL: Towards a Comprehensive Benchmark for Personalized Embodied Agents

Filippo Ziliotto, Jelin Raphael Akkara, Alessandro Daniele, Lamberto Ballan, Luciano Serafini, Tommaso Campari

机构 * University of Padova(帕多瓦大学) Fondazione Bruno Kessler(布鲁诺·科塞拉基金会)

专题命中 视觉定位与Grounding :grounding(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.19326 2025-09-25 cs.CL cs.AI 57%

Unveiling the Merits and Defects of LLMs in Automatic Review Generation for Scientific Papers

Ruochi Li, Haoxuan Zhang, Edward Gehringer, Ting Xiao, Junhua Ding, Haihua Chen

机构 * North Carolina State University(北卡罗来纳州立大学) University of North Texas(德克萨斯大学)

专题命中 视觉定位与Grounding :grounding(abstract);分类 cs.AI

Comments Accepted as short paper at 25th IEEE International Conference on Data Mining

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.19322 2025-09-25 cs.CL cs.AI 57%

Readme_AI: Dynamic Context Construction for Large Language Models

Millie Vyas, Timothy Blattner, Alden Dima

机构 * Purdue University(普渡大学) National Institute of Standards and Technology(国家标准技术研究院)

专题命中 视觉定位与Grounding :grounding(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.16180 2025-09-25 cs.CV cs.CL 57%

Redemption Score: A Multi-Modal Evaluation Framework for Image Captioning via Distributional, Perceptual, and Linguistic Signal Triangulation

Ashim Dahal, Ankit Ghimire, Saydul Akbar Murad, Nick Rahimi

机构 * University of Southern Mississippi(密苏里州南方大学)

专题命中 视觉定位与Grounding :grounding(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2501.05926 2025-09-25 cs.CL 50%

LLMs Reproduce Stereotypes of Sexual and Gender Minorities

Ruby Ostrow, Adam Lopez

机构 * University of Edinburgh(爱丁堡大学)

专题命中 视觉定位与Grounding :grounding(abstract)

Comments 13 pages, 5 figures, 9 tables (including bibliography and appendix). Accepted to Findings of EMNLP 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.19573 2025-09-25 cs.RO 50%

Chasing Stability: Humanoid Running via Control Lyapunov Function Guided Reinforcement Learning

Zachary Olkin, Kejun Li, William D. Compton, Aaron D. Ames

机构 * Technology Innovation Institute (TII)(技术创新研究所)

专题命中 视觉定位与Grounding :grounding(abstract)

Comments Submitted to ICRA 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.16254 2025-09-25 cs.HC 50%

Reassessing Collaborative Writing Theories and Frameworks in the Age of LLMs: What Still Applies and What We Must Leave Behind

Daisuke Yukita, Tim Miller, Joel Mackenzie

专题命中 视觉定位与Grounding :grounding(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏