arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

视觉大模型 / VLM

视觉语言模型、视觉推理、视觉问答、图文理解和视觉 grounding。

2025-08-20 至 2025-08-20 共收录 7 信号源:cs.CV, cs.AI, cs.LG

1. 视觉定位与Grounding 7 篇

2508.13543 2025-08-20 cs.HC 71%

"Can You See Me Think?" Grounding LLM Feedback in Keystrokes and Revision Patterns

Samra Zafar, Shifa Yousaf, Muhammad Shaheer Minhas

专题命中 视觉定位与Grounding :grounding(title)

Comments 15 pages, 4 figures, 6 tables, Submitted to IJCNLP-AACL 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2406.09767 2025-08-20 cs.RO 67%

DISCO: Language-Guided Manipulation with Diffusion Policies and Constrained Inpainting

Ce Hao, Kelvin Lin, Zhiwei Xue, Siyuan Luo, Harold Soh

机构 * School of Computing, National University of Singapore(新加坡国立大学计算机学院) Smart Systems Institute, NUS(NUS智能系统研究所)

专题命中 视觉定位与Grounding :vision-language model(abstract);VLM(abstract)

Journal ref IEEE Robotics and Automation Letters ( Volume: 10, Issue: 10, October 2025)

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.04107 2025-08-20 cs.CV cs.AI 62%

Unlocking the Potential of MLLMs in Referring Expression Segmentation via a Light-weight Mask Decoder

Jingchao Wang, Zhijian Wu, Dingjiang Huang, Yefeng Zheng, Hong Wang

专题命中 视觉定位与Grounding :MLLM(abstract);分类 cs.CV、cs.AI

Comments 9 pages, 4 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.13223 2025-08-20 cs.CV cs.AI 62%

MIRAGE: Towards AI-Generated Image Detection in the Wild

Cheng Xia, Manxi Lin, Jiexiang Tan, Xiaoxiong Du, Yang Qiu, Junjun Zheng, Xiangheng Kong, Yuning Jiang, Bo Zheng

专题命中 视觉定位与Grounding :vision-language model(abstract);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.13901 2025-08-20 cs.RO cs.CV 57%

Multimodal Data Storage and Retrieval for Embodied AI: A Survey

Yihao Lu, Hao Tang

机构 * School of Economics and Management, South China Normal University(经济管理学院,华南师范大学) State Key Laboratory of Multimedia Information Processing, School of Computer Science, Peking University(多媒体信息处理国家重点实验室,计算机学院,北京大学)

专题命中 视觉定位与Grounding :grounding(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.13530 2025-08-20 cs.AI 57%

CrafterDojo: A Suite of Foundation Models for Building Open-Ended Embodied Agents in Crafter

Junyeong Park, Hyeonseo Cho, Sungjin Ahn

机构 * CrafterDojo: A Suite of Foundation Models for Building Open-Ended Embodied Agents in Crafter(CrafterDojo:构建开放性具身智能体的基础模型集合)

专题命中 视觉定位与Grounding :grounding(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2501.04678 2025-08-20 eess.IV cs.CV 57%

RadGPT: Constructing 3D Image-Text Tumor Datasets

Pedro R. A. S. Bassi, Mehmet Can Yavuz, Kang Wang, Xiaoxi Chen, Wenxuan Li, Sergio Decherchi, Andrea Cavalli, Yang Yang, Alan Yuille, Zongwei Zhou

机构 * Johns Hopkins University(约翰霍普金斯大学) University of Bologna(博洛尼亚大学) Italian Institute of Technology(意大利理工学院) University of California, San Francisco(加州大学旧金山分校) Istanbul Medipol University(伊斯坦布尔Medipol大学) University of Zurich(苏黎世大学) ETH AI Center(ETH人工智能中心) University of Illinois Urbana-Champaign(伊利诺伊大学厄巴纳-香槟分校) École Polytechnique Fédérale de Lausanne(洛桑联邦理工学院)

专题命中 视觉定位与Grounding :vision-language model(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏