arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

视觉大模型 / VLM

视觉语言模型、视觉推理、视觉问答、图文理解和视觉 grounding。

2025-11-03 至 2025-11-03 共收录 5 信号源:cs.CV, cs.AI, cs.LG

1. 视觉定位与Grounding 5 篇

2510.27196 2025-11-03 cs.CL cs.AI 83%

MemeArena: Automating Context-Aware Unbiased Evaluation of Harmfulness Understanding for Multimodal Large Language Models

Zixin Chen, Hongzhan Lin, Kaixin Li, Ziyang Luo, Yayue Deng, Jing Ma

机构 * Hong Kong Baptist University(香港 Baptist 大学) Beijing University of Posts and Telecommunications(北京邮电大学) National University of Singapore(新加坡国立大学)

专题命中 视觉定位与Grounding :multimodal large language model(title,abstract);MLLM(abstract);分类 cs.AI

Comments EMNLP 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.27164 2025-11-03 cs.CV cs.AI 73%

Generating Accurate and Detailed Captions for High-Resolution Images

Hankyeol Lee, Gawon Seo, Kyounggyu Lee, Dogun Kim, Kyungwoo Song, Jiyoung Jung

机构 * Department of Artificial Intelligence, University of Seoul(首尔大学人工智能系) Department of Computer Science and Engineering, POSTECH(POSTECH计算机科学与工程系) Department of Applied Statistics, Yonsei University(延世大学应用统计系)

专题命中 视觉定位与Grounding :vision-language model(abstract);VLM(abstract);分类 cs.CV、cs.AI

Comments Work conducted in 2024; released for archival purposes

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.26996 2025-11-03 cs.CV 57%

MoME: Mixture of Visual Language Medical Experts for Medical Imaging Segmentation

Arghavan Rezvani, Xiangyi Yan, Anthony T. Wu, Kun Han, Pooya Khosravi, Xiaohui Xie

机构 * Department of Computer Science, University of California, Irvine(加州大学尔湾分校计算机科学系) School of Medicine, University of California, Irvine(加州大学尔湾分校医学院)

专题命中 视觉定位与Grounding :vision-language model(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.26915 2025-11-03 cs.RO cs.AI 57%

Heterogeneous Robot Collaboration in Unstructured Environments with Grounded Generative Intelligence

Zachary Ravichandran, Fernando Cladera, Ankit Prabhu, Jason Hughes, Varun Murali, Camillo Taylor, George J. Pappas, Vijay Kumar

机构 * Dept. of Electrical and Computer Engineering at Texas A&M University(德克萨斯A&M大学电气与计算机工程系) GRASP Laboratory at the University of Pennsylvania(宾夕法尼亚大学GRASP实验室)

专题命中 视觉定位与Grounding :grounding(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.21657 2025-11-03 cs.CV 57%

FantasyWorld: Geometry-Consistent World Modeling via Unified Video and 3D Prediction

Yixiang Dai, Fan Jiang, Chiyu Wang, Mu Xu, Yonggang Qi

机构 * AMAP, Alibaba Group(阿里集团AMAP) Beijing University of Posts and Telecommunications(北京邮电大学)

专题命中 视觉定位与Grounding :grounding(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏