arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

视觉大模型 / VLM

视觉语言模型、视觉推理、视觉问答、图文理解和视觉 grounding。

2025-10-24 至 2025-10-24 共收录 3 信号源:cs.CV, cs.AI, cs.LG

1. 视觉问答 3 篇

2510.20287 2025-10-24 cs.CV cs.AI cs.LG 75%

Breakdance Video classification in the age of Generative AI

Sauptik Dhar, Naveen Ramakrishnan, Michelle Munson

机构 * Eluvio AI Labs(Eluvio AI实验室)

专题命中 视觉问答 :vision language model(abstract);visual question answering(abstract);分类 cs.CV、cs.AI、cs.LG

Comments 11 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.20193 2025-10-24 cs.IR cs.CL cs.CV cs.LG 62%

Multimedia-Aware Question Answering: A Review of Retrieval and Cross-Modal Reasoning Architectures

Rahul Raja, Arpita Vats

机构 * Carnegie Mellon University(卡内基梅隆大学) Boston University(波士顿大学)

专题命中 视觉问答 :grounding(abstract);分类 cs.CV、cs.LG

Comments In Proceedings of the 2nd ACM Workshop in AI-powered Question and Answering Systems (AIQAM '25), October 27-28, 2025, Dublin, Ireland. ACM, New York, NY, USA, 8 pages. https://doi.org/10.1145/3746274.3760393

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.11520 2025-10-24 cs.CV 57%

mmWalk: Towards Multi-modal Multi-view Walking Assistance

Kedi Ying, Ruiping Liu, Chongyan Chen, Mingzhe Tao, Hao Shi, Kailun Yang, Jiaming Zhang, Rainer Stiefelhagen

机构 * CV:HCI, KIT(KIT计算机视觉与人机交互中心) Hunan University(湖南大学) ETH Zurich(苏黎世联邦理工学院) University of Texas at Austin(德克萨斯大学奥斯汀分校) Zhejiang University(浙江大学)

专题命中 视觉问答 :vision-language model(abstract);分类 cs.CV

Comments Accepted by NeurIPS 2025 Datasets and Benchmarks Track. Data and Code: https://github.com/KediYing/mmWalk

详情

展开后加载摘要…

URL PDF HTML 收藏