Breakdance Video classification in the age of Generative AI
机构 * Eluvio AI Labs(Eluvio AI实验室)
专题命中 视觉问答 :vision language model(abstract);visual question answering(abstract);分类 cs.CV、cs.AI、cs.LG
Comments 11 pages
AI 大模型
视觉语言模型、视觉推理、视觉问答、图文理解和视觉 grounding。
机构 * Eluvio AI Labs(Eluvio AI实验室)
专题命中 视觉问答 :vision language model(abstract);visual question answering(abstract);分类 cs.CV、cs.AI、cs.LG
Comments 11 pages
机构 * Carnegie Mellon University(卡内基梅隆大学) ; Boston University(波士顿大学)
专题命中 视觉问答 :grounding(abstract);分类 cs.CV、cs.LG
Comments In Proceedings of the 2nd ACM Workshop in AI-powered Question and Answering Systems (AIQAM '25), October 27-28, 2025, Dublin, Ireland. ACM, New York, NY, USA, 8 pages. https://doi.org/10.1145/3746274.3760393
机构 * CV:HCI, KIT(KIT计算机视觉与人机交互中心) ; Hunan University(湖南大学) ; ETH Zurich(苏黎世联邦理工学院) ; University of Texas at Austin(德克萨斯大学奥斯汀分校) ; Zhejiang University(浙江大学)
专题命中 视觉问答 :vision-language model(abstract);分类 cs.CV
Comments Accepted by NeurIPS 2025 Datasets and Benchmarks Track. Data and Code: https://github.com/KediYing/mmWalk