StoryTeller: Improving Long Video Description through Global Audio-Visual Character Identification
专题命中 视频多模态 :audio-visual(title,abstract);multimodal(abstract);分类 cs.CV、cs.AI
AI 大模型
跨文本、图像、视频、音频等模态的大模型与学习方法。
专题命中 视频多模态 :audio-visual(title,abstract);multimodal(abstract);分类 cs.CV、cs.AI
专题命中 视频多模态 :cross-modal(title,abstract);multimodal(abstract);分类 cs.CL
Comments arXiv admin note: This paper has been withdrawn by arXiv due to disputed and unverifiable authorship and affiliation
机构 * University of Central Florida(佛罗里达中央大学) ; Case Western Reserve University(凯斯西储大学)
专题命中 视频多模态 :multimodal(title,abstract);分类 cs.CL
Comments Accepted to the 24th International Semantic Web Conference Resource Track (ISWC 2025)
机构 * Meta AI
专题命中 视频多模态 :multimodal(abstract)