arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

AI 大模型

多模态大模型

跨文本、图像、视频、音频等模态的大模型与学习方法。

2025-09-18 至 2025-09-18 共收录 4 信号源:cs.CV, cs.CL, cs.AI, cs.MM, eess.AS

1. 视频多模态 4 篇

2509.13515 2025-09-18 cs.CV 79%

Multimodal Hate Detection Using Dual-Stream Graph Neural Networks

Jiangbei Yue, Shuonan Yang, Tailin Chen, Jianbo Jiao, Zeyu Fu

机构 * Multimodal Intelligence Lab, Department of Computer Science University of Exeter Exeter, UK(埃克塞特大学计算机科学系多模态智能实验室) School of Computer Science University of Birmingham Birmingham, UK(伯明翰大学计算机科学学院)

专题命中 视频多模态 :multimodal(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2408.15864 2025-09-18 cs.RO 78%

FlowAct: A Proactive Multimodal Human-robot Interaction System with Continuous Flow of Perception and Modular Action Sub-systems

Timothée Dhaussy, Bassam Jabaian, Fabrice Lefèvre

机构 * Laboratoire Informatique d'Avignon, Avignon University, France(阿维尼翁信息实验室,阿维尼翁大学,法国)

专题命中 视频多模态 :multimodal(title,abstract)

Comments Paper accepted at ICPRAM 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2408.11915 2025-09-18 cs.SD cs.CV cs.LG cs.MM eess.AS 67%

Video-Foley: Two-Stage Video-To-Sound Generation via Temporal Event Condition For Foley Sound

Junwon Lee, Jaekwon Im, Dabin Kim, Juhan Nam

机构 * Graduate School of AI, KAIST(韩国国立庆熙大学人工智能研究生院) Graduate School of CT, KAIST(韩国国立庆熙大学CT研究生院)

专题命中 视频多模态 :audio-visual(abstract);分类 cs.CV、cs.MM、eess.AS

Comments Accepted at IEEE/ACM Transactions on Audio, Speech and Language Processing (TASLP)

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.13722 2025-09-18 cs.CV cs.AI 62%

Mitigating Query Selection Bias in Referring Video Object Segmentation

Dingwei Zhang, Dong Zhang, Jinhui Tang

机构 * Nanjing University of Science and Technology(南京理工大学) The Hong Kong University of Science and Technology(香港科学大学) Nanjing Forestry University(南京林业大学)

专题命中 视频多模态 :cross-modal(abstract);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏