arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

AI 大模型

多模态大模型

跨文本、图像、视频、音频等模态的大模型与学习方法。

2025-10-03 至 2025-10-03 共收录 7 信号源:cs.CV, cs.CL, cs.AI, cs.MM, eess.AS

1. 音频语音多模态 7 篇

2510.02165 2025-10-03 cs.SE 82%

Towards fairer public transit: Real-time tensor-based multimodal fare evasion and fraud detection

Peter Wauyo, Dalia Bwiza, Alain Murara, Edwin Mugume, Eric Umuhoza

专题命中 音频语音多模态 :multimodal(title,abstract);cross-modal(abstract)

Comments 10 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.01690 2025-10-03 cs.GR cs.HC 82%

Multimodal Feedback for Task Guidance in Augmented Reality

Hu Guo, Lily Patel, Rohan Gupt

专题命中 音频语音多模态 :multimodal(title,abstract);cross-modal(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2405.18213 2025-10-03 cs.SD cs.CV eess.AS 73%

NeRAF: 3D Scene Infused Neural Radiance and Acoustic Fields

Amandine Brunetto, Sascha Hornauer, Fabien Moutarde

机构 * Center for Robotics, Mines Paris - PSL University Paris, France(机器人中心,巴黎 Mines Paris - PSL 大学)

专题命中 音频语音多模态 :cross-modal(abstract);audio-visual(abstract);分类 cs.CV、eess.AS

Comments ICLR 2025 (Poster). Camera ready version. Project Page: https://amandinebtto.github.io/NeRAF; 24 pages, 13 figures

Journal ref The Thirteenth International Conference on Learning Representations, 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.02313 2025-10-03 cs.CV 57%

Clink! Chop! Thud! -- Learning Object Sounds from Real-World Interactions

Mengyu Yang, Yiming Chen, Haozheng Pei, Siddhant Agarwal, Arun Balajee Vasudevan, James Hays

机构 * Georgia Institute of Technology(佐治亚理工学院) Carnegie Mellon University(卡内基梅隆大学)

专题命中 音频语音多模态 :multimodal(abstract);分类 cs.CV

Comments ICCV 2025. Project page: https://clink-chop-thud.github.io/

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.02110 2025-10-03 cs.SD cs.LG eess.AS 57%

SoundReactor: Frame-level Online Video-to-Audio Generation

Koichi Saito, Julian Tanke, Christian Simon, Masato Ishii, Kazuki Shimada, Zachary Novack, Zhi Zhong, Akio Hayakawa, Takashi Shibuya, Yuki Mitsufuji

专题命中 音频语音多模态 :audio-visual(abstract);分类 eess.AS

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.11835 2025-10-03 cs.LG cs.CV 57%

How Can Time Series Analysis Benefit From Multiple Modalities? A Survey and Outlook

Haoxin Liu, Harshavardhan Kamarthi, Zhiyuan Zhao, Shangqing Xu, Shiyu Wang, Qingsong Wen, Tom Hartvigsen, Fei Wang, B. Aditya Prakash

机构 * Georgia Institute of Technology(佐治亚理工学院) Bytedance Inc.(字节跳动公司) Squirrel AI, USA(squirrel AI 美国分公司) The University of Virginia(弗吉尼亚大学) Cornell University(康奈尔大学)

专题命中 音频语音多模态 :multimodal(abstract);分类 cs.CV

Comments Github Repo: https://github.com/AdityaLab/MM4TSA Updated to include papers accepted by IJCAI25, KDD25, ICML25, NeurIPS25 4 figures or tables, 19 pages, 251 references

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.02206 2025-10-03 cs.LG 50%

Poolformer: Recurrent Networks with Pooling for Long-Sequence Modeling

Daniel Gallo Fernández

机构 * MSc Artificial Intelligence Master Thesis(人工智能硕士论文)

专题命中 音频语音多模态 :multi-modal(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏