arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

AI 大模型

多模态大模型

跨文本、图像、视频、音频等模态的大模型与学习方法。

2025-09-08 至 2025-09-08 共收录 5 信号源:cs.CV, cs.CL, cs.AI, cs.MM, eess.AS

1. 音频语音多模态 5 篇

2509.05205 2025-09-08 eess.AS cs.SD 79%

MEAN-RIR: Multi-Modal Environment-Aware Network for Robust Room Impulse Response Estimation

Jiajian Chen, Jiakang Chen, Hang Chen, Qing Wang, Yu Gao, Jun Du

机构 * University of Science and Technology of China(科学技术大学) AI Research Center, Midea Group (Shanghai) Co.,Ltd.(美的集团(上海)有限公司人工智能研究中心)

专题命中 音频语音多模态 :multi-modal(title,abstract);分类 eess.AS

Comments Accepted by ASRU 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.04605 2025-09-08 cs.CL 79%

Spoken in Jest, Detected in Earnest: A Systematic Review of Sarcasm Recognition -- Multimodal Fusion, Challenges, and Future Prospects

Xiyuan Gao, Shekhar Nayak, Matt Coler

机构 * Campus Fryslân, University of Groningen(格罗宁根大学弗里桑校区)

专题命中 音频语音多模态 :multimodal(title,abstract);分类 cs.CL

Comments 20 pages, 7 figures, Submitted to IEEE Transactions on Affective Computing

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.04606 2025-09-08 cs.CL cs.AI cs.CV 75%

Sample-efficient Integration of New Modalities into Large Language Models

Osman Batur İnce, André F. T. Martins, Oisin Mac Aodha, Edoardo M. Ponti

机构 * University of Edinburgh(爱丁堡大学) Instituto de Telecomunicações(电信研究所) Instituto Superior Técnico, Universidade de Lisboa(里斯本大学技术学院) Unbabel

专题命中 音频语音多模态 :multimodal(abstract);multimodal foundation model(abstract);分类 cs.CV、cs.CL、cs.AI

Comments Pre-print

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.22809 2025-09-08 cs.CL cs.AI cs.HC 62%

First Steps Towards Overhearing LLM Agents: A Case Study With Dungeons & Dragons Gameplay

Andrew Zhu, Evan Osgood, Chris Callison-Burch

机构 * University of Pennsylvania(宾夕法尼亚大学)

专题命中 音频语音多模态 :multimodal(abstract);分类 cs.CL、cs.AI

Comments 9 pages, 5 figures. COLM 2025 Workshop on AI Agents

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.04736 2025-09-08 cs.CV 61%

WatchHAR: Real-time On-device Human Activity Recognition System for Smartwatches

Taeyoung Yeon, Vasco Xu, Henry Hoffmann, Karan Ahuja

机构 * Northwestern University(西北大学) University of Chicago(芝加哥大学)

专题命中 音频语音多模态 :multimodal(abstract,comments);分类 cs.CV

Comments 8 pages, 4 figures, ICMI '25 (27th International Conference on Multimodal Interaction), October 13-17, 2025, Canberra, ACT, Australia

详情

展开后加载摘要…

URL PDF HTML 收藏