arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

多模态大模型

跨文本、图像、视频、音频等模态的大模型与学习方法。

2025-07-28 至 2025-07-28 共收录 2 信号源:cs.CV, cs.CL, cs.AI, cs.MM, eess.AS

1. 音频语音多模态 2 篇

2507.19356 2025-07-28 cs.CL cs.SD eess.AS 62%

Enhancing Speech Emotion Recognition Leveraging Aligning Timestamps of ASR Transcripts and Speaker Diarization

Hsuan-Yu Wang, Pei-Ying Lee, Berlin Chen

机构 * Department of English(英语系) National Taiwan Normal University(台湾师范大学) Department of Computer Science and Information Engineering(计算机科学与信息工程系)

专题命中 音频语音多模态 :multimodal(abstract);分类 cs.CL、eess.AS

Comments 6 pages, 3 figures, to appear in the Proceedings of the 2025 International Conference on Asian Language Processing (IALP)

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.18380 2025-07-28 cs.AI cs.LG 57%

RedactOR: An LLM-Powered Framework for Automatic Clinical Data De-Identification

Praphul Singh, Charlotte Dzialo, Jangwon Kim, Sumana Srivatsa, Irfan Bulu, Sri Gadde, Krishnaram Kenthapadi

机构 * Oracle Health & AI(Oracle健康与人工智能)

专题命中 音频语音多模态 :multi-modal(abstract);分类 cs.AI

Comments Accepted to ACL 2025 Industry Track. To appear

详情

展开后加载摘要…

URL PDF HTML 收藏