arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

多模态大模型

跨文本、图像、视频、音频等模态的大模型与学习方法。

共收录 46430 信号源:cs.CV, cs.CL, cs.AI, cs.MM, eess.AS

1. 音频语音多模态 4597 篇

2406.10056 2024-06-17 cs.SD eess.AS 70%

UniAudio 1.5: Large Language Model-driven Audio Codec is A Few-shot Audio Task Learner

Dongchao Yang, Haohan Guo, Yuanyuan Wang, Rongjie Huang, Xiang Li, Xu Tan, Xixin Wu, Helen Meng

专题命中 音频语音多模态 :multi-modal(abstract);cross-modal(abstract);分类 eess.AS

详情

展开后加载摘要…

URL PDF HTML 收藏
2401.06395 2024-04-05 cs.CV 70%

ModaVerse: Efficiently Transforming Modalities with LLMs

Xinyu Wang, Bohan Zhuang, Qi Wu

专题命中 音频语音多模态 :multi-modal(abstract);MLLM(abstract);分类 cs.CV

Comments CVPR2024

详情

展开后加载摘要…

URL PDF HTML 收藏
2309.08730 2024-04-03 eess.AS cs.AI cs.CL cs.MM cs.SD 70%

MusiLingo: Bridging Music and Text with Pre-trained Language Models for Music Captioning and Query Response

Zihao Deng, Yinghao Ma, Yudong Liu, Rongchen Guo, Ge Zhang, Wenhu Chen, Wenhao Huang, Emmanouil Benetos

专题命中 音频语音多模态 :multimodal(abstract);分类 cs.CL、cs.AI、cs.MM

Journal ref 2024 Annual Conference of the North American Chapter of the Association for Computational Linguistics

详情

展开后加载摘要…

URL PDF HTML 收藏
2403.07938 2024-03-14 cs.SD cs.AI cs.CV cs.LG cs.MM eess.AS 70%

Text-to-Audio Generation Synchronized with Videos

Shentong Mo, Jing Shi, Yapeng Tian

专题命中 音频语音多模态 :audio-visual(abstract);分类 cs.CV、cs.AI、cs.MM

Comments arXiv admin note: text overlap with arXiv:2305.12903

详情

展开后加载摘要…

URL PDF HTML 收藏
2401.03697 2024-03-08 cs.SD eess.AS 70%

An audio-quality-based multi-strategy approach for target speaker extraction in the MISP 2023 Challenge

Runduo Han, Xiaopeng Yan, Weiming Xu, Pengcheng Guo, Jiayao Sun, He Wang, Quan Lu, Ning Jiang, Lei Xie

专题命中 音频语音多模态 :multi-modal(abstract);audio-visual(abstract);分类 eess.AS

Comments Accepted by ICASSP 2024

详情

展开后加载摘要…

URL PDF HTML 收藏
2306.02972 2024-03-08 eess.AS cs.LG 70%

Simultaneous or Sequential Training? How Speech Representations Cooperate in a Multi-Task Self-Supervised Learning System

Khazar Khorrami, María Andrea Cruz Blandón, Tuomas Virtanen, Okko Räsänen

专题命中 音频语音多模态 :cross-modal(abstract);audio-visual(abstract);分类 eess.AS

Comments 5 pages, accepted by EUSIPCO 2023

详情

展开后加载摘要…

URL PDF HTML 收藏
2403.01226 2024-03-05 cs.CV 70%

DiffSal: Joint Audio and Video Learning for Diffusion Saliency Prediction

Junwen Xiong, Peng Zhang, Tao You, Chuanyue Li, Wei Huang, Yufei Zha

专题命中 音频语音多模态 :multi-modal(abstract);audio-visual(abstract);分类 cs.CV

Comments 15 pages, CVPR24

详情

展开后加载摘要…

URL PDF HTML 收藏
2402.16153 2024-02-27 cs.SD cs.AI cs.CL cs.LG cs.MM eess.AS 70%

ChatMusician: Understanding and Generating Music Intrinsically with LLM

Ruibin Yuan, Hanfeng Lin, Yi Wang, Zeyue Tian, Shangda Wu, Tianhao Shen, Ge Zhang, Yuhang Wu, Cong Liu, Ziya Zhou, Ziyang Ma, Liumeng Xue, Ziyu Wang, Qin Liu, Tianyu Zheng, Yizhi Li, Yinghao Ma, Yiming Liang, Xiaowei Chi, Ruibo Liu, Zili Wang, Pengfei Li, Jingcheng Wu, Chenghua Lin, Qifeng Liu, Tao Jiang, Wenhao Huang, Wenhu Chen, Emmanouil Benetos, Jie Fu, Gus Xia, Roger Dannenberg, Wei Xue, Shiyin Kang, Yike Guo

专题命中 音频语音多模态 :multi-modal(abstract);分类 cs.CL、cs.AI、cs.MM

Comments GitHub: https://shanghaicannon.github.io/ChatMusician/

详情

展开后加载摘要…

URL PDF HTML 收藏
2310.05181 2024-01-11 eess.AS cs.GR cs.HC cs.LG cs.SD 70%

Unified speech and gesture synthesis using flow matching

Shivam Mehta, Ruibo Tu, Simon Alexanderson, Jonas Beskow, Éva Székely, Gustav Eje Henter

专题命中 音频语音多模态 :multimodal(abstract);cross-modal(abstract);分类 eess.AS

Comments 5 pages, 1 figure. Final version, accepted to IEEE ICASSP 2024

详情

展开后加载摘要…

URL PDF HTML 收藏
2312.09034 2023-12-15 eess.AS cs.SD eess.IV 70%

Fusion of Audio and Visual Embeddings for Sound Event Localization and Detection

Davide Berghi, Peipei Wu, Jinzheng Zhao, Wenwu Wang, Philip J. B. Jackson

专题命中 音频语音多模态 :cross-modal(abstract);audio-visual(abstract);分类 eess.AS

Comments ICASSP 2024 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP)

详情

展开后加载摘要…

URL PDF HTML 收藏
2308.09300 2023-12-15 cs.CV cs.AI cs.MM cs.SD eess.AS 70%

V2A-Mapper: A Lightweight Solution for Vision-to-Audio Generation by Connecting Foundation Models

Heng Wang, Jianbo Ma, Santiago Pascual, Richard Cartwright, Weidong Cai

专题命中 音频语音多模态 :cross-modal(abstract);分类 cs.CV、cs.AI、cs.MM

Comments AAAI 2024. Demo page: https://v2a-mapper.github.io/

详情

展开后加载摘要…

URL PDF HTML 收藏
2305.18500 2023-10-10 cs.CV cs.AI cs.CL cs.LG eess.AS 70%

VAST: A Vision-Audio-Subtitle-Text Omni-Modality Foundation Model and Dataset

Sihan Chen, Handong Li, Qunbo Wang, Zijia Zhao, Mingzhen Sun, Xinxin Zhu, Jing Liu

专题命中 音频语音多模态 :multi-modal(abstract);分类 cs.CV、cs.CL、cs.AI

Comments Accepted by NeurIPS 2023

详情

展开后加载摘要…

URL PDF HTML 收藏
2309.12158 2023-09-22 cs.SD cs.IR cs.LG eess.AS 70%

Towards Robust and Truly Large-Scale Audio-Sheet Music Retrieval

Luis Carvalho, Gerhard Widmer

专题命中 音频语音多模态 :multi-modal(abstract);cross-modal(abstract);分类 eess.AS

Comments Proceedings of the IEEE 6th International Conference on Multimedia Information Processing and Retrieval (MIPR)

详情

展开后加载摘要…

URL PDF HTML 收藏
2309.09501 2023-09-19 cs.CV 70%

Discovering Sounding Objects by Audio Queries for Audio Visual Segmentation

Shaofei Huang, Han Li, Yuqing Wang, Hongji Zhu, Jiao Dai, Jizhong Han, Wenge Rong, Si Liu

专题命中 音频语音多模态 :cross-modal(abstract);audio-visual(abstract);分类 cs.CV

Comments Accepted by IJCAI 2023

详情

展开后加载摘要…

URL PDF HTML 收藏
2308.07221 2023-08-28 cs.SD cs.LG eess.AS 70%

AudioFormer: Audio Transformer learns audio feature representations from discrete acoustic codes

Zhaohui Li, Haitao Wang, Xinghua Jiang

专题命中 音频语音多模态 :multimodal(abstract);audio-visual(abstract);分类 eess.AS

Comments Need to supplement more detailed experiments

详情

展开后加载摘要…

URL PDF HTML 收藏
2305.16103 2023-05-26 cs.CV cs.AI cs.CL cs.MM 70%

ChatBridge: Bridging Modalities with Large Language Model as a Language Catalyst

Zijia Zhao, Longteng Guo, Tongtian Yue, Sihan Chen, Shuai Shao, Xinxin Zhu, Zehuan Yuan, Jing Liu

专题命中 音频语音多模态 :multimodal(abstract);分类 cs.CV、cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2305.14359 2023-05-25 cs.MM cs.AI cs.CV cs.SD eess.AS 70%

Zero-shot personalized lip-to-speech synthesis with face image based voice control

Zheng-Yan Sheng, Yang Ai, Zhen-Hua Ling

专题命中 音频语音多模态 :cross-modal(abstract);分类 cs.CV、cs.AI、cs.MM

Comments ICASSP 2023

详情

展开后加载摘要…

URL PDF HTML 收藏
2305.04160 2023-05-23 cs.CL cs.AI cs.CV eess.AS 70%

X-LLM: Bootstrapping Advanced Large Language Models by Treating Multi-Modalities as Foreign Languages

Feilong Chen, Minglun Han, Haozhi Zhao, Qingyang Zhang, Jing Shi, Shuang Xu, Bo Xu

专题命中 音频语音多模态 :multimodal(abstract);分类 cs.CV、cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2305.12311 2023-05-23 cs.CL cs.AI cs.CV cs.LG eess.AS 70%

i-Code V2: An Autoregressive Generation Framework over Vision, Language, and Speech Data

Ziyi Yang, Mahmoud Khademi, Yichong Xu, Reid Pryzant, Yuwei Fang, Chenguang Zhu, Dongdong Chen, Yao Qian, Mei Gao, Yi-Ling Chen, Robert Gmyr, Naoyuki Kanda, Noel Codella, Bin Xiao, Yu Shi, Lu Yuan, Takuya Yoshioka, Michael Zeng, Xuedong Huang

专题命中 音频语音多模态 :multimodal(abstract);分类 cs.CV、cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2305.10763 2023-05-19 cs.SD eess.AS 70%

CLAPSpeech: Learning Prosody from Text Context with Contrastive Language-Audio Pre-training

Zhenhui Ye, Rongjie Huang, Yi Ren, Ziyue Jiang, Jinglin Liu, Jinzheng He, Xiang Yin, Zhou Zhao

专题命中 音频语音多模态 :multi-modal(abstract);cross-modal(abstract);分类 eess.AS

Comments Accepted by ACL 2023 (Main Conference)

详情

展开后加载摘要…

URL PDF HTML 收藏
2208.14114 2023-04-26 cs.CV 70%

Robust Sound-Guided Image Manipulation

Seung Hyun Lee, Gyeongrok Oh, Wonmin Byeon, Sang Ho Yoon, Jinkyu Kim, Sangpil Kim

专题命中 音频语音多模态 :multi-modal(abstract);image-text(abstract);分类 cs.CV

Comments arXiv admin note: text overlap with arXiv:2112.00007

详情

展开后加载摘要…

URL PDF HTML 收藏
2301.02379 2023-04-04 cs.CV 70%

CodeTalker: Speech-Driven 3D Facial Animation with Discrete Motion Prior

Jinbo Xing, Menghan Xia, Yuechen Zhang, Xiaodong Cun, Jue Wang, Tien-Tsin Wong

专题命中 音频语音多模态 :cross-modal(abstract);audio-visual(abstract);分类 cs.CV

Comments CVPR2023 Camera-Ready. Project Page: https://doubiiu.github.io/projects/codetalker/, Code: https://github.com/Doubiiu/CodeTalker

详情

展开后加载摘要…

URL PDF HTML 收藏
2303.00091 2023-03-02 eess.AS cs.AI cs.CL cs.CV cs.SD eess.IV 70%

Improving Medical Speech-to-Text Accuracy with Vision-Language Pre-training Model

Jaeyoung Huh, Sangjoon Park, Jeong Eun Lee, Jong Chul Ye

专题命中 音频语音多模态 :multi-modal(abstract);分类 cs.CV、cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2302.02845 2023-02-07 cs.SD cs.LG eess.AS 70%

Audio Representation Learning by Distilling Video as Privileged Information

Amirhossein Hajavi, Ali Etemad

专题命中 音频语音多模态 :multi-modal(abstract);audio-visual(abstract);分类 eess.AS

详情

展开后加载摘要…

URL PDF HTML 收藏
2211.01950 2022-12-23 cs.NE cs.CV cs.LG 70%

Unlocking the potential of two-point cells for energy-efficient and resilient training of deep nets

Ahsan Adeel, Adewale Adetomi, Khubaib Ahmed, Amir Hussain, Tughrul Arslan, W. A. Phillips

专题命中 音频语音多模态 :multimodal(abstract);audio-visual(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2211.14419 2022-11-29 cs.CV 70%

Panoramic Video Salient Object Detection with Ambisonic Audio Guidance

Xiang Li, Haoyuan Cao, Shijie Zhao, Junlin Li, Li Zhang, Bhiksha Raj

专题命中 音频语音多模态 :multimodal(abstract);audio-visual(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2107.03465 2022-11-04 cs.CV 70%

An audiovisual and contextual approach for categorical and continuous emotion recognition in-the-wild

Panagiotis Antoniadis, Ioannis Pikoulis, Panagiotis P. Filntisis, Petros Maragos

专题命中 音频语音多模态 :multi-modal(abstract);audio-visual(abstract);分类 cs.CV

Comments 7 pages, 1 figure, 3 tables, accepted to the 2nd Workshop and Competition on Affective Behavior Analysis In-the-Wild (ABAW2)

Journal ref 2021 IEEE/CVF International Conference on Computer Vision Workshops (ICCVW)

详情

展开后加载摘要…

URL PDF HTML 收藏
2210.05076 2022-10-12 cs.SD cs.IR eess.AS 70%

ConchShell: A Generative Adversarial Networks that Turns Pictures into Piano Music

Wanpeng Fan, Yuanzhi Su, Yuxin Huang

专题命中 音频语音多模态 :multimodal(abstract);multi-modal(abstract);分类 eess.AS

Comments 5 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2208.10238 2022-08-23 cs.CV 70%

Learning Branched Fusion and Orthogonal Projection for Face-Voice Association

Muhammad Saad Saeed, Shah Nawaz, Muhammad Haris Khan, Sajid Javed, Muhammad Haroon Yousaf, Alessio Del Bue

专题命中 音频语音多模态 :cross-modal(abstract);audio-visual(abstract);分类 cs.CV

Comments Submitted: IEEE Transactions on Multimedia. arXiv admin note: substantial text overlap with arXiv:2112.10483

详情

展开后加载摘要…

URL PDF HTML 收藏
2207.03048 2022-08-15 cs.CV 70%

AV-Gaze: A Study on the Effectiveness of Audio Guided Visual Attention Estimation for Non-Profilic Faces

Shreya Ghosh, Abhinav Dhall, Munawar Hayat, Jarrod Knibbe

专题命中 音频语音多模态 :cross-modal(abstract);audio-visual(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏