arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

多模态大模型

跨文本、图像、视频、音频等模态的大模型与学习方法。

共收录 4587 信号源:cs.CV, cs.CL, cs.AI, cs.MM, eess.AS

1. 音频语音多模态 4587 篇

2101.09725 2021-03-15 cs.CR 78%

Audio-Visual Biometric Recognition and Presentation Attack Detection: A Comprehensive Survey

Hareesh Mandalapu, P N Aravinda Reddy, Raghavendra Ramachandra, K Sreenivasa Rao, Pabitra Mitra, S R Mahadeva Prasanna, Christoph Busch

专题命中 音频语音多模态 :audio-visual(title,abstract)

Journal ref in IEEE Access, vol. 9, pp. 37431-37455, 2021

详情

展开后加载摘要…

URL PDF HTML 收藏
2103.05309 2021-03-10 cs.HC cs.CY 78%

Trade-offs in the Design of Multimodal Interaction for Older Adults

Gianluca Schiavo, Ornella Mich, Michela Ferron, Nadia Mana

专题命中 音频语音多模态 :multimodal(title,abstract)

Journal ref Behaviour & Information Technology, 2020

详情

展开后加载摘要…

URL PDF HTML 收藏
2012.01583 2020-12-04 cs.RO 78%

Multimodal Contact Detection using Auditory and Force Features for Reliable Object Placing in Household Environments

Jaime Maldonado, Asil Kaan Bozcuoğlu, Christoph Zetzsche

专题命中 音频语音多模态 :multimodal(title,abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2012.00922 2020-12-03 cs.HC 78%

Cross-Modal Terrains: Navigating Sonic Space through Haptic Feedback

Gabriella Isaac, Lauren Hayes, Todd Ingalls

专题命中 音频语音多模态 :cross-modal(title);multimodal(abstract)

Journal ref Proceedings of the International Conference on New Interfaces for Musical Expression, 2017

详情

展开后加载摘要…

URL PDF HTML 收藏
2004.05488 2020-09-03 cs.NE cs.LG q-bio.NC 78%

Brain-inspired self-organization with cellular neuromorphic computing for multimodal unsupervised learning

Lyes Khacef, Laurent Rodriguez, Benoit Miramond

专题命中 音频语音多模态 :multimodal(title,abstract)

Comments Preprint

详情

展开后加载摘要…

URL PDF HTML 收藏
2008.13510 2020-09-01 cs.HC 78%

Multi-Modal End-User Programming of Web-Based Virtual Assistant Skills

Michael H. Fischer, Giovanni Campagna, Euirim Choi, Monica S. Lam

专题命中 音频语音多模态 :multi-modal(title,abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2008.11834 2020-08-28 cs.HC 78%

Conversations On Multimodal Input Design With Older Adults

Adam S. Williams, Sarah Coler, Francisco Ortega

专题命中 音频语音多模态 :multimodal(title,abstract)

Comments Presented at the CHI 2020 Designing Interactions for the Ageing Populations Addressing Global Challenges Workshop

详情

展开后加载摘要…

URL PDF HTML 收藏
2005.08271 2020-08-12 cs.CV cs.CL cs.LG cs.SD eess.AS 78%

A Better Use of Audio-Visual Cues: Dense Video Captioning with Bi-modal Transformer

Vladimir Iashin, Esa Rahtu

专题命中 音频语音多模态 :audio-visual(title);分类 cs.CV、cs.CL、eess.AS

Comments Accepted by BMVC 2020. More experiments. Code: https://github.com/v-iashin/bmt Project page: https://v-iashin.github.io/bmt

详情

展开后加载摘要…

URL PDF HTML 收藏
2006.03813 2020-06-09 cs.HC 78%

Multimodal Systems: Taxonomy, Methods, and Challenges

Muhammad Zeeshan Baig, Manolya Kavakli

专题命中 音频语音多模态 :multimodal(title,abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2003.09670 2020-06-02 cs.CY cs.HC cs.LG 78%

Identifying At-Risk K-12 Students in Multimodal Online Environments: A Machine Learning Approach

Hang Li, Wenbiao Ding, Zitao Liu

专题命中 音频语音多模态 :multimodal(title,abstract)

Comments The 13th International Conference on Educational Data Mining (EDM 2020), 2020

详情

展开后加载摘要…

URL PDF HTML 收藏
2004.14505 2020-05-01 cs.HC 78%

Touch? Speech? or Touch and Speech? Investigating Multimodal Interaction for Visual Network Exploration and Analysis

Ayshwarya Saktheeswaran, Arjun Srinivasan, John Stasko

专题命中 音频语音多模态 :multimodal(title,abstract)

Comments 12 pages, 3 figures, 3 tables

Journal ref IEEE Transactions on Visualization and Computer Graphics. January 2020

详情

展开后加载摘要…

URL PDF HTML 收藏
2004.13908 2020-04-30 cs.HC 78%

Interactive Rainbow Score: A Visual-centered Multimodal Flute Tutoring System

Daniel Chin, Yian Zhang, Tianyu Zhang, Jake Zhao, Gus G. Xia

专题命中 音频语音多模态 :multimodal(title,abstract)

Comments NIME 2020 poster presentation. 6 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2004.10428 2020-04-23 cs.HC 78%

Interweaving Multimodal Interaction with Flexible Unit Visualizations for Data Exploration

Arjun Srinivasan, Bongshin Lee, John Stasko

专题命中 音频语音多模态 :multimodal(title,abstract)

Comments 15 pages, 9 figures, 3 tables

Journal ref IEEE Transactions on Visualization and Computer Graphics. March 2020

详情

展开后加载摘要…

URL PDF HTML 收藏
2003.14392 2020-04-01 cs.HC cs.RO 78%

Multimodal Interfaces for Effective Teleoperation

Eleftherios Triantafyllidis, Christopher McGreavy, Jiacheng Gu, Zhibin Li

专题命中 音频语音多模态 :multimodal(title);multi-modal(abstract)

Comments 15 pages, 28 figures, 5 tables, 5 equations

详情

展开后加载摘要…

URL PDF HTML 收藏
2002.05070 2020-02-13 cs.CV cs.LG cs.MM cs.SD eess.AS 78%

AlignNet: A Unifying Approach to Audio-Visual Alignment

Jianren Wang, Zhaoyuan Fang, Hang Zhao

专题命中 音频语音多模态 :audio-visual(title);分类 cs.CV、cs.MM、eess.AS

Comments WACV2020. Project video and code are available at https://jianrenw.github.io/AlignNet

详情

展开后加载摘要…

URL PDF HTML 收藏
2002.02671 2020-02-10 cs.GR cs.HC cs.NE 78%

Audio-Visual-Olfactory Resource Allocation for Tri-modal Virtual Environments

Efstratios Doukakis, Kurt Debattista, Thomas Bashford-Rogers, Amar Dhokia, Ali Asadipour, Alan Chalmers, Carlo Harvey

专题命中 音频语音多模态 :audio-visual(title);multi-modal(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2001.06423 2020-01-20 cs.HC 78%

InChorus: Designing Consistent Multimodal Interactions for Data Visualization on Tablet Devices

Arjun Srinivasan, Bongshin Lee, Nathalie Henry Riche, Steven M. Drucker, Ken Hinckley

专题命中 音频语音多模态 :multimodal(title,abstract)

Comments To appear in ACM CHI 2020 Conference on Human Factors in Computing Systems; 13 pages (10 content + 3 references); 4 Figures, 1 Table

详情

展开后加载摘要…

URL PDF HTML 收藏
1911.05544 2019-12-03 cs.LG stat.ML 78%

Learning Relationships between Text, Audio, and Video via Deep Canonical Correlation for Multimodal Language Analysis

Zhongkai Sun, Prathusha Sarma, William Sethares, Yingyu Liang

专题命中 音频语音多模态 :multimodal(title,abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
1907.00758 2019-11-28 cs.CL cs.CV cs.LG cs.SD eess.AS eess.IV 78%

Synchronising audio and ultrasound by learning cross-modal embeddings

Aciel Eshky, Manuel Sam Ribeiro, Korin Richmond, Steve Renals

专题命中 音频语音多模态 :cross-modal(title);分类 cs.CV、cs.CL、eess.AS

Comments 5 pages, 1 figure, 4 tables; Interspeech 2019 with the following edits: 1) Loss and accuracy upon convergence were accidentally reported from an older model. Now updated with model described throughout the paper. All other results remain unchanged. 2) Max true offset in the training data corrected from 179ms to 1789ms. 3) Detectability "boundary/range" renamed to detectability "thresholds"

详情

展开后加载摘要…

URL PDF HTML 收藏
1907.02426 2019-07-05 cs.RO cs.HC cs.LG stat.ML 78%

Multimodal Uncertainty Reduction for Intention Recognition in Human-Robot Interaction

Susanne Trick, Dorothea Koert, Jan Peters, Constantin Rothkopf

专题命中 音频语音多模态 :multimodal(title,abstract)

Comments Submitted to IROS 2019

详情

展开后加载摘要…

URL PDF HTML 收藏
1906.01810 2019-06-06 cs.NI eess.SP 78%

A Sustainable Multi-modal Multi-layer Emotion-aware Service at the Edge

Long Hu, Wei Li, Jun Yang, Giancarlo Fortino, Min Chen

专题命中 音频语音多模态 :multi-modal(title,abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
1906.01197 2019-06-05 cs.HC 78%

Adaptive Multimodal Music Learning via Interactive-haptic Instrument

Yian Zhang, Yinmiao Li, Daniel Chin, Gus Xia

专题命中 音频语音多模态 :multimodal(title,abstract)

Comments 6 pages, 14 figures, 2 tables. This paper is accepted by NIME 2019(New Interface for Musical Expression)

详情

展开后加载摘要…

URL PDF HTML 收藏
1903.12133 2019-03-29 cs.HC 78%

A Multimodal Emotion Sensing Platform for Building Emotion-Aware Applications

Daniel McDuff, Kael Rowan, Piali Choudhury, Jessica Wolk, ThuVan Pham, Mary Czerwinski

专题命中 音频语音多模态 :multimodal(title,abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
1903.07171 2019-03-19 cs.CY 78%

Responsible and Representative Multimodal Data Acquisition and Analysis: On Auditability, Benchmarking, Confidence, Data-Reliance & Explainability

Alice Baird, Simone Hantke, Björn Schuller

专题命中 音频语音多模态 :multimodal(title,abstract)

Comments 4 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
1807.08804 2018-07-25 cs.DB cs.PF 78%

GPU-based Commonsense Paradigms Reasoning for Real-Time Query Answering and Multimodal Analysis

Nguyen Ha Tran, Erik Cambria

专题命中 音频语音多模态 :multimodal(title,abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
1805.01369 2018-05-04 cs.AI cs.CL cs.CV 78%

Framewise approach in multimodal emotion recognition in OMG challenge

Grigoriy Sterling, Andrey Belyaev, Maxim Ryabov

专题命中 音频语音多模态 :multimodal(title);分类 cs.CV、cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
1803.04713 2018-03-14 cs.HC 78%

A Gaze-Assisted Multimodal Approach to Rich and Accessible Human-Computer Interaction

Vijay Rajanna, Tracy Hammond

专题命中 音频语音多模态 :multimodal(title,abstract)

Comments 4 pages, 5 figures, ACM Richard Tapia Conference, Atlanta, 2017

详情

展开后加载摘要…

URL PDF HTML 收藏
1711.10886 2017-11-30 cs.CY 78%

A Multimodal Assistive System for Helping Visually Impaired in Social Interactions

M. Saquib Sarfraz, Angela Constantinescu, Melanie Zuzej, Rainer Stiefelhagen

专题命中 音频语音多模态 :multimodal(title);multi-modal(abstract)

Journal ref Informatik Spectrum, Springer volume 40,No. 6. 2017

详情

展开后加载摘要…

URL PDF HTML 收藏
1706.09739 2017-07-25 cs.IR cs.LG 78%

A Deep Multimodal Approach for Cold-start Music Recommendation

Sergio Oramas, Oriol Nieto, Mohamed Sordo, Xavier Serra

专题命中 音频语音多模态 :multimodal(title,abstract)

Comments In Proceedings of the 2nd Workshop on Deep Learning for Recommender Systems (DLRS 2017), collocated with RecSys 2017

详情

展开后加载摘要…

URL PDF HTML 收藏
1511.03690 2015-11-13 cs.CV cs.AI cs.CL 78%

Deep Multimodal Semantic Embeddings for Speech and Images

David Harwath, James Glass

专题命中 音频语音多模态 :multimodal(title);分类 cs.CV、cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏