arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

AI 大模型

多模态大模型

跨文本、图像、视频、音频等模态的大模型与学习方法。

共收录 46430 信号源:cs.CV, cs.CL, cs.AI, cs.MM, eess.AS

1. 音频语音多模态 4597 篇

2601.14357 2026-01-26 astro-ph.IM 50%

One Attempt at Building an Inclusive & Accessible Hybrid Astronomy Conference: FRB 2025

一次构建包容与可及性混合天文学会议的尝试:FRB 2025

Alice P. Curtin, Reshma Anna-Thomas, Amanda M. Cook, Carolina Cruz-Vinaccia, Jason Hessels, Robert Main, Inés Pastor Marazuela, Lauren Rhodes, Vishwangi Shah

专题命中 音频语音多模态 :audio-visual(abstract)

AI总结 FRB 2025会议通过混合模式、低费用和教育活动,为早期研究人员和欠资地区提供可及性,展示了低成本高包容性的会议可能性。

Comments 10 pages, 4 figures; Conference Reflection

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.08422 2026-01-22 cs.RO 50%

Teaching Robots Like Dogs: Learning Agile Navigation from Luring, Gesture, and Speech

教机器人像教狗一样:从诱饵、手势和言语中学习敏捷导航

Taerim Yoon, Dongho Kang, Jin Cheng, Fatemeh Zargarbashi, Yijiang Huang, Minsung Ahn, Stelian Coros, Sungjoon Choi

机构 * Department of Artificial Intelligence, Korea University(人工智能系,韩国大学) Department of Computer Science, ETH Zurich(计算机科学系,苏黎世联邦理工学院) Department of Mechanical and Aerospace Engineering, UCLA(机械与航空航天工程系,加州大学洛杉矶分校)

专题命中 音频语音多模态 :multimodal(abstract)

AI总结 本研究提出了一种人机协同框架,通过手势、言语和诱饵等多模态输入,使机器人高效学习敏捷导航,实验显示在少于1小时的数据下任务成功率高达97.15%。

Comments 10 pages, 7 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.13355 2026-01-21 cs.HC 50%

Remote Triggers: Misophonia, Technology Non-Use, and Design for Inclusive Digital Spaces

远程触发:对声音厌恶、技术不使用与包容性数字空间的设计

Tawfiq Ammari, Samantha Gilgan

专题命中 音频语音多模态 :audio-visual(abstract)

AI总结 本研究探讨声音厌恶者在数字空间中的体验,提出设计干预以减少技术不使用和排斥问题。

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.11328 2026-01-21 cs.HC 50%

ProjecTA: A Semi-Humanoid Robotic Teaching Assistant with In-Situ Projection for Guided Tours

ProjecTA:一种配备现场投影的半人形教学助手用于导览

Hanqing Zhou, Yichuan Zhang, Zihan Zhang, Wei Zhang, Chao Wang, Pengcheng An

专题命中 音频语音多模态 :cross-modal(abstract)

AI总结 ProjecTA通过现场投影与手势协调,减少认知负荷并提升学习体验,为移动学习提供新的设计方向。

Comments 35 pages, 12 figures, 2 appendixes, 3 supplementary meterials, to appear at ACM CHI'26

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.11262 2026-01-19 cs.SD cs.IR cs.LG 50%

Scalable Music Cover Retrieval Using Lyrics-Aligned Audio Embeddings

基于歌词对齐音频嵌入的可扩展音乐封面检索

Joanne Affolter, Benjamin Martin, Elena V. Epure, Gabriel Meseguer-Brocal, Frédéric Kaplan

机构 * Deezer Research, Paris, France(DeepZoom研究机构,法国巴黎) EPFL, Lausanne, Switzerland(瑞士洛桑联邦理工学院)

专题命中 音频语音多模态 :multimodal(abstract)

AI总结 LIVI通过利用歌词信息提升音乐封面检索的准确性和效率,减少对复杂模型的依赖。

Comments Published at ECIR 2026 (European Conference of Information Retrieval)

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.10119 2026-01-16 cs.CR 50%

Advanced Encryption Technique for Multimedia Data Using Sudoku-Based Algorithms for Enhanced Security

基于数独算法的多媒体数据高级加密技术

Mithil Bavishi, Anuj Bohra, Kushal Vadodaria, Abhinav Bohra, Neha Katre, Ramchandra Mangrulkar, Vinaya Sawant

专题命中 音频语音多模态 :multimodal(abstract)

AI总结 本文提出了一种基于数独算法的高级加密技术,用于增强多媒体数据的安全性,通过时间戳依赖密钥生成并支持多种媒体类型。

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.06774 2026-01-13 cs.HC 50%

ImmuniFraug: A Metacognitive Intervention Anti-Fraud Approach to Enhance Undergraduate Students' Cyber Fraud Awareness

ImmuniFraug:一种基于元认知干预的反欺诈方法,以提高本科生的网络欺诈意识

Xiangzhe Yuan, Jiajun Wang, Huanchen Wang, Qian Wan, Siying Hu

专题命中 音频语音多模态 :multimodal(abstract)

AI总结 ImmuniFraug通过基于LLM的多模态欺诈模拟,提升本科生网络欺诈意识和自我效能感,提供更有效的反欺诈教育方法。

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.17609 2026-01-01 cs.SD cs.LG 50%

Audio Super-Resolution with Latent Bridge Models

基于潜在桥模型的音频超分辨率

Chang Li, Zehua Chen, Liyuan Wang, Jun Zhu

机构 * Department of CST, Tsinghua University, Beijing, China(计算机科学与技术系,清华大学,北京,中国) Shengshu AI, Beijing, China(盛舒人工智能,北京,中国) USTC, Hefei, China(中国科学技术大学,合肥,中国)

专题命中 音频语音多模态 :any-to-any(abstract)

AI总结 本文提出基于潜在桥模型的音频超分辨率方法,通过压缩音频到潜在空间并设计桥模型,实现高质量的任意到48kHz和192kHz音频上采样。

Comments Accepted at NeurIPS 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.21566 2025-12-29 physics.optics 50%

Broadband tunable microwave photonic radar for simultaneous detection of human respiration, heartbeat, and speech with deep learning-based speech recognition

宽带可调微波光子雷达用于同时检测人体呼吸、心跳和语音的系统,结合深度学习语音识别

Lei Gao, Dingding Liang, Jiawei Gao, Chulun Lin, Zhiqiang Huang, Taixia Shi, Yang Chen

专题命中 音频语音多模态 :multimodal(abstract)

AI总结 本研究提出一种宽带可调微波光子雷达系统,结合深度学习实现对呼吸、心跳和语音的同步检测,实验验证其高精度语音识别和多模态生命体征监测能力。

Comments 40 pages, 14 figures, 3 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.17241 2025-12-22 cs.RO cs.HC 50%

A Service Robot's Guide to Interacting with Busy Customers

服务机器人与忙碌顾客互动指南

Suraj Nukala, Meera Sushma, Leimin Tian, Akansel Cosgun, Dana Kulic

机构 * Faculty of Engineering, Monash University(蒙纳什大学工程学院) CSIRO Robotics(CSIRO机器人研究中心) School of Information Technology, Deakin University(德金大学信息技术学院)

专题命中 音频语音多模态 :multimodal(abstract)

AI总结 本研究探讨了服务机器人在模拟餐厅场景中通过不同沟通方式与忙碌顾客互动的效果,发现视觉显示在传达意图上效果最佳,而语音在捕捉注意力方面表现突出。

Comments Presented at ACRA 2025. 10 pages, 4 figures. Includes a user study (N=24) using the Temi robot evaluating speech, visual, and micromotion modalities

Journal ref Proceedings of the 2025 Australasian Conference on Robotics and Automation (ACRA 2025)

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.17228 2025-12-22 cs.HC 50%

LUMIA: A Handheld Vision-to-Music System for Real-Time, Embodied Composition

LUMIA: 一种用于实时具身创作的手持视觉到音乐系统

Chung-Ta Huang, Connie Cheng, Vealy Lai

专题命中 音频语音多模态 :multimodal(abstract)

AI总结 LUMIA通过视觉到音乐的实时具身创作系统,将环境互动与生成式AI结合,实现基于感知的即兴音乐创作。

Comments 6 pages, 15 pages with appendix, NeurIPS 2025 Creative AI track

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.13981 2025-12-17 cs.RO 50%

Impact of Robot Facial-Audio Expressions on Human Robot Trust Dynamics and Trust Repair

机器人面部-语音表达对人类机器人信任动态及信任修复的影响

Hossein Naderi, Alireza Shojaei, Philip Agee, Kereshmeh Afsari, Abiola Akanmu

专题命中 音频语音多模态 :multimodal(abstract)

AI总结 本研究探讨了机器人面部-语音表达对人类信任动态及修复的影响,发现成功提升信任,失败导致信任下降,道歉表达部分恢复信任,且年龄和先前态度影响信任变化的持续时间。

详情

展开后加载摘要…

URL PDF HTML 收藏
2210.17139 2025-12-04 cs.LO cs.DM math.LO 50%

Nested Sequents for Intuitionistic Grammar Logics via Structural Refinement

通过结构细化方法为直觉语法学逻辑构建嵌套序列为

Tim S. Lyon

专题命中 音频语音多模态 :multi-modal(abstract)

AI总结 本文通过结构细化方法,为直觉语法学逻辑构建了无割嵌套序列表演系统,并证明了其保守性、不可判定性和可判定子类。

Comments This paper is currently under review. arXiv admin note: text overlap with arXiv:2107.01998

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.02515 2025-12-03 cs.SD 50%

VibOmni: Towards Scalable Bone-conduction Speech Enhancement on Earables

VibOmni: 向耳戴设备的可扩展骨传导语音增强迈进

Lixing He, Yunqi Guo, Haozheng Hou, Zhenyu Yan

机构 * department of information engineering, The Chinese University of Hong Kong(信息工程系,香港中文大学)

专题命中 音频语音多模态 :multi-modal(abstract)

AI总结 VibOmni通过骨传导振动与音频融合,提升耳戴设备在嘈杂环境中的语音质量与噪声抑制能力。

Comments Submitted to TMC

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.19947 2025-11-26 cs.IT eess.SP math.IT 50%

Towards Edge General Intelligence: Knowledge Distillation for Mobile Agentic AI

迈向边缘泛智能:面向移动代理AI的知识蒸馏

Yuxuan Wu, Linghan Ma, Ruichen Zhang, Yinqiu Liu, Dusit Niyato, Shunpu Tang, Zehui Xiong, Zhu Han, Zhaohui Yang, Kaibin Huang, Zhaoyang Zhang, Kai-Kit Wong

专题命中 音频语音多模态 :multimodal(abstract)

AI总结 本文探讨了将知识蒸馏整合到边缘泛智能中的方法,旨在提升移动边缘计算中智能代理的效率与可扩展性。

Comments 21 pages, 6 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.02792 2025-11-25 cs.HC 50%

Where Do Passengers Gaze? Impact of Passengers' Personality Traits on Their Gaze Pattern Toward Pedestrians During APMV-Pedestrian Interactions with Diverse eHMIs

乘客注视方向在哪里?乘客个性特征对其在APMV-行人互动中对行人注视模式的影响

Hailong Liu, Zhe Zeng, Takahiro Wada

专题命中 音频语音多模态 :multimodal(abstract)

AI总结 研究探讨了乘客个性特征如何影响其在APMV与行人互动中的注视模式,发现不同eHMI设计对乘客注视行为有显著影响。

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.16769 2025-11-24 cs.HC 50%

Trust in AI emerges from distrust in humans: A machine learning study on decision-making guidance

对人工智能的信任源于对人类的不信任:一项关于决策指导的机器学习研究

Johan Sebastián Galindez-Acosta, Juan José Giraldo-Huertas

专题命中 音频语音多模态 :multimodal(abstract)

AI总结 本研究通过机器学习探讨了人工智能在决策指导中的信任机制,发现对人类的不信任会促使人们转向AI,且AI在事实性情景中更受青睐。

Comments 36 pages, 6 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.10108 2025-11-21 cs.LG q-bio.NC 50%

NeuroXVocal: Detection and Explanation of Alzheimer's Disease through Non-invasive Analysis of Picture-prompted Speech

NeuroXVocal:通过非侵入性分析提示语音检测和解释阿尔茨海默病

Nikolaos Ntampakis, Konstantinos Diamantaras, Ioanna Chouvarda, Magda Tsolaki, Vasileios Argyriou, Panagiotis Sarigianndis

机构 * International Hellenic University, Sindos, Greece MetaMind Innovations, Kozani, Greece Aristotle University of Thessaloniki, Thessaloniki, Greece Greek Association of Alzheimer’s Disease \& Related Disorders, Thessaloniki, Greece Kingston University London, London, UK University of Western Macedonia, Kozani, Greece

专题命中 音频语音多模态 :cross-modal(abstract)

AI总结 NeuroXVocal通过非侵入性语音分析实现阿尔茨海默病的检测与解释,结合高精度分类和文献支持的可解释性,提升临床诊断效果。

Journal ref Medical Image Computing and Computer Assisted Intervention - MICCAI 2025. Lecture Notes in Computer Science, vol 15973. Springer, Cham (2025)

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.04211 2025-11-07 cs.DL 50%

From data to corpus: semiotic and documentary issues in audiovisual archives

Peter Stockinger

专题命中 音频语音多模态 :multimodal(abstract)

Comments in French language

Journal ref Corpus audiovisuels. Quelles approches ? Quels usages ?, Editions des archives contemporaines, 2022, 9782813003799

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.00611 2025-11-04 math.OC 50%

From Generality to Specificity: Prior-Driven Optimal Sparse Transformation in Compressed Sensing

Zhihan Zhu, Yanhao Zhang, Yong Xia

专题命中 音频语音多模态 :multimodal(abstract)

Comments 47 pages, 10 figures, 1 table

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.26096 2025-10-31 cs.SD cs.CR cs.LG 50%

ALMGuard: Safety Shortcuts and Where to Find Them as Guardrails for Audio-Language Models

Weifei Jin, Yuxin Cao, Junjie Su, Minhui Xue, Jie Hao, Ke Xu, Jin Song Dong, Derui Wang

机构 * Beijing University of Posts and Telecommunications(北京邮电大学) National University of Singapore(新加坡国立大学) CSIRO’s Data61(CSIRO数据61) Responsible AI Research (RAIR) Centre, The University of Adelaide(负责任人工智能研究(RAIR)中心,阿德莱德大学) Tsinghua University(清华大学)

专题命中 音频语音多模态 :multimodal(abstract)

Comments Accepted to NeurIPS 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.23887 2025-10-29 cs.HC 50%

MORA: AI-Mediated Story-Based practice for Speech Sound Disorder from Clinic to Home

Sumin Hong, Xavier Briggs, Qingxiao Zheng, Yao Du, Jinjun Xiong, Toby Jia-jun Li

专题命中 音频语音多模态 :multimodal(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.24180 2025-10-29 cs.LG 50%

V-SAT: Video Subtitle Annotation Tool

Arpita Kundu, Joyita Chakraborty, Anindita Desarkar, Aritra Sen, Srushti Anil Patil, Vishwanathan Raman

机构 * LTIMindTree

专题命中 音频语音多模态 :audio-visual(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.23203 2025-10-24 cs.RO 50%

CE-Nav: Flow-Guided Reinforcement Refinement for Cross-Embodiment Local Navigation

Kai Yang, Tianlin Zhang, Zhengbo Wang, Zedong Chu, Xiaolong Wu, Yang Cai, Mu Xu

机构 * AMAP, Alibaba Group(阿里集团AMAP)

专题命中 音频语音多模态 :multi-modal(abstract)

Comments Project Page: https://ce-nav.github.io/. Code is available at https://github.com/amap-cvlab/CE-Nav

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.13558 2025-10-16 cs.SD 50%

Steer-MoE: Efficient Audio-Language Alignment with a Mixture-of-Experts Steering Module

Ruitao Feng, Bixi Zhang, Sheng Liang, Zheng Yuan

机构 * The University of Hong Kong, Fauclty of Science, Hong Kong(香港大学科学学院) Aix-Marseille University, Laboratoire Parole et Langage (LPL), France(艾克斯-马赛大学语言与言语实验室(LPL))

专题命中 音频语音多模态 :multimodal(abstract)

Comments 5 pages, 1 figures. Code is available at: https://github.com/forfrt/SteerMoE. Submitted to ICASSP 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.05249 2025-10-08 cs.HC 50%

CLAd-VR: Cognitive Load-based Adaptive Training for Machining Tasks in Virtual Reality

Bhavya Matam, Adamay Mann, Kachina Studer, Christian Gabbianelli, Sonia Castelo, John Liu, Claudio Silva, Dishita Turakhia

专题命中 音频语音多模态 :multimodal(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.02206 2025-10-03 cs.LG 50%

Poolformer: Recurrent Networks with Pooling for Long-Sequence Modeling

Daniel Gallo Fernández

机构 * MSc Artificial Intelligence Master Thesis(人工智能硕士论文)

专题命中 音频语音多模态 :multi-modal(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.26593 2025-10-01 cs.HC 50%

Exploring Large Language Model as an Interactive Sports Coach: Lessons from a Single-Subject Half Marathon Preparation

Kichang Lee

专题命中 音频语音多模态 :multimodal(abstract)

Comments 23 pages, 21 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2409.14608 2025-10-01 cs.RO 50%

Visual-auditory Extrinsic Contact Estimation

Xili Yi, Jayjun Lee, Nima Fazeli

机构 * Robotics Department, University of Michigan(密歇根大学机器人系)

专题命中 音频语音多模态 :multimodal(abstract)

Comments 8 pages, 7 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.16286 2025-09-23 cs.CY 50%

What's Not on the Plate? Rethinking Food Computing through Indigenous Indian Datasets

Pamir Gogoi, Neha Joshi, Ayushi Pandey, Deepthi Sudharsan, Saransh Kumar Gupta, Lipika Dey, Partha Pratim Das, Kalika Bali, Vivek Seshadri

专题命中 音频语音多模态 :multimodal(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏