arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

AI 大模型

多模态大模型

跨文本、图像、视频、音频等模态的大模型与学习方法。

2026-02-17 至 2026-02-17 共收录 101 信号源:cs.CV, cs.CL, cs.AI, cs.MM, eess.AS

1. 图文多模态 9 篇

2602.13758 2026-02-17 cs.CV cs.AI 86%

OmniScience: A Large-scale Multi-modal Dataset for Scientific Image Understanding

OmniScience: 一个大规模多模态数据集用于科学图像理解

Haoyi Tao, Chaozheng Huang, Nan Wang, Han Lyu, Linfeng Zhang, Guolin Ke, Xi Fang

机构 * DP Technology(DP技术)

专题命中 图文多模态 :multi-modal(title,abstract);multimodal(abstract);image-text(abstract);分类 cs.CV、cs.AI

AI总结 OmniScience是一个大规模多模态数据集,通过动态模型路由生成高信息密度的图像标题,提升多模态模型在科学图像理解上的性能。

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.14889 2026-02-17 cs.LG cs.CV cs.ET cs.HC cs.NE 79%

Web-Scale Multimodal Summarization using CLIP-Based Semantic Alignment

基于CLIP的语义对齐的网络级多模态摘要

Mounvik K, N Harshit

机构 * School of Computer Science Engineering(计算机科学与工程学院) VIT-AP University(VIT-AP大学)

专题命中 图文多模态 :multimodal(title,abstract);分类 cs.CV

AI总结 本文提出基于CLIP的语义对齐网络级多模态摘要框架,通过结合网络文本和图像数据生成摘要,实现高准确率的多模态对齐。

详情

展开后加载摘要…

URL PDF HTML 收藏
2412.20110 2026-02-17 cs.CV 79%

Cross-Modal Mapping: Mitigating the Modality Gap for Few-Shot Image Classification

跨模态映射:缓解模态差距以实现少样本图像分类

Xi Yang, Pai Peng, Wulin Xie, Xiaohuan Lu, Jie Wen

专题命中 图文多模态 :cross-modal(title,abstract);分类 cs.CV

AI总结 本文提出跨模态映射方法,通过全局对齐和三元组损失优化,缓解模态差距,提升少样本图像分类性能。

Comments The authors request withdrawal of this article. This version was submitted in error. Compared to the intended final version, it contains inaccuracies and fails to accurately reflect the authors' work and conclusions

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.12916 2026-02-17 cs.CV cs.LG 77%

Reliable Thinking with Images

基于图像的可靠思考

Haobin Li, Yutong Yang, Yijie Lin, Xiang Dai, Mouxing Yang, Xi Peng

机构 * College of Computer Science, Sichuan University(四川大学计算机科学学院) Southwest China Institute of Electronic Technology(西南中国电子技术研究所) National Key Laboratory of Fundamental Algorithms(国家基础算法重点实验室) Models for Engineering Simulation, Sichuan University(工程模拟模型研究所,四川大学)

专题命中 图文多模态 :multimodal(abstract);multi-modal(abstract);image-text(abstract);分类 cs.CV

AI总结 本文提出RTWI方法,通过估计视觉提示和文本Co T的可靠性,以解决多模态推理中噪声思考问题,提升多模态大语言模型的性能。

Comments 26 pages, 19 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.14225 2026-02-17 cs.AI 70%

Text Before Vision: Staged Knowledge Injection Matters for Agentic RLVR in Ultra-High-Resolution Remote Sensing Understanding

文本优先于视觉:针对超高清遥感理解的代理强化学习在超高清遥感理解中的知识注入至关重要

Fengxiang Wang, Mingshuo Chen, Yueying Li, Yajie Yang, Yuhao Zhou, Di Wang, Yifan Zhang, Haoyu Wang, Haiyan Zhao, Hongda Sun, Long Lan, Jun Song, Yulin Wang, Jing Zhang, Wenlong Zhang, Bo Du

机构 * National University of Defense Technology, China(国防科技大学) Beijing University of Posts and Telecommunications, China(北京邮电大学) University of the Chinese Academy of Sciences, China(中国科学院大学) Sichuan University, China(四川大学) Wuhan University, China(武汉大学) Chinese Academy of Science, China(中国科学院) Tsinghua University, China(清华大学) Shanghai Artificial Intelligence Laboratory, China(上海人工智能实验室) Renmin University of China, China(中国人民大学)

专题命中 图文多模态 :multimodal(abstract);image-text(abstract);分类 cs.AI

AI总结 本文提出分阶段知识注入方法,利用文本引导提升超高清遥感理解的视觉推理能力,实现XLRS-Bench上的高准确率。

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.13352 2026-02-17 cs.CV cs.AI cs.CL 67%

Using Deep Learning to Generate Semantically Correct Hindi Captions

使用深度学习生成语义正确的印地语描述

Wasim Akram Khan, Anil Kumar Vuppala

专题命中 图文多模态 :multi-modal(abstract);分类 cs.CV、cs.CL、cs.AI

AI总结 本研究利用深度学习生成印地语图像描述,采用VGG16和双向LSTM结合注意力机制,实现语义准确的描述生成。

Comments 34 pages, 12 figures, 3 tables. Master's thesis, Liverpool John Moores University, November 2022

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.14425 2026-02-17 cs.CV 57%

Hierarchical Vision-Language Interaction for Facial Action Unit Detection

层次化视觉-语言交互用于面部动作单元检测

Yong Li, Yi Ren, Yizhe Zhang, Wenhua Zhang, Tianyi Zhang, Muyun Jiang, Guo-Sen Xie, Cuntai Guan

机构 * Key Laboratory of Child Development and Learning Science (Ministry of Education), School of Biological Sciences and Medical Engineering, Southeast University(儿童发展与学习科学重点实验室(教育部),生物科学与医学工程学院,东南大学) School of Computer Science and Engineering, Nanjing University of Science and Technology(计算机科学与工程学院,南京理工大学) School of Computer Science and Engineering, Nanyang Technological University(计算机科学与工程学院,南洋理工大学)

专题命中 图文多模态 :cross-modal(abstract);分类 cs.CV

AI总结 HiVA通过层次化视觉-语言交互方法,利用文本描述和多模态注意力机制提升面部动作单元检测的鲁棒性和语义丰富性。

Comments Accepted to IEEE Transaction on Affective Computing 2026

Journal ref IEEE Transaction on Affective Computing 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.15704 2026-02-17 cs.CV 57%

Pyramid Token Pruning for High-Resolution Large Vision-Language Models via Region, Token, and Instruction-Guided Importance

通过区域、令牌和指令引导的重要性进行高分辨率大视觉-语言模型的金字塔令牌修剪

Yuxuan Liang, Xu Li, Xiaolei Chen, Yi Zheng, Haotian Chen, Bin Li, Xiangyang Xue

机构 * Shanghai Key Laboratory of Intelligent Information Processing(上海智能信息处理重点实验室) College of Computer Science and Artificial Intelligence(计算机科学与人工智能学院)

专题命中 图文多模态 :multimodal(abstract);分类 cs.CV

AI总结 通过区域、令牌和指令引导的重要性进行高分辨率大视觉-语言模型的金字塔令牌修剪,有效降低计算和内存开销,同时保持性能

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.13157 2026-02-17 eess.SP 50%

Seeing Radio: From Zero RF Priors to Explainable Modulation Recognition with Vision Language Models

看见无线电:从零RF先验到可解释的调制识别与视觉语言模型

Hang Zou, Bohao Wang, Yu Tian, Lina Bariah, Chongwen Huang, Samson Lasaulce, Mérouane Debbah

专题命中 图文多模态 :multimodal(abstract)

AI总结 本文提出利用视觉语言模型直接感知无线电波信号并推断调制模式,通过转换IQ流为图像数据,提升模型准确性至90%,并实现可解释的调制识别。

详情

展开后加载摘要…

URL PDF HTML 收藏

2. 音频语音多模态 12 篇

2602.13640 2026-02-17 cs.RO cs.AI 83%

Hierarchical Audio-Visual-Proprioceptive Fusion for Precise Robotic Manipulation

层次化音频-视觉-本体感知融合用于精确机器人操作

Siyuan Li, Jiani Lu, Yu Song, Xianren Li, Bo An, Peng Liu

机构 * Harbin Institute of Technology, China(哈尔滨工业大学) Nanyang Technological University, Singapore(南洋理工大学)

专题命中 音频语音多模态 :audio-visual(title);multimodal(abstract);cross-modal(abstract);分类 cs.AI

AI总结 本文提出层次化多模态融合框架,通过整合音频、视觉和本体感知信息,提升机器人精确操作性能。

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.05847 2026-02-17 cs.AI cs.CV 81%

OmniVideo-R1: Reinforcing Audio-visual Reasoning with Query Intention and Modality Attention

OmniVideo-R1: 通过查询意图和模态注意力强化音频视觉推理

Zhangquan Chen, Jiale Tao, Ruihuang Li, Yihao Hu, Ruitao Chen, Zhantao Yang, Xinlei Yu, Haodong Jing, Manyuan Zhang, Shuai Shao, Biao Wang, Qinglin Lu, Ruqi Huang

机构 * tencent(腾讯)

专题命中 音频语音多模态 :audio-visual(title,abstract);分类 cs.CV、cs.AI

AI总结 OmniVideo-R1通过查询意图和模态注意力机制,提升多模态推理能力,在多个基准上超越现有基线模型。

Comments 19 pages, 12 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.13263 2026-02-17 cs.CL cs.SD eess.AS 81%

Multimodal Consistency-Guided Reference-Free Data Selection for ASR Accent Adaptation

多模态一致性引导的参考自由数据选择用于ASR口音适应

Ligong Lei, Wenwen Lu, Xudong Pang, Zaokere Kadeer, Aishan Wumaier

机构 * School of Computer Science and Technology(计算机科学与技术学院) Xinjiang University(新疆大学)

专题命中 音频语音多模态 :multimodal(title,abstract);分类 cs.CL、eess.AS

AI总结 本文提出一种多模态一致性引导的参考自由数据选择方法,用于提升ASR在口音适应中的性能,通过减少噪声伪标签和优化查询相关性,实现更高效的口音适应。

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.13787 2026-02-17 cs.SD eess.AS 74%

Enhancing spatial hearing with cochlear implants: exploring the role of AI, multimodal interaction and perceptual training

通过 cochlear implants 增强空间听觉:探索人工智能、多模态交互和感知训练的作用

Lorenzo Picinali, Robert Baumgartner, Valerie Gaveau, Antonino Greco, Stefanie Liebe, Paul Oomen, Christoph Braun

机构 * Acoustics Research Institute, Austrian Academy of Sciences(奥地利科学院声学研究所) Centre de Recherche en Neurosciences de Lyon Inserm(里昂神经科学研究中心(Inserm)) Department of Neural Dynamics and Magnetoencephalography, Hertie Institute for Clinical Brain Research, University of Tübingen(图宾根大学神经动力学与脑磁图部门) Centre for Integrative Neuroscience, University of Tübingen(图宾根大学整合神经科学中心) NEMO Labs Nonprofit Kft., Budapest(布达佩斯NEMO实验室(非营利公司))

专题命中 音频语音多模态 :multimodal(title);分类 eess.AS

AI总结 本文提出多学科合作框架,通过人工智能、多模态交互和感知训练提升耳蜗植入物用户的空间听觉能力。

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.14655 2026-02-17 cs.CL cs.AI 73%

Breaking Data Efficiency Dilemma: A Federated and Augmented Learning Framework For Alzheimer's Disease Detection via Speech

突破数据效率困境:一种联邦学习与增强学习框架用于通过语音检测阿尔茨海默病

Xiao Wei, Bin Wen, Yuqin Lin, Kai Li, Mingyang gu, Xiaobao Wang, Longbiao Wang, Jianwu Dang

机构 * Tianjin Key Laboratory of Cognitive Computing and Application(认知计算与应用天津重点实验室) College of Intelligence and Computing, Tianjin University(智能计算学院,天津大学) Shenzhen Institutes of Advanced Technology, Chinese Academy of Sciences(深圳先进技术研究院,中国科学院) College of Computer and Data Science, Fuzhou University(计算机与数据科学学院,福州大学) Huiyan Technology (Tianjin) Co., Ltd(慧研科技(天津)有限公司)

专题命中 音频语音多模态 :multi-modal(abstract);cross-modal(abstract);分类 cs.CL、cs.AI

AI总结 FAL-AD通过联邦学习与数据增强框架,实现阿尔茨海默病语音检测中的数据效率提升,达到91.52%的多模态准确率。

Comments 5 pages, 1 figures, accepted by ICASSP 2026 conference

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.13954 2026-02-17 cs.SD cs.AI 70%

Eureka-Audio: Triggering Audio Intelligence in Compact Language Models

Eureka-Audio: 在紧凑语言模型中触发音频智能

Dan Zhang, Yishu Lei, Jing Hu, Shuwei He, Songhe Deng, Xianlong Luo, Danxiang Zhu, Shikun Feng, Rui Liu, Jingzhou He, Yu Sun, Hua Wu, Haifeng Wang

专题命中 音频语音多模态 :cross-modal(abstract);omni-modal(abstract);分类 cs.AI

AI总结 Eureka-Audio通过紧凑架构和统一端到端设计,在低参数量下实现高性能音频理解,匹配甚至超越大模型表现。

Comments 23 pages, 4 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.09424 2026-02-17 cs.CL cs.AI cs.LG eess.AS 67%

The Speech-LLM Takes It All: A Truly Fully End-to-End Spoken Dialogue State Tracking Approach

语音大语言模型一网打尽:一种真正端到端的语音对话状态跟踪方法

Nizar El Ghazal, Antoine Caubrière, Valentin Vielzeuf

专题命中 音频语音多模态 :multimodal(abstract);分类 cs.CL、cs.AI、eess.AS

AI总结 本文提出利用语音大语言模型实现端到端语音对话状态跟踪,通过比较不同上下文管理策略,发现完整语音对话输入能获得最佳性能,同时压缩历史信息能保持准确性与效率的平衡。

Comments Accepted for presentation at LREC 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.14127 2026-02-17 cs.SD 67%

MUKA: Multi Kernel Audio Adaptation Of Audio-Language Models

MUKA:多核音频适应音频-语言模型

Reda Bensaid, Amine Ouasfi, Yassir Bendou, Ilyass Moummad, Vincent Gripon, François Leduc-Primeau, Adnane Boukhayma

机构 * Inria, University Rennes, IRISA, CNRS(Inria、里昂大学、IRISA、CNRS)

专题命中 音频语音多模态 :multimodal(abstract);multimodal foundation model(abstract)

AI总结 MUKA通过结合细粒度表示与全局语义表示,实现音频-语言模型的少样本适应,展现了在适应性和效率上的平衡性能。

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.14224 2026-02-17 cs.SD cs.CL cs.MM 62%

The Interspeech 2026 Audio Reasoning Challenge: Evaluating Reasoning Process Quality for Audio Reasoning Models and Agents

Interspeech 2026音频推理挑战:评估音频推理模型和代理的推理过程质量

Ziyang Ma, Ruiyang Xu, Yinghao Ma, Chao-Han Huck Yang, Bohan Li, Jaeyeon Kim, Jin Xu, Jinyu Li, Carlos Busso, Kai Yu, Eng Siong Chng, Xie Chen

机构 * Shanghai Jiao Tong University(上海交通大学) Nanyang Technological University(南洋理工大学) Queen Mary University of London(伦敦大学Queen Mary) NVIDIA(NVIDIA公司) Carnegie Mellon University(卡内基梅隆大学) Qwen Team, Alibaba Group(通义实验室,阿里巴巴集团) Microsoft Corporation(微软公司)

专题命中 音频语音多模态 :cross-modal(abstract);分类 cs.CL、cs.MM

AI总结 Interspeech 2026音频推理挑战通过评估推理过程质量,探讨了音频推理模型和代理在事实性和逻辑性方面的表现及改进方向。

Comments The official website of the Audio Reasoning Challenge: https://audio-reasoning-challenge.github.io

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.13455 2026-02-17 cs.CL cs.AI cs.HC 62%

Using Machine Learning to Enhance the Detection of Obfuscated Abusive Words in Swahili: A Focus on Child Safety

利用机器学习增强斯瓦希里语中隐晦侮辱性词汇的检测:聚焦儿童安全

Phyllis Nabangi, Abdul-Jalil Zakaria, Jema David Ndibwile

专题命中 音频语音多模态 :multimodal(abstract);分类 cs.CL、cs.AI

AI总结 本研究利用机器学习方法提升斯瓦希里语中隐晦侮辱性词汇的检测能力,旨在提高儿童网络环境的安全性。

Comments Accepted at the Second IJCAI AI for Good Symposium in Africa, hosted by Deep Learning Indaba, 7 pages, 1 figure

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.13790 2026-02-17 cs.CL 57%

How Do Lexical Senses Correspond Between Spoken German and German Sign Language?

德语口语中的词义如何与德国手语对应?

Melis Çelikkol, Wei Zhao

机构 * Institute for Computational Linguistics, University of Heidelberg(计算语言学研究所,海德堡大学) Department of Computing Science, University of Aberdeen(计算机科学系,阿伯丁大学)

专题命中 音频语音多模态 :cross-modal(abstract);分类 cs.CL

AI总结 本研究通过分析德语和德国手语的词义对应关系,构建了首个跨模态词义对应标注数据集,利用语义相似性方法提高了词义映射的识别效果。

Comments EACL'26 (Student Research Workshop)

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.13330 2026-02-17 cs.CV 57%

Zwitscherkasten -- DIY Audiovisual bird monitoring

Zwitscherkasten -- DIY音频视觉鸟类监测

Dominik Blum, Elias Häring, Fabian Jirges, Martin Schäffer, David Schick, Florian Schulenberg, Torsten Schön

专题命中 音频语音多模态 :multimodal(abstract);分类 cs.CV

AI总结 Zwitscherkasten是一种利用边缘设备进行音频和视觉数据采集的DIY鸟类监测系统,通过深度学习实现实时无创的鸟类物种识别,支持生物多样性监测和公民科学应用。

Comments Project Report of the Applied Artificial Intelligence Degree Program at Technische Hochschule Ingolstadt

详情

展开后加载摘要…

URL PDF HTML 收藏

3. 视频多模态 13 篇

2602.10551 2026-02-17 cs.CV cs.AI 81%

C^2ROPE: Causal Continuous Rotary Positional Encoding for 3D Large Multimodal-Models Reasoning

C^2ROPE: 3D 大多模态模型推理中的因果连续旋转位置编码

Guanting Ye, Qiyan Zhao, Wenhao Yu, Xiaofeng Zhang, Jianmin Ji, Yanyong Zhang, Ka-Veng Yuen

机构 * State Key Laboratory of Internet of Things for Smart City, University of Macau(物联网智能城市国家重点实验室,澳门大学) Department of Automation, Shanghai Jiaotong University(上海交通大学自动化系) Institute of Advanced Technology, University of Science and Technology of China(中国科学技术大学先进技术研究院) School of Computer Science and Technology, USTC(中国科学技术大学计算机科学与技术学院) School of Artificial Intelligence and Data Science, USTC(中国科学技术大学人工智能与数据科学学院)

专题命中 视频多模态 :multimodal(title,abstract);分类 cs.CV、cs.AI

AI总结 C^2ROPE通过引入空间-时间连续位置编码和切比雪夫因果掩码,解决3D多模态模型中视觉特征连续性和因果关系建模问题。

Comments Accepted in ICRA 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.04641 2026-02-17 cs.CV cs.AI cs.LG 81%

Simulating the Real World: A Unified Survey of Multimodal Generative Models

模拟现实世界:多模态生成模型的统一综述

Yuqi Hu, Longguang Wang, Xian Liu, Ling-Hao Chen, Yuwei Guo, Yukai Shi, Ce Liu, Anyi Rao, Zeyu Wang, Hui Xiong

机构 * Thrust of Artificial Intelligence, The Hong Kong University of Science and Technology (Guangzhou)(人工智能前沿技术研究所,香港科学与技术大学(广州)) Department of Computer Science and Engineering, The Hong Kong University of Science and Technology Hong Kong SAR(计算机科学与工程系,香港科学与技术大学香港特别行政区) MMLab, The Hong Kong University of Science and Technology(多模态实验室,香港科学与技术大学) School of Electronics and Communication Engineering, Shenzhen Campus of Sun Yat-sen University(电子与通信工程学院,中山大学深圳校区) The Chinese University of Hong Kong, Hong Kong, China(香港中文大学,香港,中国) Tsinghua University, Guangdong, China(清华大学,广东,中国) Bosch (China) Investment Co., Ltd., Shanghai, China(博世(中国)投资有限公司,上海,中国)

专题命中 视频多模态 :multimodal(title,abstract);分类 cs.CV、cs.AI

AI总结 本文首次系统性地统一研究了2D、视频、3D和4D生成,为多模态生成模型和现实世界模拟提供了统一框架的综述。

Comments Repository for the related papers at https://github.com/ALEEEHU/World-Simulator

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.14589 2026-02-17 cs.AI cs.CL cs.LG 81%

MATEO: A Multimodal Benchmark for Temporal Reasoning and Planning in LVLMs

MATEO:一种多模态基准,用于LVLMs中的时间推理和规划

Gabriel Roccabruna, Olha Khomyn, Giuseppe Riccardi

机构 * Signals and Interactive Systems Lab, University of Trento, Italy(特伦托大学信号与交互系统实验室) University of Trento(特伦托大学) Amazon(亚马逊)

专题命中 视频多模态 :multimodal(title,abstract);分类 cs.CL、cs.AI

AI总结 MATEO是一个多模态基准,用于评估和提升大型视觉语言模型在时间推理和规划方面的能力。

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.21842 2026-02-17 cs.CV cs.CR 79%

Modal Aphasia: Can Unified Multimodal Models Describe Images From Memory?

模态失语:统一多模态模型能否从记忆中描述图像?

Michael Aerni, Joshua Swanson, Kristina Nikolić, Florian Tramèr

机构 * Michael Aerni(独立研究者) Joshua Swanson(独立研究者) Kristina Nikolić(独立研究者) Florian Tramèr(独立研究者)

专题命中 视频多模态 :multimodal(title,abstract);分类 cs.CV

AI总结 研究发现统一多模态模型在视觉记忆与文本表达间存在系统性缺陷,导致安全框架可能因单一模态防护而失效。

Comments Accepted to ICLR 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2110.12433 2026-02-17 cs.RO cs.SY eess.SY 78%

Model Predictive Control with Gaussian Processes for Flexible Multi-Modal Physical Human Robot Interaction

基于高斯过程的模型预测控制用于柔性多模态人机协作交互

Kevin Haninger, Christian Hegeler, Luka Peternel

专题命中 视频多模态 :multi-modal(title,abstract)

AI总结 本文提出基于高斯过程的模型预测控制方法,用于多模态人机协作交互,通过贝叶斯推断和在线控制提升任务灵活性和效率。

Comments Submitted, ICRA 2022. Video: https://youtu.be/0GT1pPpXvt8 Data and code: https://owncloud.fraunhofer.de/index.php/s/kmCZvlKOghclHy9

Journal ref 2022 IEEE International Conference on Robotics and Automation (ICRA), May 2022

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.13459 2026-02-17 eess.SP 78%

Towards Causality-Aware Modeling for Multimodal Brain-Muscle Interactions

迈向因果意识的多模态脑-肌相互作用建模

Farwa Abbas, Wei Dai, Zoran Cvetkovic, Verity McClelland

专题命中 视频多模态 :multimodal(title,abstract)

AI总结 本文提出一种结合几何流形重建与概率时间建模的DBN启发CCM框架,用于多模态脑-肌交互的因果建模,揭示了肌张力障碍中特定频率的通路重组织,并展示了其在生物标志物开发和神经调节干预中的应用潜力。

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.13329 2026-02-17 cs.CV cs.AI cs.RO 62%

HiST-VLA: A Hierarchical Spatio-Temporal Vision-Language-Action Model for End-to-End Autonomous Driving

HiST-VLA:一种用于端到端自动驾驶的分层时空视觉-语言-动作模型

Yiru Wang, Zichong Gu, Yu Gao, Anqing Jiang, Zhigang Sun, Shuo Wang, Yuwen Heng, Hao Sun

机构 * Bosch Corporate Research(博世企业研究) School of Communication and Information Engineering(信息工程学院)

专题命中 视频多模态 :multimodal(abstract);分类 cs.CV、cs.AI

AI总结 HiST-VLA通过分层时空视觉-语言-动作模型提升自动驾驶轨迹生成的精度与效率,实现端到端的自动驾驶系统。

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.23232 2026-02-17 cs.CV cs.AI 62%

ShotFinder: Imagination-Driven Open-Domain Video Shot Retrieval via Web Search

ShotFinder: 通过网络搜索驱动的开放域视频镜头检索

Tao Yu, Haopeng Jin, Hao Wang, Shenghua Chai, Yujia Yang, Junhao Gong, Jiaming Guo, Minghui Zhang, Xinlong Chen, Zhenghao Zhang, Yuxuan Zhou, Yufei Xiong, Shanbin Zhang, Jiabing Yang, Hongzhu Yi, Xinming Wang, Cheng Zhong, Xiao Ma, Zhang Zhang, Yan Huang, Liang Wang

机构 * CASIA(中国科学院自动化研究所) UCAS(中国科学院大学) Lenovo(联想集团) Peking University(北京大学) Tsinghua University(清华大学)

专题命中 视频多模态 :multimodal(abstract);分类 cs.CV、cs.AI

AI总结 ShotFinder通过网络搜索驱动开放域视频镜头检索,提出三阶段检索流程并揭示多模态大模型在开放域视频检索中的性能差距。

Comments 28 pages, 7 figures, Project website: https://github.com/yutao1024/ShotFinder

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.14653 2026-02-17 cs.CL 57%

Is Information Density Uniform when Utterances are Grounded on Perception and Discourse?

当话语基于感知和话语进行 grounding 时,信息密度是否均匀?

Matteo Gay, Coleman Haley, Mario Giulianelli, Edoardo Ponti

专题命中 视频多模态 :multimodal(abstract);分类 cs.CL

AI总结 本研究首次探讨了基于感知和话语的视觉环境对信息密度均匀性的影响,发现 grounding 能提高信息分布的均匀性,并在话语单元开始处产生最大的惊奇度降低。

Comments Accepted as main paper at EACL 2026

详情

展开后加载摘要…

URL PDF HTML 收藏