arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

多模态大模型

跨文本、图像、视频、音频等模态的大模型与学习方法。

2026-01-22 至 2026-01-22 共收录 7 信号源:cs.CV, cs.CL, cs.AI, cs.MM, eess.AS

1. 视频多模态 7 篇

2601.14799 2026-01-22 cs.CV 83%

UBATrack: Spatio-Temporal State Space Model for General Multi-Modal Tracking

UBATrack: 一种用于通用多模态跟踪的时空状态空间模型

Qihua Liang, Liang Chen, Yaozong Zheng, Jian Nong, Zhiyi Mo, Bineng Zhong

机构 * Key Laboratory of Education Blockchain and Intelligent Technology, Ministry of Education, Guangxi Normal University(教育区块链与智能技术重点实验室,教育部,广西师范大学) Guangxi Key Laboratory of Multi-Source Information Mining and Security, Guangxi Normal University(广西多源信息挖掘与安全重点实验室,广西师范大学) University Engineering Research Center of Educational Intelligent Technology, Guangxi Normal University(教育智能技术大学工程研究中心,广西师范大学) Guangxi Key Laboratory of Machine Vision and Intelligent Control, Wuzhou University(广西机器视觉与智能控制重点实验室,梧州大学)

专题命中 视频多模态 :multi-modal(title,abstract);cross-modal(abstract);分类 cs.CV

AI总结 UBATrack提出基于mamba状态空间模型的多模态跟踪框架,通过时空Mamba适配器和动态多模态特征混合器提升跟踪鲁棒性,有效捕捉时空线索并提高训练效率。

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.14777 2026-01-22 cs.CV cs.AI 79%

FunCineForge: A Unified Dataset Toolkit and Model for Zero-Shot Movie Dubbing in Diverse Cinematic Scenes

FunCineForge: 一个统一的数据集工具包和模型,用于多样的影视场景零样本配音

Jiaxuan Liu, Yang Xiang, Han Zhao, Xiangang Li, Zhenhua Ling

机构 * Alibaba Group(阿里巴巴集团) Tongyi Lab(通义实验室)

专题命中 视频多模态 :multimodal(abstract);MLLM(abstract);audio-visual(abstract);分类 cs.CV、cs.AI

AI总结 FunCineForge提出了一种统一的数据集工具包和模型,用于多样的影视场景零样本配音,通过构建大规模数据集和基于大规模语言模型的模型,提升了音频质量、唇形同步和情绪表达性能。

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.14550 2026-01-22 cs.RO 78%

TacUMI: A Multi-Modal Universal Manipulation Interface for Contact-Rich Tasks

TacUMI: 一种多模态通用操控接口用于接触密集型任务

Tailai Cheng, Kejia Chen, Lingyun Chen, Liding Zhang, Yue Zhang, Yao Ling, Mahdi Hamad, Zhenshan Bing, Fan Wu, Karan Sharma, Alois Knoll

机构 * School of Computation, Information and Technology, Technical University of Munich(技术大学慕尼黑计算、信息与技术学院) Agile Robots SE(敏捷机器人公司) State Key Laboratory for Novel Software Technology and the School of Science and Technology, Nanjing University (Suzhou Campus)(南京大学软件新技术国家重点实验室及科学与技术学院(苏州校区)) Shanghai University(上海大学)

专题命中 视频多模态 :multi-modal(title,abstract)

AI总结 TacUMI通过整合多种传感器,提供了一种多模态数据采集系统,用于提高接触密集型任务中多模态演示的分割和收集效率。

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.22661 2026-01-22 cs.IR cs.AI 74%

Next Point-of-interest (POI) Recommendation Model Based on Multi-modal Spatio-temporal Context Feature Embedding

基于多模态时空上下文特征嵌入的下一个兴趣点(POI)推荐模型

Lingyu Zhang, Pengfei Xu, Rui Ban, Zhenchao Zhang, Songtao Liu, Yan Wang, Yunhai Wang

机构 * Institute of Trustworthy Autonomous Systems, Southern University of Science and Technology(SUSTech)(可信自主系统研究院,南方科技大学) School of Information Sciences and Technology, Northwest University(信息科学与技术学院,西北大学) China Information Technology Designing & Consulting Institute Co., Ltd.(中国信息科技设计与咨询研究院有限公司) Renmin University of China(中国人民大学)

专题命中 视频多模态 :multi-modal(title);分类 cs.AI

AI总结 本文提出基于多模态时空上下文特征嵌入的POI推荐模型,通过双流时空注意力机制有效区分长期习惯与短期意图,提升个性化出行预测性能。

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.14875 2026-01-22 cs.CV cs.AI 62%

GAT-NeRF: Geometry-Aware-Transformer Enhanced Neural Radiance Fields for High-Fidelity 4D Facial Avatars

GAT-NeRF:基于几何感知的Transformer增强神经辐射场用于高保真的4D面部虚拟人物

Zhe Chang, Haodong Jin, Ying Sun, Yan Song, Hui Yu

机构 * Department of Control Science and Engineering, University of Shanghai for Science and Technology(控制科学与工程系,上海科学技术大学) Business School, University of Shanghai for Science and Technology(商学院,上海科学技术大学) School of Psychology and Neuroscience, University of Glasgow(心理学与神经科学学院,格拉斯哥大学)

专题命中 视频多模态 :multi-modal(abstract);分类 cs.CV、cs.AI

AI总结 GAT-NeRF通过整合Transformer机制与几何感知模块,提升从单目视频重建高保真4D动态面部虚拟人物的能力。

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.15115 2026-01-22 cs.CV 57%

Training-Free and Interpretable Hateful Video Detection via Multi-stage Adversarial Reasoning

无需训练的多阶段对抗推理 hateful 视频检测

Shuonan Yang, Yuchen Zhang, Zeyu Fu

机构 * Multimodal Intelligence Lab, Department of Computer Science, University of Exeter, United Kingdom(埃克塞特大学计算机科学系多模态智能实验室) Institute for Analytics and Data Science, University of Essex, United Kingdom(埃塞克斯大学分析与数据科学研究所)

专题命中 视频多模态 :multimodal(abstract);分类 cs.CV

AI总结 MARS通过多阶段对抗推理框架实现无需训练的可解释仇恨视频检测,提升检测可靠性与透明度。

Comments Accepted at ICASSP 2026. \c{opyright} 2026 IEEE. This is the author accepted manuscript. The final published version will be available via IEEE Xplore

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.14917 2026-01-22 cs.LG cs.AI 57%

Tailoring Adverse Event Prediction in Type 1 Diabetes with Patient-Specific Deep Learning Models

为1型糖尿病定制不良事件预测的患者特异性深度学习模型

Giorgia Rigamonti, Mirko Paolo Barbato, Davide Marelli, Paolo Napoletano

机构 * Department of Informatics, Systems and Communication(信息学、系统与通信系)

专题命中 视频多模态 :multimodal(abstract);分类 cs.AI

AI总结 本文提出一种基于深度学习的个性化血糖预测方法,利用患者特定数据提高预测准确性,以改善1型糖尿病的管理。

详情

展开后加载摘要…

URL PDF HTML 收藏