arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

多模态大模型

跨文本、图像、视频、音频等模态的大模型与学习方法。

2025-10-29 至 2025-10-29 共收录 9 信号源:cs.CV, cs.CL, cs.AI, cs.MM, eess.AS

1. 视频多模态 9 篇

2510.15870 2025-10-29 cs.CV cs.AI cs.CL 85%

OmniVinci: Enhancing Architecture and Data for Omni-Modal Understanding LLM

Hanrong Ye, Chao-Han Huck Yang, Arushi Goel, Wei Huang, Ligeng Zhu, Yuanhang Su, Sean Lin, An-Chieh Cheng, Zhen Wan, Jinchuan Tian, Yuming Lou, Dong Yang, Zhijian Liu, Yukang Chen, Ambrish Dantrey, Ehsan Jahangiri, Sreyan Ghosh, Daguang Xu, Ehsan Hosseini-Asl, Danial Mohseni Taheri, Vidya Murali, Sifei Liu, Yao Lu, Oluwatobi Olabiyi, Yu-Chiang Frank Wang, Rafael Valle, Bryan Catanzaro, Andrew Tao, Song Han, Jan Kautz, Hongxu Yin, Pavlo Molchanov

机构 * NVIDIA

专题命中 视频多模态 :omni-modal(title,abstract);cross-modal(abstract);分类 cs.CV、cs.CL、cs.AI

Comments Technical Report. Code: https://github.com/NVlabs/OmniVinci

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.17637 2025-10-29 cs.LG 82%

Causal Spatio-Temporal Prediction: An Effective and Efficient Multi-Modal Approach

Yuting Huang, Ziquan Fang, Zhihao Zeng, Lu Chen, Yunjun Gao

机构 * Zhejiang University(浙江大学)

专题命中 视频多模态 :multi-modal(title,abstract);cross-modal(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.23934 2025-10-29 cs.CY cs.AI cs.ET 79%

MFiSP: A Multimodal Fire Spread Prediction Framework

Alec Sathiyamoorthy, Wenhao Zhou, Xiangmin Zhou, Xiaodong Li, Iqbal Gondal

机构 * School of Computing Technologies, RMIT University(计算技术学院,皇家墨尔本理工大学)

专题命中 视频多模态 :multimodal(title,abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2501.03727 2025-10-29 eess.AS cs.LG 79%

Detecting Neurocognitive Disorders through Analyses of Topic Evolution and Cross-modal Consistency in Visual-Stimulated Narratives

Jinchao Li, Yuejiao Wang, Junan Li, Jiawen Kang, Bo Zheng, Ka Ho Wong, Brian Mak, Helene H. Fung, Jean Woo, Man-Wai Mak, Timothy Kwok, Vincent Mok, Xianmin Gong, Xixin Wu, Xunying Liu, Patrick C. M. Wong, Helen Meng

机构 * The Chinese University of Hong Kong(香港中文大学) The Hong Kong Polytechnic University(香港理工大学) The Hong Kong University of Science and Technology(香港科技大学)

专题命中 视频多模态 :cross-modal(title,abstract);分类 eess.AS

Comments 16 pages, 5 figures, accepted by "IEEE Journal of Selected Topics in Signal Processing"

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.24238 2025-10-29 cond-mat.mtrl-sci 78%

Unlocking Dynamic Luminescent Mapping of pH with Sustainable Lignin-Derived Carbon Dots with Multimodal Readout Capacity

Maja Szymczak, Jan Hočevar, Jernej Iskra, Darja Lisjak, Jelena Papan Djaniš, Lukasz Marciniak, Karolina Elzbieciak-Piecka

专题命中 视频多模态 :multimodal(title,abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.23630 2025-10-29 cs.LG cs.AI cs.CL 62%

NUM2EVENT: Interpretable Event Reasoning from Numerical time-series

Ninghui Feng, Yiyan Qi

机构 * International Digital Economy Academy (IDEA)(国际数字经济学院(IDEA)) University of Nottingham Ningbo(诺丁汉大学宁波校区)

专题命中 视频多模态 :multimodal(abstract);分类 cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.22092 2025-10-29 cs.AI 57%

VIRAL: Vision-grounded Integration for Reward design And Learning

Valentin Cuzin-Rambaud, Emilien Komlenovic, Alexandre Faure, Bruno Yun

机构 * Université Claude Bernard Lyon 1(克莱尔伯恩大学里昂1分校) CNRS(国家科学研究中心) Ecole Centrale de Lyon(里昂中央理工学院) INSA Lyon(里昂工业高等学院) Université Lumière Lyon 2(里昂2大学卢米埃尔分校) LIRIS(图像研究所)

专题命中 视频多模态 :multi-modal(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.24055 2025-10-29 cs.RO cs.LG 50%

Language-Conditioned Representations and Mixture-of-Experts Policy for Robust Multi-Task Robotic Manipulation

Xiucheng Zhang, Yang Jiang, Hongwei Qing, Jiashuo Bai

机构 * Xiucheng Zhang(未知) Yang Jiang(未知) Jiashuo Bai(未知)

专题命中 视频多模态 :multimodal(abstract)

Comments 8 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.05868 2025-10-29 cs.SI 50%

Detecting Coordinated Behaviour on Video-First Platforms: The Challenge of Multimodality and Complex Similarity on TikTok

Inga K. Wohlert, Davide Vega, Matteo Magnani, Alexandra Segerberg

专题命中 视频多模态 :multimodal(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏