arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

多模态大模型

跨文本、图像、视频、音频等模态的大模型与学习方法。

共收录 4729 信号源:cs.CV, cs.CL, cs.AI, cs.MM, eess.AS

1. 视频多模态 4729 篇

2508.06892 2025-08-12 astro-ph.SR physics.space-ph 50%

Large Model Driven Solar Activity AI Forecaster: A Scalable Dual Data-Model Framework

Jingjing Wang, Pengyu Liang, Tingyu Wang, Ming Li, Yanmei Cui, Siwei Liu, Xin Huang, Xiang Li, Minghui Zhang, Yunshi Zeng, Zhu Cao, Jiekang Feng, Qinghua Hu, Bingxian Luo, Bing Cao

专题命中 视频多模态 :multi-modal(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2412.10855 2025-08-08 cs.RO cs.LG 50%

Fast and Robust Visuomotor Riemannian Flow Matching Policy

Haoran Ding, Noémie Jaquier, Jan Peters, Leonel Rozo

机构 * Bosch Center for Artificial Intelligence(博世人工智能中心) Division of Robotics, Perception, and Learning, KTH Royal Institute of Technology(机器人、感知与学习 division,皇家理工学院) Computer Science Department of the Technische Universität Darmstadt(达姆施塔特技术大学计算机科学系)

专题命中 视频多模态 :multi-modal(abstract)

Comments Accepted for publication in IEEE T-RO. Project website: https://sites.google.com/view/rfmp 17 pages, 12 figures, 12 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.03518 2025-08-06 cs.IR cs.LG 50%

Parameter-Efficient Single Collaborative Branch for Recommendation

Marta Moscati, Shah Nawaz, Markus Schedl

机构 * Institute of Computational Perception, Johannes Kepler University Linz(计算感知研究所,林茨约瑟夫·施密特大学) AI Lab, Linz Institute of Technology(林茨技术学院人工智能实验室)

专题命中 视频多模态 :multimodal(abstract)

Comments 5 pages

Journal ref Proceedings of the Nineteenth ACM Conference on Recommender Systems (RecSys'25), September 22-26, 2025, Prague, Czech Republic. ACM, New York, NY, USA

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.22229 2025-07-31 cs.LG 50%

TRIBE: TRImodal Brain Encoder for whole-brain fMRI response prediction

Stéphane d'Ascoli, Jérémy Rapin, Yohann Benchetrit, Hubert Banville, Jean-Rémi King

机构 * Meta AI

专题命中 视频多模态 :multimodal(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.21016 2025-07-29 cs.LG q-bio.NC 50%

Predicting Cognition from fMRI:A Comparative Study of Graph, Transformer, and Kernel Models Across Task and Rest Conditions

Jagruti Patel, Mikkel Schöttner, Thomas A. W. Bolton, Patric Hagmann

专题命中 视频多模态 :multimodal(abstract)

Comments Preliminary version; a revised version will be uploaded later

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.20866 2025-07-29 physics.comp-ph 50%

Neuromorphic Photonic Processing and Memory with Spiking Resonant Tunnelling Diode Neurons and Neural Networks

Dafydd Owen-Newns, Joshua Robertson, Giovanni Donati, Jose Figueiredo, Edward Wasige, Kathy Ludge, Bruno Romeira, Antonio Hurtado

专题命中 视频多模态 :multi-modal(abstract)

Comments 19 pages, 11 figures, submitted to Advanced Intelligent Systems

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.19566 2025-07-29 eess.IV 50%

SLENet: A Novel Multiscale CNN-Based Network for Detecting the Rats Estrous Cycle

Qinyang Wang, Hoileong Lee, Xiaodi Pu, Yuanming Lai, Yiming Ma

专题命中 视频多模态 :multi-modal(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.19082 2025-07-28 cs.RO 50%

Bot Appétit! Exploring how Robot Morphology Shapes Perceived Affordances via a Mise en Place Scenario in a VR Kitchen

Rachel Ringe, Leandra Thiele, Mihai Pomarlan, Nima Zargham, Robin Nolte, Lars Hurrelbrink, Rainer Malaka

机构 * Digital Media Lab, University of Bremen(柏林布雷门大学数字媒体实验室) Department of Linguistics, University of Bremen(柏林布雷门大学语言学系)

专题命中 视频多模态 :multimodal(abstract)

Comments Copyright 2025 IEEE. Personal use of this material is permitted. Permission from IEEE must be obtained for all other uses, in any current or future media, including reprinting/republishing this material for advertising or promotional purposes, creating new collective works, for resale or redistribution to servers or lists, or reuse of any copyrighted component of this work in other works

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.18401 2025-07-25 cs.HC q-bio.NC 50%

Multisensory Integration and Sensory Substitution Across Vision, Audition, and Haptics: Answering the What, Which, and When in Study Protocols

Andrew Jeyathasan, Swati Banerjee

专题命中 视频多模态 :cross-modal(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.17795 2025-07-25 cs.LG 50%

LSDM: LLM-Enhanced Spatio-temporal Diffusion Model for Service-Level Mobile Traffic Prediction

Shiyuan Zhang, Tong Li, Zhu Xiao, Hongyang Du, Kaibin Huang

机构 * College of Computer Science and Electronic Engineering, Hunan University(计算机科学与电子工程学院,湖南大学) Department of Electrical and Electronic Engineering, University of Hong Kong(电气与电子工程系,香港大学)

专题命中 视频多模态 :multimodal(abstract)

Comments 14 pages, 9 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.17189 2025-07-24 cs.LG 50%

Met$^2$Net: A Decoupled Two-Stage Spatio-Temporal Forecasting Model for Complex Meteorological Systems

Shaohan Li, Hao Yang, Min Chen, Xiaolin Qin

机构 * Chengdu University of Information Technology(成都信息学院) Chengdu Institute of Computer Applications, Chinese Academy of Sciences(中国科学院成都计算机应用研究所) University of Chinese Academy of Sciences(中国科学院大学)

专题命中 视频多模态 :multimodal(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.13718 2025-07-21 cs.LG 50%

Bi-GRU Based Deception Detection using EEG Signals

Danilo Avola, Muhammad Yasir Bilal, Emad Emam, Cristina Lakasz, Daniele Pannone, Amedeo Ranaldi

机构 * Department of Computer Science, Sapienza University of Rome(计算机科学系,罗马萨皮恩扎大学)

专题命中 视频多模态 :multimodal(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2410.03993 2025-07-21 cs.HC 50%

TR-LLM: Integrating Trajectory Data for Scene-Aware LLM-Based Human Action Prediction

Kojiro Takeyama, Yimeng Liu, Misha Sra

专题命中 视频多模态 :multimodal(abstract)

Comments Accepted to IROS 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.13309 2025-07-18 cs.HC 50%

FocusView: Understanding and Customizing Informational Video Watching Experiences for Viewers with ADHD

Hanxiu 'Hazel' Zhu, Ruijia Chen, Yuhang Zhao

专题命中 视频多模态 :multimodal(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.18212 2025-07-17 cs.RO 50%

Haptic-Informed ACT with a Soft Gripper and Recovery-Informed Training for Pseudo Oocyte Manipulation

Pedro Miguel Uriguen Eljuri, Hironobu Shibata, Maeyama Katsuyoshi, Yuanyuan Jia, Tadahiro Taniguchi

机构 * Kyoto University(京都大学) Ritsumeikan University(立命馆大学)

专题命中 视频多模态 :multimodal(abstract)

Comments Accepted at IEEE/RSJ International Conference on Intelligent Robots and Systems (IROS2025) Project website https://tanichu-laboratory.github.io/pedro_haptic_act_iros2025/

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.09959 2025-07-15 cs.HC 50%

Branch Explorer: Leveraging Branching Narratives to Support Interactive 360° Video Viewing for Blind and Low Vision Users

Shuchang Xu, Xiaofu Jin, Wenshuo Zhang, Huamin Qu, Yukang Yan

专题命中 视频多模态 :multi-modal(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.08283 2025-07-14 cs.DB 50%

TableCopilot: A Table Assistant Empowered by Natural Language Conditional Table Discovery

Lingxi Cui, Guanyu Jiang, Huan Li, Ke Chen, Lidan Shou, Gang Chen

专题命中 视频多模态 :cross-modal(abstract)

Comments Accepted by VLDB'25

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.06247 2025-07-10 physics.flu-dyn physics.ins-det 50%

FED-PV: A Large-Scale Synthetic Frame/Event Dataset for Particle-Based Velocimetry

Fan Wu, Xiang Feng, Aoyu Zhang, Yong Lee

专题命中 视频多模态 :cross-modal(abstract)

Comments This work has been accepted as a conference paper at the 16th International Symposium on Particle Image Velocimetry (ISPIV 2025)

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.22473 2025-07-01 cs.RO eess.SP 50%

Unsupervised Discovery of Behavioral Primitives from Sensorimotor Dynamic Functional Connectivity

Fernando Diaz Ledezma, Valentin Marcel, Matej Hoffmann

机构 * Munich Institute of Robotics and Machine Intelligence, Technical University of Munich(慕尼黑机器人与机器智能研究所,慕尼黑技术大学) Faculty of Electrical Engineering, Czech Technical University in Prague(布拉格捷克技术大学电子工程学院)

专题命中 视频多模态 :multimodal(abstract)

Comments 8 pages with 6 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.20065 2025-06-26 cs.LG stat.AP 50%

Supervised Coupled Matrix-Tensor Factorization (SCMTF) for Computational Phenotyping of Patient Reported Outcomes in Ulcerative Colitis

Cristian Minoccheri, Sophia Tesic, Kayvan Najarian, Ryan Stidham

专题命中 视频多模态 :multi-modal(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.16257 2025-06-25 astro-ph.IM astro-ph.GA math.OC 50%

Multiobjective optimization for scattering mitigation and scattering screen reconstruction in VLBI observations of the Galactic Center

Alejandro Mus, Teresa Toscano, Hendrik Müller, Guang-Yao Zhao, Andrei Lobanov, Ciriaco Goddi

专题命中 视频多模态 :multi-modal(abstract)

Comments To appear in A&A

Journal ref A&A 698, A299 (2025)

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.18094 2025-06-24 eess.SY cs.SY 50%

G-SEED: A Spatio-temporal Encoding Framework for Forest and Grassland Data Based on GeoSOT

Xuan Ouyang, Xinwen Yu, Yan Chen, Guang Deng, Xuanxin Liu

专题命中 视频多模态 :multimodal(abstract)

Comments 11 pages, 2 figures. Previously submitted to a non-academic conference (ICGARSA 2025) and formally withdrawn

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.11634 2025-06-16 q-bio.NC 50%

Differences in Neurovascular Coupling in Patients with Major Depressive Disorder: Evidence from Simultaneous Resting-State EEG-fNIRS

Feng Yan, Xiaobin Wang, Yao Zhao, Shuyi Yang, Zhiren Wang

专题命中 视频多模态 :multimodal(abstract)

Comments 19 pages,9 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2409.10312 2025-06-03 cond-mat.dis-nn 50%

Soft modes in vector spin glass models on sparse random graphs

Silvio Franz, Cosimo Lupo, Flavio Nicoletti, Giorgio Parisi, Federico Ricci-Tersenghi

专题命中 视频多模态 :multi-modal(abstract)

Comments 15 pages, 8 figures

Journal ref Phys. Rev. B 111, 014203 (2025)

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.24309 2025-06-02 cs.SE cs.DC 50%

Supporting Long-term Transactions in Smart Contracts Generated from Business Process Model and Notation (BPMN) Models

Christian Gang Liu

专题命中 视频多模态 :multi-modal(abstract)

Comments Ph.D. Dissertation Ph.D. Dissertation

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.04445 2025-05-08 cs.IR 50%

M2Rec: Multi-scale Mamba for Efficient Sequential Recommendation

Qianru Zhang, Liang Qu, Honggang Wen, Dong Huang, Siu-Ming Yiu, Nguyen Quoc Viet Hung, Hongzhi Yin

专题命中 视频多模态 :multimodal(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.04237 2025-05-06 cs.IR 50%

Short Video Segment-level User Dynamic Interests Modeling in Personalized Recommendation

Zhiyu He, Zhixin Ling, Jiayu Li, Zhiqiang Guo, Weizhi Ma, Xinchen Luo, Min Zhang, Guorui Zhou

专题命中 视频多模态 :multi-modal(abstract)

Comments This paper has been accepted by SIGIR 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.15512 2025-04-29 cs.CR cs.LG 50%

T2VShield: Model-Agnostic Jailbreak Defense for Text-to-Video Models

Siyuan Liang, Jiayang Liu, Jiecheng Zhai, Tianmeng Fang, Rongcheng Tu, Aishan Liu, Xiaochun Cao, Dacheng Tao

机构 * Nanyang Technological University(南洋理工大学) Beijing Jiaotong University(北京交通大学) National University of Singapore(国立新加坡大学) Beihang University(北航) Sun Yat-sen University(中山大学)

专题命中 视频多模态 :multimodal(abstract)

Comments 33 pages, 9 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.18189 2025-04-28 cs.HC 50%

ClassComet: Exploring and Designing AI-generated Danmaku in Educational Videos to Enhance Online Learning

Zipeng Ji, Pengcheng An, Jian Zhao

专题命中 视频多模态 :multimodal(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.15561 2025-04-23 cs.RO cs.LG 50%

SPECI: Skill Prompts based Hierarchical Continual Imitation Learning for Robot Manipulation

Jingkai Xu, Xiangli Nie

机构 * State Key Laboratory of Multimodal Artificial Intelligence Systems, Institute of Automation, Chinese Academy of Sciences(多模态人工智能系统国家重点实验室,自动化研究所,中国科学院) School of Artificial Intelligence, University of Chinese Academy of Sciences(人工智能学院,中国科学院大学)

专题命中 视频多模态 :multimodal(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏