arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

多模态大模型

跨文本、图像、视频、音频等模态的大模型与学习方法。

共收录 4729 信号源:cs.CV, cs.CL, cs.AI, cs.MM, eess.AS

1. 视频多模态 4729 篇

2506.05868 2025-10-29 cs.SI 50%

Detecting Coordinated Behaviour on Video-First Platforms: The Challenge of Multimodality and Complex Similarity on TikTok

Inga K. Wohlert, Davide Vega, Matteo Magnani, Alexandra Segerberg

专题命中 视频多模态 :multimodal(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.22339 2025-10-28 cs.RO 50%

Estimating Continuum Robot Shape under External Loading using Spatiotemporal Neural Networks

Enyi Wang, Zhen Deng, Chuanchuan Pan, Bingwei He, Jianwei Zhang

机构 * Hamlyn Centre for Robotic Surgery, Institute of Global Health Innovation, Imperial College London(帝国理工学院伦敦校区全球健康创新研究所机器人手术中心) Department of Mechanical Engineering and Automation, Fuzhou University(福州大学机械工程与自动化学院) TAMS Group, Informatics, University of Hamburg(汉堡大学信息学院TAMS集团)

专题命中 视频多模态 :multi-modal(abstract)

Comments 2025 IEEE/RSJ International Conference on Intelligent Robots and Systems (IROS)

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.15745 2025-10-27 eess.IV cs.LG 50%

InfiniPot-V: Memory-Constrained KV Cache Compression for Streaming Video Understanding

Minsoo Kim, Kyuhong Shim, Jungwook Choi, Simyung Chang

专题命中 视频多模态 :multimodal(abstract)

Comments NeurIPS 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.22205 2025-10-22 cs.RO 50%

From Watch to Imagine: Steering Long-horizon Manipulation via Human Demonstration and Future Envisionment

Ke Ye, Jiaming Zhou, Yuanfeng Qiu, Jiayi Liu, Shihui Zhou, Kun-Yu Lin, Junwei Liang

机构 * The Hong Kong University of Science and Technology (Guangzhou)(香港科学与技术大学(广州)) The University of Hong Kong(香港大学) The Hong Kong University of Science and Technology(香港科学与技术大学)

专题命中 视频多模态 :multimodal(abstract)

Comments More details and videos can be found at: https://yipko.com/super-mimic

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.16980 2025-10-21 cs.LG 50%

Towards Interpretable and Trustworthy Time Series Reasoning: A BlueSky Vision

Kanghui Ning, Zijie Pan, Yushan Jiang, Anderson Schneider, Yuriy Nevmyvaka, Dongjin Song

机构 * School of Computing University of Connecticut Storrs, CT(计算学院 美国康涅狄格大学 斯托尔斯分校) Department of Machine Learning Research Morgan Stanley New York, NY(机器学习研究部 花旗集团 新 York)

专题命中 视频多模态 :multi-modal(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.19949 2025-10-21 eess.IV cs.LG 50%

Automated Video-EEG Analysis in Epilepsy Studies: Advances and Challenges

Valerii A. Zuev, Elena G. Salmagambetova, Stepan N. Djakov, Lev V. Utkin

机构 * Peter the Great St.Petersburg Polytechnic University(彼得大帝圣彼得堡理工大学)

专题命中 视频多模态 :multimodal(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.14286 2025-10-17 cs.LG 50%

Stable Prediction of Adverse Events in Medical Time-Series Data

Mayank Keoliya, Seewon Choi, Rajeev Alur, Mayur Naik, Eric Wong

机构 * University of Pennsylvania(宾夕法尼亚大学)

专题命中 视频多模态 :multi-modal(abstract)

Comments 18 pages, 3 Figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.13019 2025-10-16 physics.optics 50%

Phase Matching of Orbital Angular Momentum in Rare Earth Ion Doped Solid State Systems

Owen R. Wolfe, Joshua Dugre, Grant Kirkland, R. Krishna Mohan

专题命中 视频多模态 :multi-modal(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.10392 2025-10-14 cs.RO cs.SY eess.SY 50%

MicroRoboScope: A Portable and Integrated Mechatronic Platform for Magnetic and Acoustic Microrobotic Experimentation

Max Sokolich, Yanda Yang, Subrahmanyam Cherukumilli, Fatma Ceren Kirmizitas, Sambeeta Das

机构 * Department of Mechanical Engineering, University of Delaware(机械工程系,德雷克塞尔大学) Departments of Animal & Food Sciences, Biological Sciences, and Medical & Molecular Sciences, University of Delaware(动物与食品科学系、生物科学系和医学与分子科学系,德雷克塞尔大学)

专题命中 视频多模态 :multi-modal(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.06657 2025-10-09 cs.IR 50%

LLM-Powered Nuanced Video Attribute Annotation for Enhanced Recommendations

Boyuan Long, Yueqi Wang, Hiloni Mehta, Mick Zomnir, Omkar Pathak, Changping Meng, Ruolin Jia, Yajun Peng, Dapeng Hong, Xia Wu, Mingyan Gao, Onkar Dalal, Ningren Han

专题命中 视频多模态 :multimodal(abstract)

Comments RecSys 2025 Industry Track

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.05533 2025-10-08 q-fin.PM 50%

The New Quant: A Survey of Large Language Models in Financial Prediction and Trading

Weilong Fu

专题命中 视频多模态 :multimodal(abstract)

Comments 21 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.02702 2025-10-06 cs.CE cs.SI stat.ML 50%

VisitHGNN: Heterogeneous Graph Neural Networks for Modeling Point-of-Interest Visit Patterns

Lin Pang, Jidong J. Yang

专题命中 视频多模态 :multimodal(abstract)

Comments 16 pages, 9 figures, 5 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.04528 2025-10-02 cs.LG 50%

Federated Dynamic Modeling and Learning for Spatiotemporal Data Forecasting

Thien Pham, Angelo Furno, Faïcel Chamroukhi, Latifa Oukhellou

机构 * COSYS-GRETTIA, Gustave Eiffel University, 77420 France(COSYS-GRETTIA,巴黎-伊夫林大学) ENTPE, University of Lyon(ENTPE,里昂大学) the LICIT-ECO7 University Gustave Eiffel, France(LICIT-ECO7 巴黎-伊夫林大学)

专题命中 视频多模态 :multimodal(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.02371 2025-09-30 cond-mat.mtrl-sci cond-mat.mes-hall cond-mat.str-el 50%

Time- and Polarization-Resolved Extreme Ultraviolet Momentum Microscopy

Sotirios Fragkos, Quentin Courtade, Olena Tkach, Jérôme Gaudin, Dominique Descamps, Guillaume Barrette, Stéphane Petit, Gerd Schönhense, Yann Mairesse, Samuel Beaulieu

专题命中 视频多模态 :multimodal(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.23112 2025-09-30 cs.RO 50%

FTACT: Force Torque aware Action Chunking Transformer for Pick-and-Reorient Bottle Task

Ryo Watanabe, Maxime Alvarez, Pablo Ferreiro, Pavel Savkin, Genki Sano

机构 * TELEXISTENCE Inc, Foundation Model Division(TELEXISTENCE公司,基础模型部门) The University of Tokyo(东京大学)

专题命中 视频多模态 :multimodal(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2303.16829 2025-09-30 q-bio.NC cond-mat.stat-mech 50%

Neural spikes as rare events

Siddharth Kackar

专题命中 视频多模态 :multimodal(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.19892 2025-09-25 cs.RO 50%

D3Grasp: Diverse and Deformable Dexterous Grasping for General Objects

Keyu Wang, Bingcong Lu, Zhengxue Cheng, Hengdi Zhang, Li Song

机构 * Shanghai Jiao Tong University(上海交通大学)

专题命中 视频多模态 :multimodal(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.19169 2025-09-24 cs.RO 50%

MagiClaw: A Dual-Use, Vision-Based Soft Gripper for Bridging the Human Demonstration to Robotic Deployment Gap

Tianyu Wu, Xudong Han, Haoran Sun, Zishang Zhang, Bangchao Huang, Chaoyang Song, Fang Wan

机构 * Design + Learning Research Group(设计+学习研究组) Southern University of Science and Technology(南方科技大学)

专题命中 视频多模态 :multi-modal(abstract)

Comments 8 pages, 4 figures, accepted to Data@CoRL2025 Workshop

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.15481 2025-09-22 cs.LG cs.SI 50%

Solar Forecasting with Causality: A Graph-Transformer Approach to Spatiotemporal Dependencies

Yanan Niu, Demetri Psaltis, Christophe Moser, Luisa Lambertini

机构 * EPFL(苏黎世联邦理工学院)

专题命中 视频多模态 :multimodal(abstract)

Comments Accepted to CIKM 2025

Journal ref Proceedings of the 34th ACM International Conference on Information and Knowledge Management (CIKM '25), November 10--14, 2025, Seoul, Republic of Korea

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.11460 2025-09-19 q-bio.QM stat.AP 50%

Mechanistic inference of stochastic gene expression from structured single-cell data

Christopher E. Miles

专题命中 视频多模态 :multimodal(abstract)

Comments submitted invited review for the `Identifiability, estimation and uncertainty in mathematical modelling' issue in Current Opinion in Systems Biology

Journal ref Curr. Opin. Syst. Biol. 42, 100555 (2025)

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.11713 2025-09-16 cs.LG cs.NI 50%

Beyond Regularity: Modeling Chaotic Mobility Patterns for Next Location Prediction

Yuqian Wu, Yuhong Peng, Jiapeng Yu, Xiangyu Liu, Zeting Yan, Kang Lin, Weifeng Su, Bingqing Qu, Raymond Lee, Dingqi Yang

机构 * Beijing Normal-Hong Kong Baptist University(北京师范大学-香港 Baptist大学) University of Warwick(沃里克大学) University of Macau(澳门大学)

专题命中 视频多模态 :multimodal(abstract)

Comments 12 pages, 5 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.10864 2025-09-16 cs.LG 50%

CogGNN: Cognitive Graph Neural Networks in Generative Connectomics

Mayssa Soussia, Yijun Lin, Mohamed Ali Mahjoub, Islem Rekik

机构 * National Engineering School of Sousse, University of Sousse, LATIS- Laboratory of Advanced Technology and Intelligent Systems(突尼斯苏塞国立工程学校,苏塞大学,先进技术与智能系统实验室) BASIRA Lab, Imperial-X(BASIRA实验室,Imperial-X) Department of Computing, Imperial College London, UK(计算系,伦敦帝国学院,英国)

专题命中 视频多模态 :multimodal(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.10552 2025-09-16 q-bio.NC cs.LG 50%

Trial-Level Time-frequency EEG Desynchronization as a Neural Marker of Pain

D. A. Blanco-Mora, A. Dierolf, J. Gonçalves, M. van Der Meulen

机构 * Luxembourg Centre for Systems Biomedicine, University of Luxembourg(卢森堡系统生物医学研究中心,卢森堡大学) Department of Behavioural and Cognitive Sciences, University of Luxembourg(行为与认知科学系,卢森堡大学)

专题命中 视频多模态 :multimodal(abstract)

Comments 7 pages, 3 Figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.09889 2025-09-15 cs.RO cs.HC 50%

Using the Pepper Robot to Support Sign Language Communication

Giulia Botta, Marco Botta, Cristina Gena, Alessandro Mazzei, Massimo Donini, Alberto Lillo

专题命中 视频多模态 :multi-modal(abstract)

Comments paper presented at ICSR2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.02353 2025-09-03 cs.RO 50%

SAVOR: Skill Affordance Learning from Visuo-Haptic Perception for Robot-Assisted Bite Acquisition

Zhanxin Wu, Bo Ai, Tom Silver, Tapomayukh Bhattacharjee

机构 * Cornell University(康奈尔大学) UC San Diego(南加州大学)

专题命中 视频多模态 :multi-modal(abstract)

Comments Conference on Robot Learning, Oral

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.19958 2025-08-29 cs.RO 50%

Long-VLA: Unleashing Long-Horizon Capability of Vision Language Action Model for Robot Manipulation

Yiguo Fan, Pengxiang Ding, Shuanghao Bai, Xinyang Tong, Yuyang Zhu, Hongchao Lu, Fengqi Dai, Wei Zhao, Yang Liu, Siteng Huang, Zhaoxin Fan, Badong Chen, Donglin Wang

专题命中 视频多模态 :multimodal(abstract)

Comments Accepted to CoRL 2025; Github Page: https://long-vla.github.io

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.19916 2025-08-28 cond-mat.mtrl-sci 50%

Microscale optoelectronic reservoir networks of halide perovskite for in-sensor computing

Jeroen J. de Boer, Agustin O. Alvarez, Moritz C. Schmidt, Bruno Ehrler

专题命中 视频多模态 :multimodal(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.18725 2025-08-27 cs.NI cs.IT math.IT 50%

Toward Edge General Intelligence with Agentic AI and Agentification: Concepts, Technologies, and Future Directions

Ruichen Zhang, Guangyuan Liu, Yinqiu Liu, Changyuan Zhao, Jiacheng Wang, Yunting Xu, Dusit Niyato, Jiawen Kang, Yonghui Li, Shiwen Mao, Sumei Sun, Xuemin Shen, Dong In Kim

专题命中 视频多模态 :multimodal(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.14942 2025-08-22 cs.LG 50%

Structure-Aware Temporal Modeling for Chronic Disease Progression Prediction

Jiacheng Hu, Bo Zhang, Ting Xu, Haifeng Yang, Min Gao

专题命中 视频多模态 :multimodal(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.14432 2025-08-21 cs.LG 50%

Personalized Counterfactual Framework: Generating Potential Outcomes from Wearable Data

Ajan Subramanian, Amir M. Rahmani

机构 * Dept. of Computer Science, University of California, Irvine(加州大学伊市分校计算机科学系) School of Nursing, University of California, Irvine(加州大学伊市分校护理学院)

专题命中 视频多模态 :multi-modal(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏