arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

AI 大模型

多模态大模型

跨文本、图像、视频、音频等模态的大模型与学习方法。

2025-09-23 至 2025-09-23 共收录 9 信号源:cs.CV, cs.CL, cs.AI, cs.MM, eess.AS

1. 视频多模态 9 篇

2506.08283 2025-09-23 cs.IR 82%

Serendipitous Recommendation with Multimodal LLM

Haoting Wang, Jianling Wang, Hao Li, Fangjun Yi, Mengyu Fu, Youwei Zhang, Yifan Liu, Liang Liu, Minmin Chen, Ed H. Chi, Lichan Hong, Haokai Lu

专题命中 视频多模态 :multimodal(title,abstract);MLLM(abstract)

Comments Accepted by 2025 Recsys EARL Workshop

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.17712 2025-09-23 cs.CV 79%

RCTDistill: Cross-Modal Knowledge Distillation Framework for Radar-Camera 3D Object Detection with Temporal Fusion

Geonho Bang, Minjae Seong, Jisong Kim, Geunju Baek, Daye Oh, Junhyung Kim, Junho Koh, Jun Won Choi

机构 * Seoul National University(首尔国立大学) Hanyang University(翰阳大学) Hyundai Motor Company(现代汽车公司)

专题命中 视频多模态 :cross-modal(title,abstract);分类 cs.CV

Comments Accepted at ICCV 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.15564 2025-09-23 cs.CV 79%

Show-o2: Improved Native Unified Multimodal Models

Jinheng Xie, Zhenheng Yang, Mike Zheng Shou

机构 * Show Lab, National University of Singapore(新加坡国立大学Show实验室) ByteDance(字节跳动)

专题命中 视频多模态 :multimodal(title,abstract);分类 cs.CV

Comments NeurIPS 2025. (v3: update to include video understanding, OneIG, and more ablation study results)

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.18020 2025-09-23 cs.HC 78%

ClassMind: Scaling Classroom Observation and Instructional Feedback with Multimodal AI

Ao Qu, Yuxi Wen, Jiayi Zhang, Yunge Wen, Yibo Zhao, Alok Prakash, Andrés F. Salazar-Gómez, Paul Pu Liang, Jinhua Zhao

专题命中 视频多模态 :multimodal(title,abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.17532 2025-09-23 cs.DC 78%

TACTFL: Temporal Contrastive Training for Multi-modal Federated Learning with Similarity-guided Model Aggregation

Guanxiong Sun, Majid Mirmehdi, Zahraa Abdallah, Raul Santos-Rodriguez, Ian Craddock, Telmo de Menezes e Silva Filho

专题命中 视频多模态 :multi-modal(title,abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.04603 2025-09-23 cs.RO 78%

Diffusion-Based Approximate MPC: Fast and Consistent Imitation of Multi-Modal Action Distributions

Pau Marquez Julbe, Julian Nubert, Henrik Hose, Sebastian Trimpe, Katherine J. Kuchenbecker

机构 * Max Planck ETH CLS(马克斯·普朗克-ETH CLS) German Research Foundation(德国研究基金会) Max Planck Institute for Intelligent Systems(马克斯·普朗克智能系统研究所) ETH Zürich(苏黎世联邦理工学院) Institute for Data Science in Mechanical Engineering (DSME)(机械工程数据科学研究所) RWTH Aachen University(亚琛工业大学)

专题命中 视频多模态 :multi-modal(title,abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.03319 2025-09-23 cs.CV cs.AI cs.MM 75%

SD-VSum: A Method and Dataset for Script-Driven Video Summarization

Manolis Mylonas, Evlampios Apostolidis, Vasileios Mezaris

机构 * ITI, CERTH(ITI、CERTH)

专题命中 视频多模态 :multimodal(abstract);cross-modal(abstract);分类 cs.CV、cs.AI、cs.MM

Comments In ACM Multimedia 2025, DOI:10.1145/3746027.3755821

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.17888 2025-09-23 cs.CV cs.AI 62%

Trainee Action Recognition through Interaction Analysis in CCATT Mixed-Reality Training

Divya Mereddy, Marcos Quinones-Grueiro, Ashwin T S, Eduardo Davalos, Gautam Biswas, Kent Etherton, Tyler Davis, Katelyn Kay, Jill Lear, Benjamin Goldberg

专题命中 视频多模态 :multimodal(abstract);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.08771 2025-09-23 cs.IR cs.AI 57%

Generate the browsing process for short-video recommendation

Chao Feng, Yanze Zhang, Chenghao Zhang

机构 * Kuaishou Technology(快手科技)

专题命中 视频多模态 :multimodal(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏