arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

AI 大模型

多模态大模型

跨文本、图像、视频、音频等模态的大模型与学习方法。

2025-10-03 至 2025-10-03 共收录 7 信号源:cs.CV, cs.CL, cs.AI, cs.MM, eess.AS

1. 多模态生成 7 篇

2510.01284 2025-10-03 cs.MM cs.CV cs.SD eess.AS 85%

Ovi: Twin Backbone Cross-Modal Fusion for Audio-Video Generation

Chetwin Low, Weimin Wang, Calder Katyal

机构 * Character AI Yale University(耶鲁大学)

专题命中 多模态生成 :cross-modal(title,abstract);multimodal(abstract);分类 cs.CV、cs.MM、eess.AS

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.00015 2025-10-03 cs.CL 79%

Design and Application of Multimodal Large Language Model Based System for End to End Automation of Accident Dataset Generation

MD Thamed Bin Zaman Chowdhury, Moazzem Hossain

专题命中 多模态生成 :multimodal(title,abstract);分类 cs.CL

Comments This paper is accepted for presentation in TRB annual meeting 2026. The version presented here is the preprint version before peer review process

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.01622 2025-10-03 cs.RO cs.LG 71%

VFP: Variational Flow-Matching Policy for Multi-Modal Robot Manipulation

Xuanran Zhai, Qianyou Zhao, Qiaojun Yu, Ce Hao

机构 * National University of Singapore(新加坡国立大学) Shanghai Jiao Tong University(上海交通大学) Shanghai AI Lab(上海人工智能实验室)

专题命中 多模态生成 :multi-modal(title)

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.23426 2025-10-03 cs.LG cs.AI 57%

Enhanced DACER Algorithm with High Diffusion Efficiency

Yinuo Wang, Likun Wang, Mining Tan, Wenjun Zou, Xujie Song, Wenxuan Wang, Tong Liu, Guojian Zhan, Tianze Zhu, Shiqi Liu, Zeyu He, Feihong Zhang, Jingliang Duan, Shengbo Eben Li

机构 * School of Vehicle and Mobility & College of AI, Tsinghua University(车辆与移动学院及人工智能学院,清华大学) School of Artificial Intelligence, University of Chinese Academy of Sciences(人工智能学院,中国科学院大学) School of Mechanical Engineering, University of Science and Technology Beijing(机械工程学院,北京科技大学)

专题命中 多模态生成 :multimodal(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2501.14399 2025-10-03 cs.IR cs.AI cs.DB cs.LG cs.SI 57%

Handling Heterophily in Recommender Systems with Wavelet Hypergraph Diffusion

Darnbi Sakong, Thanh Tam Nguyen

专题命中 多模态生成 :multi-modal(abstract);分类 cs.AI

Comments Fixed and extended results

详情

展开后加载摘要…

URL PDF HTML 收藏
2406.02842 2025-10-03 cs.CV 57%

DiffCut: Catalyzing Zero-Shot Semantic Segmentation with Diffusion Features and Recursive Normalized Cut

Paul Couairon, Mustafa Shukor, Jean-Emmanuel Haugeard, Matthieu Cord, Nicolas Thome

专题命中 多模态生成 :multimodal(abstract);分类 cs.CV

Comments NeurIPS 2024. Project page at https://diffcut-segmentation.github.io. Code at https://github.com/PaulCouairon/DiffCut

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.04584 2025-10-03 cs.HC 50%

SlideItRight: Using AI to Find Relevant Slides and Provide Feedback for Open-Ended Questions

Chloe Qianhui Zhao, Jie Cao, Eason Chen, Kenneth R. Koedinger, Jionghao Lin

专题命中 多模态生成 :multimodal(abstract)

Comments 14 pages, to be published at the 26th International Conference on Artificial Intelligence in Education (AIED '25)

详情

展开后加载摘要…

URL PDF HTML 收藏