arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

多模态大模型

跨文本、图像、视频、音频等模态的大模型与学习方法。

2025-10-23 至 2025-10-23 共收录 9 信号源:cs.CV, cs.CL, cs.AI, cs.MM, eess.AS

1. 多模态生成 9 篇

2510.18879 2025-10-23 cs.HC 78%

FIRETWIN: Digital Twin Advancing Multi-Modal Sensing, Interactive Analytics for Wildfire Response

Mayamin Hamid Raha, Ali Reza Tavakkoli, Chris Webb, Mobin Habibpour, Janice Coen, Eric Rowell, Fatemeh Afghah

专题命中 多模态生成 :multi-modal(title);multimodal(abstract)

Comments 8 pages, 6 figures, accepted in IEEE International Workshop on Computer-Aided Modeling and Design of Communication Links and Networks (CAMAD)

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.19808 2025-10-23 cs.CV cs.CL cs.LG 73%

Pico-Banana-400K: A Large-Scale Dataset for Text-Guided Image Editing

Yusu Qian, Eli Bocek-Rivele, Liangchen Song, Jialing Tong, Yinfei Yang, Jiasen Lu, Wenze Hu, Zhe Gan

机构 * Apple(苹果公司)

专题命中 多模态生成 :multimodal(abstract);MLLM(abstract);分类 cs.CV、cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.19641 2025-10-23 cs.CL cs.AI 62%

Style Attack Disguise: When Fonts Become a Camouflage for Adversarial Intent

Yangshijie Zhang, Xinda Wang, Jialin Liu, Wenqiang Wang, Zhicong Ma, Xingxing Jia

机构 * Lanzhou University(兰州大学) Peking University(北京大学) Sun Yat-sen University(中山大学)

专题命中 多模态生成 :multimodal(abstract);分类 cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.17519 2025-10-23 cs.CV cs.AI 62%

MUG-V 10B: High-efficiency Training Pipeline for Large Video Generation Models

Yongshun Zhang, Zhongyi Fan, Yonghang Zhang, Zhangzikang Li, Weifeng Chen, Zhongwei Feng, Chaoyue Wang, Peng Hou, Anxiang Zeng

机构 * LLM Team, Shopee Pte. Ltd.(Shopee 股份有限公司语言模型团队)

专题命中 多模态生成 :cross-modal(abstract);分类 cs.CV、cs.AI

Comments Technical Report; Project Page: https://github.com/Shopee-MUG/MUG-V

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.00939 2025-10-23 cs.CV cs.CL 62%

WikiVideo: Article Generation from Multiple Videos

Alexander Martin, Reno Kriz, William Gantt Walden, Kate Sanders, Hannah Recknor, Eugene Yang, Francis Ferraro, Benjamin Van Durme

机构 * Johns Hopkins University(约翰霍普金斯大学) Human Language Technology Center of Excellence(人机语言技术卓越中心) University of Maryland Baltimore County(马里兰大学巴尔的摩县分校)

专题命中 多模态生成 :multimodal(abstract);分类 cs.CV、cs.CL

Comments Repo can be found here: https://github.com/alexmartin1722/wikivideo

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.18911 2025-10-23 physics.chem-ph cs.AI 57%

Prospects for Using Artificial Intelligence to Understand Intrinsic Kinetics of Heterogeneous Catalytic Reactions

Andrew J. Medford, Todd N. Whittaker, Bjarne Kreitz, David W. Flaherty, John R. Kitchin

机构 * organization= School of Chemical \& Biomolecular Engineering, Georgia Institute of Technology , addressline= 311 Ferst Drive NW , city= Atlanta , postcode= 30332 , state= GA , country= USA organization= Department of Chemical Engineering, Carnegie Mellon University , addressline= 5000 Forbes Street , city= Pittsburgh , postcode= 15213 , state= PA , country= USA

专题命中 多模态生成 :multimodal(abstract);分类 cs.AI

Comments Submitted to "Current Opinion in Chemical Engineering" for peer review

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.12323 2025-10-23 cs.CV 57%

Doctor Approved: Generating Medically Accurate Skin Disease Images through AI-Expert Feedback

Janet Wang, Yunbei Zhang, Zhengming Ding, Jihun Hamm

机构 * Tulane University(路易斯安那州立大学)

专题命中 多模态生成 :multimodal(abstract);分类 cs.CV

Comments NeurIPS 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2410.05051 2025-10-23 cs.CV cs.RO 57%

ComDrive: Comfort-Oriented End-to-End Autonomous Driving

Junming Wang, Xingyu Zhang, Zebin Xing, Songen Gu, Xiaoyang Guo, Yang Hu, Ziying Song, Qian Zhang, Xiaoxiao Long, Wei Yin

机构 * Horizon Robotics University of Hong Kong(香港大学) University of the Chinese Academy of Sciences(中国科学院大学) Nanjing University(南京大学) Beijing Jiaotong University(北京交通大学)

专题命中 多模态生成 :multi-modal(abstract);分类 cs.CV

Comments IROS 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2501.12407 2025-10-23 cs.DC cs.LG 50%

The Streaming Batch Model for Efficient and Fault-Tolerant Heterogeneous Execution

Frank Sifei Luan, Ron Yifeng Wang, Yile Gu, Ziming Mao, Charlotte Lin, Amog Kamsetty, Hao Chen, Cheng Su, Balaji Veeramani, Scott Lee, SangBin Cho, Clark Zinzow, Eric Liang, Ion Stoica, Stephanie Wang

机构 * UC Berkeley(加州大学伯克利分校) University of Washington(华盛顿大学) Anyscale

专题命中 多模态生成 :multimodal(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏