arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

AI 大模型

多模态大模型

跨文本、图像、视频、音频等模态的大模型与学习方法。

2025-10-07 至 2025-10-07 共收录 15 信号源:cs.CV, cs.CL, cs.AI, cs.MM, eess.AS

1. 多模态生成 15 篇

2510.04765 2025-10-07 cs.AI 79%

LMM-Incentive: Large Multimodal Model-based Incentive Design for User-Generated Content in Web 3.0

Jinbo Wen, Jiawen Kang, Linfeng Zhang, Xiaoying Tang, Jianhang Tang, Yang Zhang, Zhaohui Yang, Dusit Niyato

机构 * Nanjing University of Aeronautics and Astronautics(南京航空航天大学) Guangdong University of Technology(广东工业大学) The Hong Kong Polytechnic University(香港理工大学) The Chinese University of Hong Kong(香港中文大学) Guizhou University(贵州大学) Zhejiang University(浙江大学) Nanyang Technological University(南洋理工大学)

专题命中 多模态生成 :multimodal(title,abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.03341 2025-10-07 cs.CV 70%

OpusAnimation: Code-Based Dynamic Chart Generation

Bozheng Li, Miao Yang, Zhenhan Chen, Jiawang Cao, Mushui Liu, Yi Lu, Yongliang Wu, Bin Zhang, Yangguang Ji, Licheng Tang, Jay Wu, Wenbo Zhu

机构 * Opus AI Research(Opus人工智能研究机构) Brown University(布朗大学) Zhejiang University(浙江大学) University of Toronto(多伦多大学)

专题命中 多模态生成 :multi-modal(abstract);MLLM(abstract);分类 cs.CV

Comments working in progress

详情

展开后加载摘要…

URL PDF HTML 收藏
2412.07730 2025-10-07 cs.CV cs.AI cs.LG cs.MM 67%

STIV: Scalable Text and Image Conditioned Video Generation

Zongyu Lin, Wei Liu, Chen Chen, Jiasen Lu, Wenze Hu, Tsu-Jui Fu, Jesse Allardice, Zhengfeng Lai, Liangchen Song, Bowen Zhang, Cha Chen, Yiran Fei, Lezhi Li, Yizhou Sun, Kai-Wei Chang, Yinfei Yang

机构 * Apple(苹果公司) University of California, Los Angeles(加州大学洛杉矶分校)

专题命中 多模态生成 :image-text(abstract);分类 cs.CV、cs.AI、cs.MM

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.04577 2025-10-07 cs.SD cs.LG cs.MM eess.AS 62%

Language Model Based Text-to-Audio Generation: Anti-Causally Aligned Collaborative Residual Transformers

Juncheng Wang, Chao Xu, Cheng Yu, Zhe Hu, Haoyu Xie, Guoqi Yu, Lei Shang, Shujun Wang

机构 * The Hong Kong Polytechnic University(香港理工大学) Alibaba Group(阿里巴巴集团)

专题命中 多模态生成 :multi-modal(abstract);分类 cs.MM、eess.AS

Comments Accepted to EMNLP 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.04498 2025-10-07 cs.CL cs.AI 62%

GenQuest: An LLM-based Text Adventure Game for Language Learners

Qiao Wang, Adnan Labib, Robert Swier, Michael Hofmeyr, Zheng Yuan

机构 * Hosei University(立命馆大学) King’s College London(伦敦大学国王学院) Kindai University(_kindai大学) Tokyo Uni. of Science(东京科学大学) University of Sheffield(谢菲尔德大学)

专题命中 多模态生成 :multi-modal(abstract);分类 cs.CL、cs.AI

Comments Workshop on Wordplay: When Language Meets Games, EMNLP 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.04201 2025-10-07 cs.CV cs.AI 62%

World-To-Image: Grounding Text-to-Image Generation with Agent-Driven World Knowledge

Moo Hyun Son, Jintaek Oh, Sun Bin Mun, Jaechul Roh, Sehyun Choi

机构 * The Hong Kong University of Science and Technology(香港科学与技术大学) Georgia Institute of Technology(佐治亚理工学院) University of Massachusetts Amherst(马萨诸塞大学阿默斯特分校) TwelveLabs

专题命中 多模态生成 :multimodal(abstract);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.24251 2025-10-07 cs.CV cs.CL 62%

Latent Visual Reasoning

Bangzheng Li, Ximeng Sun, Jiang Liu, Ze Wang, Jialian Wu, Xiaodong Yu, Hao Chen, Emad Barsoum, Muhao Chen, Zicheng Liu

机构 * University of California, Davis(加州大学戴维斯分校) Advanced Micro Devices, Inc.(先进微器件公司)

专题命中 多模态生成 :multimodal(abstract);分类 cs.CV、cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.16612 2025-10-07 cs.HC cs.AI cs.CV cs.CY 62%

Negative Shanshui: Real-time Interactive Ink Painting Synthesis

Aven-Le Zhou

机构 * The Hong Kong University of Science and Technology (Guangzhou)(香港理工大学(广州))

专题命中 多模态生成 :multimodal(abstract);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.05093 2025-10-07 cs.CV 57%

Character Mixing for Video Generation

Tingting Liao, Chongjian Ge, Guangyi Liu, Hao Li, Yi Zhou

机构 * Mohamed bin Zayed University of Artificial Intelligence(莫莫德·宾·扎耶德人工智能大学)

专题命中 多模态生成 :multimodal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.04615 2025-10-07 eess.SY cs.AI cs.SY 57%

Design Process of a Self Adaptive Smart Serious Games Ecosystem

X. Tao, P. Chen, M. Tsami, F. Khayati, M. Eckert

机构 * Research Center on Software Technologies and Multimedia Systems for Sustainability (CITSEM), Universidad Politécnica de Madrid (UPM), Spain(软件技术与多媒体系统可持续性研究所以及马德里理工大学)

专题命中 多模态生成 :multimodal(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.04125 2025-10-07 cs.CV 57%

Joint Learning of Pose Regression and Denoising Diffusion with Score Scaling Sampling for Category-level 6D Pose Estimation

Seunghyun Lee, Tae-Kyun Kim

机构 * KAIST(韩国科学技术院)

专题命中 多模态生成 :multi-modal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.03886 2025-10-07 cs.AI 57%

Rare Text Semantics Were Always There in Your Diffusion Transformer

Seil Kang, Woojung Han, Dayun Ju, Seong Jae Hwang

专题命中 多模态生成 :multi-modal(abstract);分类 cs.AI

Comments Accepted to NeurIPS 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.03591 2025-10-07 cs.CV 57%

Resolving Task Objective Conflicts in Unified Model via Task-Aware Mixture-of-Experts

Jiaxing Zhang, Hao Tang

机构 * Sichuan University(四川大学) Peking University(北京大学)

专题命中 多模态生成 :multimodal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.16876 2025-10-07 cs.SE 50%

Revolutionizing Validation and Verification: Explainable Testing Methodologies for Intelligent Automotive Decision-Making Systems

Halit Eris, Stefan Wagner

专题命中 多模态生成 :multimodal(abstract)

Comments Preprint to be published at SE4ADS

Journal ref 2025 IEEE/ACM 1st International Workshop on Software Engineering for Autonomous Driving Systems (SE4ADS), Ottawa, ON, Canada, 2025, pp. 34-37

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.03397 2025-10-07 hep-ph 50%

Foundation models for equation discovery in high energy physics

Manuel Morales-Alvarado

专题命中 多模态生成 :multimodal(abstract)

Comments 6 pages

详情

展开后加载摘要…

URL PDF HTML 收藏