arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

AI 大模型

多模态大模型

跨文本、图像、视频、音频等模态的大模型与学习方法。

2025-10-07 至 2025-10-07 共收录 98 信号源:cs.CV, cs.CL, cs.AI, cs.MM, eess.AS

1. 跨模态检索 7 篇

2510.04396 2025-10-07 cs.MM cs.IR 57%

Evaluating Keyframe Layouts for Visual Known-Item Search in Homogeneous Collections

Bastian Jäckl, Jiří Kruchina, Lucas Joos, Daniel A. Keim, Ladislav Peška, Jakub Lokoč

专题命中 跨模态检索 :multimodal(abstract);分类 cs.MM

Comments 28 Pages, 17 Figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.03614 2025-10-07 cs.LG cs.AI stat.ML 57%

Neural Bayesian Filtering

Christopher Solinas, Radovan Haluska, David Sychrovsky, Finbarr Timbers, Nolan Bard, Michael Buro, Martin Schmid, Nathan R. Sturtevant, Michael Bowling

机构 * University of Alberta(阿尔伯塔大学) Alberta Machine Intelligence Institute (Amii)(阿尔伯塔人工智能研究所) Charles University(查尔斯大学) EquiLibre Technologies, Inc.(EquiLibre技术公司) Allen Institute for AI(艾伦人工智能研究所) Sony AI(索尼人工智能)

专题命中 跨模态检索 :multimodal(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.02605 2025-10-07 cs.CV 57%

ReMoMask: Retrieval-Augmented Masked Motion Generation

Zhengdao Li, Siheng Wang, Zeyu Zhang, Hao Tang

专题命中 跨模态检索 :cross-modal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.01853 2025-10-07 cs.LG cs.LO 50%

Learning Representations Through Contrastive Neural Model Checking

Vladimir Krsmanovic, Matthias Cosler, Mohamed Ghanem, Bernd Finkbeiner

机构 * CISPA Helmholtz Center for Information Security(CISPA 欧洲信息安全研究中心)

专题命中 跨模态检索 :cross-modal(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏

2. 多模态生成 15 篇

2510.04765 2025-10-07 cs.AI 79%

LMM-Incentive: Large Multimodal Model-based Incentive Design for User-Generated Content in Web 3.0

Jinbo Wen, Jiawen Kang, Linfeng Zhang, Xiaoying Tang, Jianhang Tang, Yang Zhang, Zhaohui Yang, Dusit Niyato

机构 * Nanjing University of Aeronautics and Astronautics(南京航空航天大学) Guangdong University of Technology(广东工业大学) The Hong Kong Polytechnic University(香港理工大学) The Chinese University of Hong Kong(香港中文大学) Guizhou University(贵州大学) Zhejiang University(浙江大学) Nanyang Technological University(南洋理工大学)

专题命中 多模态生成 :multimodal(title,abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.03341 2025-10-07 cs.CV 70%

OpusAnimation: Code-Based Dynamic Chart Generation

Bozheng Li, Miao Yang, Zhenhan Chen, Jiawang Cao, Mushui Liu, Yi Lu, Yongliang Wu, Bin Zhang, Yangguang Ji, Licheng Tang, Jay Wu, Wenbo Zhu

机构 * Opus AI Research(Opus人工智能研究机构) Brown University(布朗大学) Zhejiang University(浙江大学) University of Toronto(多伦多大学)

专题命中 多模态生成 :multi-modal(abstract);MLLM(abstract);分类 cs.CV

Comments working in progress

详情

展开后加载摘要…

URL PDF HTML 收藏
2412.07730 2025-10-07 cs.CV cs.AI cs.LG cs.MM 67%

STIV: Scalable Text and Image Conditioned Video Generation

Zongyu Lin, Wei Liu, Chen Chen, Jiasen Lu, Wenze Hu, Tsu-Jui Fu, Jesse Allardice, Zhengfeng Lai, Liangchen Song, Bowen Zhang, Cha Chen, Yiran Fei, Lezhi Li, Yizhou Sun, Kai-Wei Chang, Yinfei Yang

机构 * Apple(苹果公司) University of California, Los Angeles(加州大学洛杉矶分校)

专题命中 多模态生成 :image-text(abstract);分类 cs.CV、cs.AI、cs.MM

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.04577 2025-10-07 cs.SD cs.LG cs.MM eess.AS 62%

Language Model Based Text-to-Audio Generation: Anti-Causally Aligned Collaborative Residual Transformers

Juncheng Wang, Chao Xu, Cheng Yu, Zhe Hu, Haoyu Xie, Guoqi Yu, Lei Shang, Shujun Wang

机构 * The Hong Kong Polytechnic University(香港理工大学) Alibaba Group(阿里巴巴集团)

专题命中 多模态生成 :multi-modal(abstract);分类 cs.MM、eess.AS

Comments Accepted to EMNLP 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.04498 2025-10-07 cs.CL cs.AI 62%

GenQuest: An LLM-based Text Adventure Game for Language Learners

Qiao Wang, Adnan Labib, Robert Swier, Michael Hofmeyr, Zheng Yuan

机构 * Hosei University(立命馆大学) King’s College London(伦敦大学国王学院) Kindai University(_kindai大学) Tokyo Uni. of Science(东京科学大学) University of Sheffield(谢菲尔德大学)

专题命中 多模态生成 :multi-modal(abstract);分类 cs.CL、cs.AI

Comments Workshop on Wordplay: When Language Meets Games, EMNLP 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.04201 2025-10-07 cs.CV cs.AI 62%

World-To-Image: Grounding Text-to-Image Generation with Agent-Driven World Knowledge

Moo Hyun Son, Jintaek Oh, Sun Bin Mun, Jaechul Roh, Sehyun Choi

机构 * The Hong Kong University of Science and Technology(香港科学与技术大学) Georgia Institute of Technology(佐治亚理工学院) University of Massachusetts Amherst(马萨诸塞大学阿默斯特分校) TwelveLabs

专题命中 多模态生成 :multimodal(abstract);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.24251 2025-10-07 cs.CV cs.CL 62%

Latent Visual Reasoning

Bangzheng Li, Ximeng Sun, Jiang Liu, Ze Wang, Jialian Wu, Xiaodong Yu, Hao Chen, Emad Barsoum, Muhao Chen, Zicheng Liu

机构 * University of California, Davis(加州大学戴维斯分校) Advanced Micro Devices, Inc.(先进微器件公司)

专题命中 多模态生成 :multimodal(abstract);分类 cs.CV、cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.16612 2025-10-07 cs.HC cs.AI cs.CV cs.CY 62%

Negative Shanshui: Real-time Interactive Ink Painting Synthesis

Aven-Le Zhou

机构 * The Hong Kong University of Science and Technology (Guangzhou)(香港理工大学(广州))

专题命中 多模态生成 :multimodal(abstract);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.05093 2025-10-07 cs.CV 57%

Character Mixing for Video Generation

Tingting Liao, Chongjian Ge, Guangyi Liu, Hao Li, Yi Zhou

机构 * Mohamed bin Zayed University of Artificial Intelligence(莫莫德·宾·扎耶德人工智能大学)

专题命中 多模态生成 :multimodal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.04615 2025-10-07 eess.SY cs.AI cs.SY 57%

Design Process of a Self Adaptive Smart Serious Games Ecosystem

X. Tao, P. Chen, M. Tsami, F. Khayati, M. Eckert

机构 * Research Center on Software Technologies and Multimedia Systems for Sustainability (CITSEM), Universidad Politécnica de Madrid (UPM), Spain(软件技术与多媒体系统可持续性研究所以及马德里理工大学)

专题命中 多模态生成 :multimodal(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.04125 2025-10-07 cs.CV 57%

Joint Learning of Pose Regression and Denoising Diffusion with Score Scaling Sampling for Category-level 6D Pose Estimation

Seunghyun Lee, Tae-Kyun Kim

机构 * KAIST(韩国科学技术院)

专题命中 多模态生成 :multi-modal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.03886 2025-10-07 cs.AI 57%

Rare Text Semantics Were Always There in Your Diffusion Transformer

Seil Kang, Woojung Han, Dayun Ju, Seong Jae Hwang

专题命中 多模态生成 :multi-modal(abstract);分类 cs.AI

Comments Accepted to NeurIPS 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.03591 2025-10-07 cs.CV 57%

Resolving Task Objective Conflicts in Unified Model via Task-Aware Mixture-of-Experts

Jiaxing Zhang, Hao Tang

机构 * Sichuan University(四川大学) Peking University(北京大学)

专题命中 多模态生成 :multimodal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.16876 2025-10-07 cs.SE 50%

Revolutionizing Validation and Verification: Explainable Testing Methodologies for Intelligent Automotive Decision-Making Systems

Halit Eris, Stefan Wagner

专题命中 多模态生成 :multimodal(abstract)

Comments Preprint to be published at SE4ADS

Journal ref 2025 IEEE/ACM 1st International Workshop on Software Engineering for Autonomous Driving Systems (SE4ADS), Ottawa, ON, Canada, 2025, pp. 34-37

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.03397 2025-10-07 hep-ph 50%

Foundation models for equation discovery in high energy physics

Manuel Morales-Alvarado

专题命中 多模态生成 :multimodal(abstract)

Comments 6 pages

详情

展开后加载摘要…

URL PDF HTML 收藏

3. 多模态评测 22 篇

2506.01713 2025-10-07 cs.CL 83%

SRPO: Enhancing Multimodal LLM Reasoning via Reflection-Aware Reinforcement Learning

Zhongwei Wan, Zhihao Dou, Che Liu, Yu Zhang, Dongfei Cui, Qinjian Zhao, Hui Shen, Jing Xiong, Yi Xin, Yifan Jiang, Chaofan Tao, Yangfan He, Mi Zhang, Shen Yan

机构 * The Ohio State University(俄亥俄州立大学)

专题命中 多模态评测 :multimodal(title,abstract);MLLM(abstract);分类 cs.CL

Comments NeurIPS 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.00883 2025-10-07 cs.CL 83%

Improve MLLM Benchmark Efficiency through Interview

Farong Wen, Yijin Guo, Junying Wang, Jiaohao Xiao, Yingjie Zhou, Ye Shen, Qi Jia, Chunyi Li, Zicheng Zhang

机构 * Shanghai Jiao Tong University(上海交通大学) Shanghai AI Laboratory(上海人工智能实验室)

专题命中 多模态评测 :MLLM(title,abstract);multimodal(abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.11625 2025-10-07 cs.CL cs.AI cs.CV cs.LG 82%

MapIQ: Evaluating Multimodal Large Language Models for Map Question Answering

Varun Srivastava, Fan Lei, Srija Mukhopadhyay, Vivek Gupta, Ross Maciejewski

机构 * School of Computing and Augmented Intelligence(计算与增强智能学院) Arizona State University(亚利桑那州立大学) Department of Computer Science(计算机科学系) International Institute of Information Technology(国际信息科技研究所)

专题命中 多模态评测 :multimodal(title,abstract);分类 cs.CV、cs.CL、cs.AI

Comments Published as a conference paper at COLM 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.03878 2025-10-07 cs.CV cs.AI 81%

Multi-Modal Oral Cancer Detection Using Weighted Ensemble Convolutional Neural Networks

Ajo Babu George, Sreehari J R Ajo Babu George, Sreehari J R Ajo Babu George, Sreehari J R

机构 * Dicemed

专题命中 多模态评测 :multi-modal(title);multimodal(abstract);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2402.07640 2025-10-07 cs.MM cs.AI 81%

Synthesizing Sentiment-Controlled Feedback For Multimodal Text and Image Data

Puneet Kumar, Sarthak Malik, Balasubramanian Raman, Xiaobai Li

机构 * Center for Machine Vision and Signal Analysis, University of Oulu(机器视觉与信号分析中心,奥卢大学) Indian Institute of Technology Roorkee(印度理工学院罗尔基分校) State Key Laboratory of Blockchain and Data Security, Zhejiang University(区块链与数据安全国家重点实验室,浙江大学)

专题命中 多模态评测 :multimodal(title,abstract);分类 cs.AI、cs.MM

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.07048 2025-10-07 cs.CV 80%

Comprehensive Evaluation of Large Multimodal Models for Nutrition Analysis: A New Benchmark Enriched with Contextual Metadata

Bruce Coburn, Jiangpeng He, Megan E. Rollo, Satvinder S. Dhaliwal, Deborah A. Kerr, Fengqing Zhu

机构 * Purdue University(普渡大学) Indiana University(印第安纳大学) Curtin University(Curtin大学)

专题命中 多模态评测 :multimodal(title,abstract);分类 cs.CV

Comments The extended full version of the accepted paper in 2025 IEEE BHI conference with title: Evaluating Large Multimodal Models for Nutrition Analysis: A New Benchmark Enriched with Contextual Metadata. Dataset is available at: https://skynet.ecn.purdue.edu/~coburn6/ACETADA/

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.04257 2025-10-07 cs.CR cs.AI 79%

AgentTypo: Adaptive Typographic Prompt Injection Attacks against Black-box Multimodal Agents

Yanjie Li, Yiming Cao, Dong Wang, Bin Xiao

机构 * Computing Department of Hong Kong Polytechnic University(香港理工大学计算机系) Computing Department, The Hong Kong Polytechnic University(香港理工大学计算机系)

专题命中 多模态评测 :multimodal(title,abstract);分类 cs.AI

Comments 13 pages, 8 figures. Submitted to IEEE Transactions on Information Forensics & Security

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.03965 2025-10-07 cs.MM 79%

FinCall-Surprise: A Large Scale Multi-modal Benchmark for Earning Surprise Prediction

Dong Shu, Yanguang Liu, Huopu Zhang, Mengnan Du

专题命中 多模态评测 :multi-modal(title,abstract);分类 cs.MM

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.12230 2025-10-07 cs.CV cs.AI cs.RO 76%

LIAM: Multimodal Transformer for Language Instructions, Images, Actions and Semantic Maps

Yihao Wang, Raphael Memmesheimer, Sven Behnke

机构 * Autonomous Intelligent Systems, Computer Science Institute VI, Center for Robotics, University of Bonn, Germany(自主智能系统,计算机科学研究所第六研究所,机器人中心,波恩大学,德国)

专题命中 多模态评测 :multimodal(title);分类 cs.CV、cs.AI

Comments 12 pages, 4 figures, 2 tables, 19th International Conference on Intelligent Autonomous Systems (IAS), Genoa, Italy, June 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.04438 2025-10-07 cs.CV cs.CL 62%

The Telephone Game: Evaluating Semantic Drift in Unified Models

Sabbir Mollah, Rohit Gupta, Sirnam Swetha, Qingyang Liu, Ahnaf Munir, Mubarak Shah

机构 * Center For Research in Computer Vision, University of Central Florida, USA(计算机视觉研究中心,中央佛罗里达大学)

专题命中 多模态评测 :cross-modal(abstract);分类 cs.CV、cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.03555 2025-10-07 cs.CV cs.AI 62%

GAS-MIL: Group-Aggregative Selection Multi-Instance Learning for Ensemble of Foundation Models in Digital Pathology Image Analysis

Peiran Quan, Zifan Gu, Zhuo Zhao, Qin Zhou, Donghan M. Yang, Ruichen Rong, Yang Xie, Guanghua Xiao

专题命中 多模态评测 :multimodal(abstract);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏