arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

多模态大模型

跨文本、图像、视频、音频等模态的大模型与学习方法。

共收录 4979 信号源:cs.CV, cs.CL, cs.AI, cs.MM, eess.AS

1. 多模态生成 4979 篇

2605.29219 2026-06-05 cs.CV 74%

SalsaAgent: A multimodal embodied language model for interactive dance generation

SalsaAgent: 一种用于交互式舞蹈生成的多模态具身语言模型

Payam Jome Yazdian, Zoe Stanley, Angelica Lim

机构 * Simon Fraser University(西蒙弗雷泽大学)

专题命中 多模态生成 :multimodal(title);分类 cs.CV

AI总结 提出SalsaAgent语言模型,通过非语言运动令牌传递和两阶段令牌到扩散管道,生成与人类领舞者及音乐背景交互的全身萨尔萨舞蹈动作。

Comments Project page: https://pjyazdian.github.io/Salsa-Agent

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.03322 2026-06-03 cs.LG cs.AI 74%

Multi-Modal Graph Neural Network with Transformer-Guided Adaptive Diffusion for Preclinical Alzheimer Classification

多模态图神经网络与Transformer引导的自适应扩散用于临床前阿尔茨海默病分类

Jaeyoon Sim, Minjae Lee, Guorong Wu, Won Hwa Kim

机构 * Pohang University of Science and Technology(浦项科学技术大学) University of North Carolina at Chapel Hill(北卡罗来纳大学教堂山分校)

专题命中 多模态生成 :multi-modal(title);分类 cs.AI

AI总结 提出一种结合扩散核与多头注意力的图神经网络框架,通过Transformer引导自适应扩散过程,有效融合多模态特征,提升临床前阿尔茨海默病分类性能并识别关键脑区。

Comments 10 pages, Accepted to MICCAI 2024

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.20861 2026-06-02 cs.IR cs.AI 74%

Deep Interest Mining for Intent-Enriched Semantic IDs in Multimodal Generative Recommendation

面向多模态生成式推荐中意图增强语义ID的深度兴趣挖掘

Yangchen Zeng, Jinze Wang

机构 * Amazon(亚马逊)

专题命中 多模态生成 :multimodal(title);分类 cs.AI

AI总结 提出DeepInterestGR框架,通过视觉线索和意图描述符丰富量化前的物品表示,并结合相关性门控语义奖励,提升基于语义ID的生成式推荐性能。

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.31093 2026-06-01 cs.CV 74%

Cross-Modal Clinical Knowledge Integration for Mammography Report Generation

跨模态临床知识整合用于乳腺X线报告生成

Jiayi Zhu, Fuxiang Huang, Yu Xie, Xi Wang, Zhixuan Chen, Yuan Guo, Qingcong Kong, Zhenhui Li, Qiong Luo, Hao Chen

机构 * The Hong Kong University of Science and Technology (Guangzhou)(香港科技大学(广州)) Lingnan University(岭南大学) The Third Affiliated Hospital of Kunming Medical University, Yunnan Cancer Hospital, Peking University Cancer Hospital Yunnan(昆明医科大学第三附属医院、云南癌症医院、北京大学肿瘤医院云南分院) Guangzhou First People's Hospital, South China University of Technology(广州第一人民医院、华南理工大学) The Third Affiliated Hospital, Sun Yat-Sen University(中山大学第三附属医院) HKUST Shenzhen-Hong Kong Collaborative Innovation Research Institute(香港科技大学深圳-香港协同创新研究院)

专题命中 多模态生成 :cross-modal(title);分类 cs.CV

AI总结 提出MammoRG框架,通过两阶段训练模拟临床报告流程,整合BI-RADS指南和先验知识,提升报告生成的临床一致性。

Comments 16 pages, 5 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.16702 2026-04-28 cs.CV 74%

Animalbooth: multimodal feature enhancement for animal subject personalization

Animalbooth:多模态特征增强用于动物主体个性化

Chen Liu, Haitao Wu, Kafeng Wang, Weiran Huang

机构 * College of Intelligence and Computing(智能与计算学院) Department of Computer Science and Technology(计算机科学与技术系) School of Computer Science(计算机科学学院)

专题命中 多模态生成 :multimodal(title);分类 cs.CV

AI总结 针对动物图像生成中丰富的外观线索和形态多样性带来的挑战,AnimalBooth通过Animal Net和自适应注意力模块增强身份保持,结合频域控制特征整合模块提升扩散过程的全局结构到细节纹理的逐步生成,同时构建AnimalBench数据集验证了其在多个基准上的优越性能。

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.16170 2026-04-20 cs.CV cs.CE 74%

neuralCAD-Edit: An Expert Benchmark for Multimodal-Instructed 3D CAD Model Editing

neuralCAD-Edit:一个多模态指令引导的3D CAD模型编辑专家基准

Toby Perrett, Matthew Bouchard, William McCarthy

机构 * Autodesk Research(Autodesk研究院)

专题命中 多模态生成 :multimodal(title);分类 cs.CV

AI总结 本文提出neuralCAD-Edit,首个基于专家CAD工程师采集的3D CAD模型编辑基准。通过捕捉专业设计师与CAD模型交互的视频,收集真实编辑请求,发现基础模型在自动指标和人工评估中均显著落后于CAD专家。

Comments Project page: https://autodeskailab.github.io/neuralCAD-Edit

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.20340 2026-04-16 cs.SE cs.AI 74%

ContractSkill: Repairable Contract-Based Skills for Multimodal Web Agents

ContractSkill:可修复的基于合同的多模态网络代理技能

Zijian Lu, Yiping Zuo, Yupeng Nie, Xin He, Weibei Fan, Lianyong Qi, Shi Jin

机构 * Nanjing University of Posts and Telecommunications(南京邮电大学) China University of Petroleum (East China)(中国石油大学(华东)) Southeast University(东南大学)

专题命中 多模态生成 :multimodal(title);分类 cs.AI

AI总结 本文提出ContractSkill框架,通过将草案技能转换为可执行的显式过程结构,解决网络代理技能隐含导致无法检查和局部修复的问题,实验证明其在真实网络环境中的有效性。

Comments 10 pages, 4 figures, 6 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.08922 2026-04-13 cs.CV 74%

Degradation-Robust Fusion: An Efficient Degradation-Aware Diffusion Framework for Multimodal Image Fusion in Arbitrary Degradation Scenarios

抗退化融合:一种高效的抗退化扩散框架,用于任意退化场景下的多模态图像融合

Yu Shi, Yu Liu, Zhong-Cheng Wu, Juan Cheng, Huafeng Li, Xun Chen

机构 * Department of Biomedical Engineering, Hefei University of Technology(合肥工业大学生物医学工程系) Faculty of Information Engineering and Automation, Kunming University of Science and Technology(昆明理工大学信息工程与自动化学院) School of Information Science and Technology, University of Science and Technology of China(中国科学技术大学信息科学技术学院)

专题命中 多模态生成 :multimodal(title);分类 cs.CV

AI总结 本文提出一种高效的抗退化扩散框架,用于处理多模态图像融合中的复杂退化场景。该框架通过隐式去噪和联合观测模型校正机制,提升了在复杂退化下的融合性能。

Comments Accepted by CVPR 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2410.22177 2026-04-09 cs.HC cs.AI 74%

Analyzing Multimodal Interaction Strategies for LLM-Assisted Manipulation of 3D Scenes

分析用于LLM辅助3D场景操作的多模态交互策略

Junlong Chen, Jens Grubert, Per Ola Kristensson

机构 * University of Cambridge(剑桥大学) Coburg University of Applied Sciences(科堡应用科学大学)

专题命中 多模态生成 :multimodal(title);分类 cs.AI

AI总结 本文通过实证研究揭示LLM辅助3D场景编辑系统中的交互模式与关键障碍,提出改进自然语言界面的设计建议,证明LLM辅助交互系统在沉浸环境中的有效性。

Comments Published in the IEEE VR (IEEE Conference on Virtual Reality and 3D User Interfaces) 2025 Proceedings

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.27666 2026-03-31 cs.CV 74%

Gated Condition Injection without Multimodal Attention: Towards Controllable Linear-Attention Transformers

门控条件注入无需多模态注意:迈向可控的线性注意变换器

Yuhe Liu, Zhenxiong Tan, Yujia Hu, Songhua Liu, Xinchao Wang

机构 * National University of Singapore(新加坡国立大学) Shanghai Jiao Tong University(上海交通大学)

专题命中 多模态生成 :multimodal(title);分类 cs.CV

AI总结 本文提出基于线性注意架构的可控扩散模型,解决现有方法在多类型条件输入灵活性和收敛速度上的不足,通过统一门控条件模块实现高效可控生成。

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.08021 2026-03-31 cs.RO cs.CV 74%

AffordGrasp: Cross-Modal Diffusion for Affordance-Aware Grasp Synthesis

AffordGrasp:跨模态扩散用于感知意识抓取合成

Xiaofei Wu, Yi Zhang, Yumeng Liu, Yuexin Ma, Yujiao Shi, Xuming He

机构 * ShanghaiTech University(上海科技大学) Shanghai Engineering Research Center of Intelligent Vision and Imaging(上海智能视觉与成像工程技术研究中心) University of Science and Technology of China(中国科学技术大学)

专题命中 多模态生成 :cross-modal(title);分类 cs.CV

AI总结 本文提出AffordGrasp,通过跨模态扩散模型生成准确反映物体几何和用户指令的抓取姿态,提升AR/VR和具身AI中的手-物交互质量。

Comments CVPR 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.20060 2026-03-27 cs.CV cs.RO 74%

MeanFuser: Fast One-Step Multi-Modal Trajectory Generation and Adaptive Reconstruction via MeanFlow for End-to-End Autonomous Driving

MeanFuser: 一种基于MeanFlow的高效多模态轨迹生成与自适应重构方法用于端到端自动驾驶

Junli Wang, Yinan Zheng, Xueyi Liu, Zebin Xing, Pengfei Li, Guang Li, Kun Ma, Guang Chen, Hangjun Ye, Zhongpu Xia, Long Chen, Qichao Zhang

机构 * SKL-MAIS, Institute of Automation, Chinese Academy of Sciences(中国科学院自动化研究所模式识别国家重点实验室) School of Artificial Intelligence, University of Chinese Academy of Sciences(中国科学院大学人工智能学院) Xiaomi EV(小米汽车) Institute for AI Industry Research (AIR), Tsinghua University(清华大学智能产业研究院)

专题命中 多模态生成 :multi-modal(title);分类 cs.CV

AI总结 本文提出MeanFuser,通过引入Gaussian Mixture Noise、MeanFlow Identity和轻量级ARM模块,实现了高效且鲁棒的多模态轨迹生成与自适应重构,提升了端到端自动驾驶的性能和效率。

Comments Accepted by CVPR 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.08998 2026-03-11 cs.CV 74%

Diffusion-Based Authentication of Copy Detection Patterns: A Multimodal Framework with Printer Signature Conditioning

基于扩散的复制检测模式认证:一种多模态框架与打印机签名条件化

Bolutife Atoki, Iuliia Tkachenko, Bertrand Kerautret, Carlos Crispim-Junior

机构 * Université Lumière Lyon 2, CNRS, INSA Lyon, Universite Claude Bernard Lyon 1, LIRIS UMR5205(里摩日大学里昂2分校、法国国家科学研究中心、里昂国立应用科学学院、里昂大学克莱尔-贝尔纳分校、LIRIS UMR5205)

专题命中 多模态生成 :multimodal(title);分类 cs.CV

AI总结 本文提出基于扩散的多模态认证框架,利用打印机签名条件化技术,实现对复制检测模式的高效认证,优于传统方法和深度学习方法。

Comments Accepted at WACV 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.18903 2026-02-24 cs.CV cs.HC 74%

SCHEMA for Gemini 3 Pro Image: A Structured Methodology for Controlled AI Image Generation on Google's Native Multimodal Model

Gemini 3 Pro图像的SCHEMA:为Google原生多模态模型的受控AI图像生成的结构化方法

Luca Cazzaniga

机构 * Independent Researcher(独立研究者)

专题命中 多模态生成 :multimodal(title);分类 cs.CV

AI总结 SCHEMA为Google Gemini 3 Pro图像提供结构化提示工程方法,通过三级系统提升AI图像生成的可控性,实现高合规率和跨领域应用。

Comments 24 pages, 8 tables. Based on SCHEMA Method v1.0 (deposited December 11, 2025). Previously published on Zenodo: doi:10.5281/zenodo.18721380

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.23208 2026-02-03 cs.CL 74%

A Structured Framework for Evaluating and Enhancing Interpretive Capabilities of Multimodal LLMs in Culturally Situated Tasks

一种评估和增强多模态大语言模型在文化情境任务中解释能力的结构框架

Haorui Yu, Ramon Ruiz-Dolz, Qiufeng Yi

机构 * DJCAD, University of Dundee, United Kingdom(邓迪大学DJCAD部门) ARG-tech, SSEN, University of Dundee, United Kingdom(邓迪大学) School of Computer Science, University of Birmingham, United Kingdom(伯明翰大学计算机科学学院)

专题命中 多模态生成 :multimodal(title);分类 cs.CL

AI总结 本研究提出了一种结构框架,用于评估和增强多模态大语言模型在文化情境任务中生成中国绘画批评的能力,通过量化评价特征和人设引导提示,揭示了VLMs在艺术批评领域的表现与局限。

Comments EMNLP 2025 submission, 10 pages, 6 figures, 5 tables

Journal ref Findings of the Association for Computational Linguistics: EMNLP 2025, pages 1945-1971, Suzhou, China

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.19774 2026-01-22 cs.LG cs.AI eess.SP 74%

PPGFlowECG: Latent Rectified Flow with Cross-Modal Encoding for PPG-Guided ECG Generation and Cardiovascular Disease Detection

PPGFlowECG: 基于跨模态编码的潜在修正流用于PPG引导的ECG生成和心血管疾病检测

Xiaocheng Fang, Jiarui Jin, Haoyu Wang, Che Liu, Jieyi Cai, Yujie Xiao, Guangkun Nie, Bo Liu, Shun Huang, Hongyan Li, Shenda Hong

机构 * National Institute of Health Data Science, Peking University, China(北京大学国家健康数据科学研究院) School of Intelligence Science and Technology, Peking University, China(北京大学智能科学与技术学院) Data Science Institute, Imperial College London, UK(伦敦帝国理工学院数据科学研究院) University of Chinese Academy of Sciences, China(中国科学院大学)

专题命中 多模态生成 :cross-modal(title);分类 cs.AI

AI总结 PPGFlowECG通过跨模态编码和潜在修正流,实现PPG引导的ECG生成,提升心血管疾病检测的可扩展性和可靠性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.04506 2026-01-09 cs.LG cs.AI cs.CE 74%

Surface-based Molecular Design with Multi-modal Flow Matching

基于表面的分子设计与多模态流匹配

Fang Wu, Zhengyuan Zhou, Shuting Jin, Xiangxiang Zeng, Jure Leskovec, Jinbo Xu

机构 * Stanford University(斯坦福大学) University of California, San Diego(加州大学圣地亚哥分校) Wuhan University of Science and Technology(武汉科技大学) Hunan University(湖南大学)

专题命中 多模态生成 :multi-modal(title);分类 cs.AI

AI总结 SurfFlow通过多模态流匹配算法实现基于分子表面的肽共同设计,提升肽与受体的结合准确性,并在PepMerge基准中优于全原子基线。

Journal ref KDD 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.00477 2025-12-24 cs.CV 74%

LAMIC: Layout-Aware Multi-Image Composition via Scalability of Multimodal Diffusion Transformer

LAMIC:基于多模态扩散变换器可扩展性的布局感知多图像合成

Yuzhuo Chen, Zehua Ma, Jianhua Wang, Kai Kang, Shunyu Yao, Weiming Zhang

机构 * University of Science and Technology of China(中国科学技术大学) Onestory Team(Onestory团队) East China Normal University(华东师范大学)

专题命中 多模态生成 :multimodal(title);分类 cs.CV

AI总结 LAMIC通过无训练方式实现多参考图像合成,提升布局控制与背景一致性。

Comments 8 pages, 5 figures, 3 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.07934 2025-11-12 cs.CV 74%

Laytrol: Preserving Pretrained Knowledge in Layout Control for Multimodal Diffusion Transformers

Sida Huang, Siqi Huang, Ping Luo, Hongyuan Zhang

专题命中 多模态生成 :multimodal(title);分类 cs.CV

Comments Accepted by AAAI 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.07816 2025-11-12 cs.CV 74%

Cancer-Net PCa-MultiSeg: Multimodal Enhancement of Prostate Cancer Lesion Segmentation Using Synthetic Correlated Diffusion Imaging

Jarett Dewbury, Chi-en Amy Tai, Alexander Wong

机构 * Systems Design Engineering University of Waterloo(水力工程系统设计系大学)

专题命中 多模态生成 :multimodal(title);分类 cs.CV

Comments Accepted at ML4H 2025 Findings

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.22439 2025-10-30 cs.SD cs.AI 74%

PromptReverb: Multimodal Room Impulse Response Generation Through Latent Rectified Flow Matching

Ali Vosoughi, Yongyi Zang, Qihui Yang, Nathan Paek, Randal Leistikow, Chenliang Xu

机构 * Smule Labs(Smule实验室) University of California, San Diego(加州大学圣地亚哥分校) University of Rochester(罗切斯特大学) Stanford University(斯坦福大学)

专题命中 多模态生成 :multimodal(title);分类 cs.AI

Comments 9 pages, 2 figures, 4 tables; v2: corrected spelling of a co-author name; no content changes

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.22431 2025-10-28 cs.MA cs.CV 74%

Hollywood Town: Long-Video Generation via Cross-Modal Multi-Agent Orchestration

Zheng Wei, Mingchen Li, Zeqian Zhang, Ruibin Yuan, Pan Hui, Huamin Qu, James Evans, Maneesh Agrawala, Anyi Rao

机构 * The Hong Kong University of Science and Technology(香港科学与技术大学) The Hong Kong University of Science and Technology (Guangzhou)(香港科学与技术大学(广州)) University of Chicago(芝加哥大学) Stanford University(斯坦福大学)

专题命中 多模态生成 :cross-modal(title);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.23635 2025-09-30 cs.CV 74%

MotionVerse: A Unified Multimodal Framework for Motion Comprehension, Generation and Editing

Ruibing Hou, Mingshuang Luo, Hongyu Pan, Hong Chang, Shiguang Shan

机构 * Key Laboratory of Intelligent Information Processing, Institute of Computing Technology (ICT), Chinese Academy of Sciences (CAS)(智能信息处理重点实验室,计算技术研究所(ICT),中国科学院(CAS)) University of the Chinese Academy of Sciences(中国科学院大学)

专题命中 多模态生成 :multimodal(title);分类 cs.CV

Comments 17 pages, 6 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.16768 2025-09-23 cs.CV 74%

MMPart: Harnessing Multi-Modal Large Language Models for Part-Aware 3D Generation

Omid Bonakdar, Nasser Mozayani

机构 * School of Computer engineering, Iran university of Science and Technology(计算机工程学院,伊朗科学技术大学)

专题命中 多模态生成 :multi-modal(title);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.13227 2025-09-18 math.OC cs.AI cs.SY eess.SY 74%

Rich Vehicle Routing Problem in Disaster Management enabling Temporally-causal Transhipments across Multi-Modal Transportation Network

Santanu Banerjee, Goutam Sen, Siddhartha Mukhopadhyay

机构 * Department of Industrial and Systems Engineering (ISE), Indian Institute of Technology (IIT) Kharagpur(工业与系统工程系,印度理工学院Kharagpur分校)

专题命中 多模态生成 :multi-modal(title);分类 cs.AI

Comments Major changes in version II: 1) Supplementary is now a separate document, 2) Algorithm steps have been updated with pseudocode in the Heuristic, 3) Explanation of the MILP formulation construction is further detailed in a supplementary section

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.20877 2025-09-12 cs.CV 74%

Deep Learning Framework for Early Detection of Pancreatic Cancer Using Multi-Modal Medical Imaging Analysis

Dennis Slobodzian, Amir Kordijazi

专题命中 多模态生成 :multi-modal(title);分类 cs.CV

Comments 21 pages, 17 figure

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.16376 2025-09-09 cs.CV 74%

LaPIG: Cross-Modal Generation of Paired Thermal and Visible Facial Images

Leyang Wang, Joice Lin

机构 * University College London(伦敦大学学院) Xiamen University(厦门大学)

专题命中 多模态生成 :cross-modal(title);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.08987 2025-08-13 cs.CV cs.HC 74%

ColorGPT: Leveraging Large Language Models for Multimodal Color Recommendation

Ding Xia, Naoto Inoue, Qianru Qiu, Kotaro Kikuchi

机构 * The University of Tokyo(东京大学) CyberAgent AI Lab(CyberAgent AI实验室)

专题命中 多模态生成 :multimodal(title);分类 cs.CV

Comments Accepted to ICDAR2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.23676 2025-08-01 cs.LG cs.CV 74%

DepMicroDiff: Diffusion-Based Dependency-Aware Multimodal Imputation for Microbiome Data

Rabeya Tus Sadia, Qiang Cheng

机构 * Department of Computer Science University of Kentucky(计算机科学系 哥伦比亚大学) Department of Computer Science, Institute for Biomedical Informatics University of Kentucky(计算机科学系 生物医学信息学研究所 哥伦比亚大学)

专题命中 多模态生成 :multimodal(title);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.21260 2025-07-30 cs.LG cs.AI q-bio.QM 74%

Adaptive Multimodal Protein Plug-and-Play with Diffusion-Based Priors

Amartya Banerjee, Xingyu Xu, Caroline Moosmüller, Harlin Lee

机构 * University of North Carolina at Chapel Hill(北卡罗来纳大学教堂山分校) Carnegie Mellon University(卡内基梅隆大学)

专题命中 多模态生成 :multimodal(title);分类 cs.AI

Comments Code: https://github.com/amartya21/Adam-PnP

详情

展开后加载摘要…

URL PDF HTML 收藏