arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

多模态大模型

跨文本、图像、视频、音频等模态的大模型与学习方法。

2026-02-04 至 2026-02-04 共收录 16 信号源:cs.CV, cs.CL, cs.AI, cs.MM, eess.AS

1. 多模态生成 16 篇

2602.00508 2026-02-04 cs.CV 85%

DuoGen: Towards General Purpose Interleaved Multimodal Generation

DuoGen:迈向通用的交错多模态生成

Min Shi, Xiaohui Zeng, Jiannan Huang, Yin Cui, Francesco Ferroni, Jialuo Li, Shubham Pachori, Zhaoshuo Li, Yogesh Balaji, Haoxiang Wang, Tsung-Yi Lin, Xiao Fu, Yue Zhao, Chieh-Yun Chen, Ming-Yu Liu, Humphrey Shi

机构 * Georgia Tech(佐治亚理工学院) NVIDIA research.nvidia.com/labs/dir/duogen(NVIDIA 研究所)

专题命中 多模态生成 :multimodal(title,abstract);MLLM(abstract);image-text(abstract);分类 cs.CV

AI总结 DuoGen通过系统性地解决数据整理、架构设计和评估问题,实现了通用的交错多模态生成,提升了文本质量、图像保真度和图像-上下文对齐性能。

Comments Technical Report. Project Page: https://research.nvidia.com/labs/dir/duogen/

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.01769 2026-02-04 cs.LG cs.AI 83%

IRIS: Implicit Reward-Guided Internal Sifting for Mitigating Multimodal Hallucination

IRIS: 隐式奖励引导的内部筛选以缓解多模态幻觉

Yuanshuai Li, Yuping Yan, Jirui Han, Fei Ming, Lingjuan Lv, Yaochu Jin

机构 * Department of Artificial Intelligence, Westlake University, Hangzhou, China(人工智能系,西湖大学,杭州,中国) Sony Research, Sony(索尼研究,索尼)

专题命中 多模态生成 :multimodal(title,abstract);MLLM(abstract);分类 cs.AI

AI总结 IRIS通过隐式奖励引导内部筛选,有效缓解多模态大语言模型的幻觉问题,无需外部反馈且性能优异。

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.02590 2026-02-04 cs.RO 82%

StepNav: Structured Trajectory Priors for Efficient and Multimodal Visual Navigation

StepNav: 为高效且多模态视觉导航引入结构化轨迹先验

Xubo Luo, Aodi Wu, Haodong Han, Xue Wan, Wei Zhang, Leizheng Shu, Ruisuo Wang

机构 * University of Chinese Academy of Sciences(中国科学院大学) Technology and Engineering Center for Space Utilization, Chinese Academy of Sciences(中国科学院空间利用技术与工程中心)

专题命中 多模态生成 :multimodal(title,abstract);multi-modal(abstract)

AI总结 StepNav通过结构化多模态轨迹先验提升视觉导航的鲁棒性、效率和安全性。

Comments 8 pages, 7 figures; Accepted by ICRA 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.09131 2026-02-04 cs.GR cs.AI cs.CV 81%

Training-Free Text-Guided Color Editing with Multi-Modal Diffusion Transformer

无需训练的文本引导颜色编辑与多模态扩散变换器

Zixin Yin, Xili Dai, Ling-Hao Chen, Deyu Zhou, Jianan Wang, Duomin Wang, Gang Yu, Lionel M. Ni, Lei Zhang, Heung-Yeung Shum

机构 * The Hong Kong University of Science and Technology(香港科学与技术大学) International Digital Economy Academy(国际数字经济学院) The Hong Kong University of Science and Technology (Guangzhou)(香港科学与技术大学(广州)) Tsinghua University(清华大学) Astribot StepFun

专题命中 多模态生成 :multi-modal(title,abstract);分类 cs.CV、cs.AI

AI总结 ColorCtrl通过多模态扩散变换器实现无需训练的文本引导颜色编辑,精准控制颜色属性并保持一致性,优于现有方法和商业模型。

Comments https://zxyin.github.io/ColorCtrl

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.12728 2026-02-04 cs.CV cs.MM 81%

SpecFLASH: A Latent-Guided Semi-autoregressive Speculative Decoding Framework for Efficient Multimodal Generation

SpecFLASH: 一种基于潜在引导的半自回归推测解码框架,用于高效多模态生成

Zihua Wang, Ruibo Li, Haozhe Du, Joey Tianyi Zhou, Yu Zhang, Xu Yang

机构 * Southeast University, Nanjing, China(东南大学) Nanyang Technological University, Singapore(南洋理工大学) A STAR Centre for Frontier AI Research (CFAR), Singapore(A STAR前沿人工智能研究中心)

专题命中 多模态生成 :multimodal(title,abstract);分类 cs.CV、cs.MM

AI总结 SpecFLASH是一种针对多模态生成的高效推测解码框架,通过潜在引导的token压缩和半自回归解码方案,显著提升视觉任务的解码速度。

Comments Under review

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.12203 2026-02-04 cs.CV 79%

LazyDrag: Enabling Stable Drag-Based Editing on Multi-Modal Diffusion Transformers via Explicit Correspondence

LazyDrag: 通过显式对应关系在多模态扩散变换器上实现稳定的拖拽编辑

Zixin Yin, Xili Dai, Duomin Wang, Xianfang Zeng, Lionel M. Ni, Gang Yu, Heung-Yeung Shum

机构 * The Hong Kong University of Science and Technology(香港科技大学) StepFun The Hong Kong University of Science and Technology (Guangzhou)(香港科技大学(广州))

专题命中 多模态生成 :multi-modal(title,abstract);分类 cs.CV

AI总结 LazyDrag通过显式对应关系实现多模态扩散变换器的稳定拖拽编辑,消除了隐式点匹配依赖,提升生成能力与精确控制。

Comments https://zxyin.github.io/LazyDrag

详情

展开后加载摘要…

URL PDF HTML 收藏
2412.20418 2026-02-04 eess.IV cs.CV 79%

Diff4MMLiTS: Advanced Multimodal Liver Tumor Segmentation via Diffusion-Based Image Synthesis and Alignment

Diff4MMLiTS: 通过基于扩散的图像合成与对齐的先进多模态肝肿瘤分割

Shiyun Chen, Li Lin, Pujin Cheng, ZhiCheng Jin, JianJian Chen, HaiDong Zhu, Kenneth K. Y. Wong, Xiaoying Tang

机构 * Department of Electronic and Electrical Engineering, Southern University of Science and Technology, Shenzhen, China(电子与电气工程系,南方科技大学,深圳,中国) Department of Electrical and Electronic Engineering, The University of Hong Kong, Hong Kong SAR, China(电气与电子工程系,香港大学,香港特别行政区,中国) Department of Radiology, Zhongda Hospital, Medical School, Southeast University, Nanjing, China(放射科,中大医院,医学院,东南大学,南京,中国) Jiaxing Research Institute, Southern University of Science and Technology, Jiaxing, China(嘉兴研究所,南方科技大学,嘉兴,中国)

专题命中 多模态生成 :multimodal(title,abstract);分类 cs.CV

AI总结 Diff4MMLiTS通过基于扩散的图像合成与对齐技术,实现肝肿瘤的多模态分割,无需严格对齐的多模态数据,提升了分割性能。

Comments International Workshop on Machine Learning in Medical Imaging, 668-678

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.02927 2026-02-04 stat.ML cs.LG 78%

Training-Free Self-Correction for Multimodal Masked Diffusion Models

无需训练的多模态掩码扩散模型自校正

Yidong Ouyang, Panwen Hu, Zhengyan Wan, Zhe Wang, Liyan Xie, Dmitriy Bespalov, Ying Nian Wu, Guang Cheng, Hongyuan Zha, Qiang Sun

机构 * University of California, Los Angeles(加州大学洛杉矶分校) Mohamed bin Zayed University of Artificial Intelligence(莫莫德·本·扎耶德人工智能大学) East China Normal University(华东师范大学) University of Virginia(弗吉尼亚大学) University of Minnesota(明尼苏达大学) Drexel university(德雷塞尔大学) The Chinese University of Hong Kong, Shenzhen(香港中文大学(深圳)) University of Toronto(多伦多大学)

专题命中 多模态生成 :multimodal(title,abstract)

AI总结 本文提出无需训练的多模态掩码扩散模型自校正方法,通过减少采样步骤提升生成质量,适用于文本到图像和多模态理解任务。

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.02033 2026-02-04 cs.CV cs.AI cs.MM 75%

One Size, Many Fits: Aligning Diverse Group-Wise Click Preferences in Large-Scale Advertising Image Generation

一个尺寸,多种适配:在大规模广告图像生成中对多样化群体点击偏好进行对齐

Shuo Lu, Haohan Wang, Wei Feng, Weizhen Wang, Shen Zhang, Yaoyu Li, Ao Ma, Zheng Zhang, Jingjing Lv, Junjie Shen, Ching Law, Bing Zhan, Yuan Xu, Huizai Yao, Yongcan Yu, Chenyang Si, Jian Liang

机构 * NLPR & MAIS, CASIA(中国科学院长春光学精密机械与物理研究所 & 中国科学院自动化所) School of AI, UCAS(中国科学院大学人工智能学院) HKUST(gz)(香港科技大学) PRLab, NJU(南京大学PRLab)

专题命中 多模态生成 :multimodal(abstract);MLLM(abstract);分类 cs.CV、cs.AI、cs.MM

AI总结 本文提出OSMF框架,通过自适应分组和群组感知多模态模型,解决广告图像生成中用户群体点击偏好多样性的优化问题。

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.03414 2026-02-04 cs.CV cs.AI 73%

Socratic-Geo: Synthetic Data Generation and Geometric Reasoning via Multi-Agent Interaction

Socratic-Geo:通过多智能体交互实现合成数据生成与几何推理

Zhengbo Jiao, Shaobo Wang, Zifan Zhang, Wei Wang, Bing Zhao, Hu Wei, Linfeng Zhang

机构 * AI DATA, Alibaba Group Holding Limited(阿里数据,阿里巴巴集团控股有限公司) EPIC Lab, Shanghai Jiao Tong University(上海交通大学EPIC实验室) Shanghai University of Finance(上海财经大学) Wuhan University(武汉大学)

专题命中 多模态生成 :multimodal(abstract);image-text(abstract);分类 cs.CV、cs.AI

AI总结 Socratic-Geo通过多智能体交互实现合成数据生成与几何推理,利用教师代理和求解代理动态耦合数据合成与模型学习,提升图像生成和推理能力。

Comments 18pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.03533 2026-02-04 cs.CV 70%

PnP-U3D: Plug-and-Play 3D Framework Bridging Autoregression and Diffusion for Unified Understanding and Generation

PnP-U3D: 无缝接入3D框架,连接自回归与扩散以实现统一的理解与生成

Yongwei Chen, Tianyi Wei, Yushi Lan, Zhaoyang Lyu, Shangchen Zhou, Xudong Xu, Xingang Pan

机构 * Shanghai Artificial Intelligence Laboratory(上海人工智能实验室)

专题命中 多模态生成 :multimodal(abstract);cross-modal(abstract);分类 cs.CV

AI总结 PnP-U3D提出结合自回归与扩散的统一3D理解和生成框架,通过轻量级Transformer实现跨模态信息交换,提升3D生成与编辑性能。

Comments Yongwei Chen and Tianyi Wei contributed equally. Project page: https://cyw-3d.github.io/PnP-U3D/

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.03069 2026-02-04 cs.DB 67%

Skill-Based Autonomous Agents for Material Creep Database Construction

基于技能的自主代理用于材料蠕变数据库构建

Yue Wu, Tianhao Su, Shunbo Hu, Deng Pan

专题命中 多模态生成 :multi-modal(abstract);cross-modal(abstract)

AI总结 本文提出基于技能的自主代理框架,用于从科学文献中自动提取高保真的材料蠕变数据,实现物理自洽的数据库构建。

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.02820 2026-02-04 cs.LG cs.AI cs.CV 62%

From Tokens to Numbers: Continuous Number Modeling for SVG Generation

从标记到数字:用于SVG生成的连续数字建模

Michael Ogezi, Martin Bell, Freda Shi, Ethan Smith

机构 * Cheriton School of Computer Science, University of Waterloo, Waterloo, ON, Canada(滑铁卢大学计算机科学学院) Vector Institute(向量研究所)

专题命中 多模态生成 :multimodal(abstract);分类 cs.CV、cs.AI

AI总结 本文提出连续数字建模(CNM)方法,通过直接建模连续数值提升SVG生成的效率与质量,实现训练速度提升30%及更高的视觉保真度。

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.03215 2026-02-04 stat.ML cs.AI cs.LG 57%

Latent Neural-ODE for Model-Informed Precision Dosing: Overcoming Structural Assumptions in Pharmacokinetics

隐式神经微分方程用于模型指导的精准给药:克服药代动力学中的结构假设

Benjamin Maurel, Agathe Guilloux, Sarah Zohar, Moreno Ursino, Jean-Baptiste Woillard

专题命中 多模态生成 :multi-modal(abstract);分类 cs.AI

AI总结 本文提出基于隐式神经ODE的模型,用于精准给药中的他克莫司AUC预测,通过模拟和临床验证展示了其在复杂生物动态建模中的优越性能。

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.02924 2026-02-04 cs.LG cs.SY eess.SY 50%

How Does the Lagrangian Guide Safe Reinforcement Learning through Diffusion Models?

如何通过扩散模型使拉格朗日量引导安全强化学习?

Xiaoyuan Cheng, Wenxuan Yuan, Boyang Li, Yuanchao Xu, Yiming Yang, Hao Liang, Bei Peng, Robert Loftin, Zhuo Sun, Yukun Hu

机构 * University College London(伦敦大学学院) Imperial College London(伦敦帝国学院) University of California, San Diego(加州大学圣地亚哥分校) Kyoto University(京都大学) King's College London(伦敦国王学院) University of Sheffield(谢菲尔德大学) Shanghai University of Finance and Economics(上海财经大学)

专题命中 多模态生成 :multimodal(abstract)

AI总结 本文提出ALGD算法,通过增强拉格朗日方法解决安全强化学习中的稳定性问题,实现稳定有效的政策生成。

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.14943 2026-02-04 cs.HC 50%

State of the Art of LLM-Enabled Interaction with Visualization

大语言模型赋能的可视化交互现状

Mathis Brossier, Tobias Isenberg, Konrad Schönborn, Jonas Unger, Mario Romero, Johanna Björklund, Anders Ynnerman, Lonni Besançon

专题命中 多模态生成 :multimodal(abstract)

AI总结 本文系统回顾了大语言模型与可视化交互的研究现状,分析了6个维度的48篇论文,突出了新兴设计模式和评估挑战,旨在指导未来LLM增强的可视化研究和系统设计。

Comments Submitted to STARs of EuroVis'26

详情

展开后加载摘要…

URL PDF HTML 收藏