arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

多模态大模型

跨文本、图像、视频、音频等模态的大模型与学习方法。

2026-02-10 至 2026-02-10 共收录 18 信号源:cs.CV, cs.CL, cs.AI, cs.MM, eess.AS

1. 多模态生成 18 篇

2602.07993 2026-02-10 cs.CV cs.AI 84%

MCIE: Multimodal LLM-Driven Complex Instruction Image Editing with Spatial Guidance

MCIE: 多模态大语言模型驱动的复杂指令图像编辑与空间引导

Xuehai Bai, Xiaoling Gu, Akide Liu, Hangjie Yuan, YiFan Zhang, Jack Ma

专题命中 多模态生成 :multimodal(title,abstract);MLLM(abstract);分类 cs.CV、cs.AI

AI总结 MCIE通过空间引导和背景一致性模块,提升复杂指令图像编辑的指令合规性,实现23.96%的性能提升。

Comments Accepted by AAAI2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.08249 2026-02-10 eess.IV cs.CV 79%

A Unified Framework for Multimodal Image Reconstruction and Synthesis using Denoising Diffusion Models

基于去噪扩散模型的多模态图像重建与合成统一框架

Weijie Gan, Xucheng Wang, Tongyao Wang, Wenshang Wang, Chunwei Ying, Yuyang Hu, Yasheng Chen, Hongyu An, Ulugbek S. Kamilov

机构 * Department of Computer Science and Engineering, Washington University in St. Louis(华盛顿大学圣路易斯分校计算机科学与工程系) Mallinckrodt Institute of Radiology, Washington University in St. Louis(华盛顿大学圣路易斯分校马林克罗德特放射医学研究所) Department of Electrical and Systems Engineering, Washington University in St. Louis(华盛顿大学圣路易斯分校电气与系统工程系) Department of Neurology, Washington University in St. Louis(华盛顿大学圣路易斯分校神经病学系) Department of Biomedical Engineering, Washington University in St. Louis(华盛顿大学圣路易斯分校生物医学工程系) Division of Biology and Biomedical Sciences, Washington University in St. Louis(华盛顿大学圣路易斯分校生物学与生物医学科学 division) Department of Electrical and Computer Engineering, University of Wisconsin–Madison(威斯康星大学麦迪逊分校电气与计算机工程系)

专题命中 多模态生成 :multimodal(title,abstract);分类 cs.CV

AI总结 Any2all通过统一框架实现多模态图像重建与合成,利用去噪扩散模型提升性能与质量。

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.09851 2026-02-10 cs.RO cs.CV 79%

Simultaneous Tactile-Visual Perception for Learning Multimodal Robot Manipulation

同时触觉-视觉感知用于学习多模态机器人操作

Yuyang Li, Yinghan Chen, Zihang Zhao, Puhao Li, Tengyu Liu, Siyuan Huang, Yixin Zhu

机构 * Institute for Artificial Intelligence, Peking University(北京大学人工智能研究院) School of Psychological and Cognitive Sciences, Peking University(北京大学心理与认知科学学院) Beijing Key Lab of Behavior and Mental Health, Peking University(北京大学行为与心理健康北京市重点实验室) Beijing Institute for General Artificial Intelligence(北京一般人工智能研究院) State Key Lab for General Artificial Intelligence(一般人工智能国家重点实验室) Embodied Intelligence Lab, PKU-Wuhan Institute for Artificial Intelligence(具身智能实验室,北京大学武汉人工智能研究院) Department of Computer Science and Technology, University of Cambridge(剑桥大学计算机科学与技术系)

专题命中 多模态生成 :multimodal(title,abstract);分类 cs.CV

AI总结 TacThru-UMI通过结合同时触觉-视觉感知与现代学习框架,实现了高精度多模态机器人操作。

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.23278 2026-02-10 cs.CV 79%

UniLiP: Adapting CLIP for Unified Multimodal Understanding, Generation and Editing

UniLiP: 适配CLIP以实现统一的多模态理解、生成与编辑

Hao Tang, Chenwei Xie, Xiaoyi Bao, Tingyu Weng, Pandeng Li, Yun Zheng, Liwei Wang

机构 * Center for Data Science, Peking University(北京大学数据科学中心) Alibaba Group(阿里巴巴集团) CASIA Center for Machine Learning Research, Peking University(北京大学机器学习研究中心) State Key Laboratory of General Artificial Intelligence, Peking University(北京大学通用人工智能国家重点实验室)

专题命中 多模态生成 :multimodal(title,abstract);分类 cs.CV

AI总结 UniLIP通过两阶段训练和双条件架构,提升CLIP在多模态理解、生成和编辑任务中的性能,实现高效参数下的高表现。

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.25682 2026-02-10 cs.CL 79%

PairUni: Pairwise Training for Unified Multimodal Language Models

PairUni: 用于统一多模态语言模型的成对训练

Jiani Zheng, Zhiyang Teng, Kunpeng Qiu, Xiangtai Li, Anran Wang, Yu Tian, Ye Tian, Haochen Wang, Zhuochen Wang

机构 * ByteDance(字节跳动)

专题命中 多模态生成 :multimodal(title);cross-modal(abstract);分类 cs.CL

AI总结 PairUni通过成对训练提升统一多模态语言模型的理解与生成能力,采用配对数据集和PairGRPO算法实现更有效的策略学习。

Comments 22 pages, 11 figures, and 10 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.08830 2026-02-10 cs.HC 78%

Enhancing Generative AI Image Refinement with Scribbles and Annotations: A Comparative Study of Multimodal Prompts

通过草图和注释增强生成式AI图像细化:多模态提示的比较研究

Hyerim Park, Phuong Thao Tran, Andre Luckow, Ceenu George, Michael Sedlmair, Malin Eiband

专题命中 多模态生成 :multimodal(title,abstract)

AI总结 本研究通过比较多模态提示,探讨草图和注释如何提升生成式AI图像细化,提出原型并揭示设计师在多模态策略中的偏好。

Comments 22 pages, 14 figures. Preprint of an accepted IUI '26 paper

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.04780 2026-02-10 cs.LG cond-mat.dis-nn 78%

Dynamical Regimes of Multimodal Diffusion Models

多模态扩散模型的动力学区域

Emil Albrychiewicz, Andrés Franco Valiente, Li-Ching Chen

机构 * Leinweber Institute for Theoretical Physics(莱因韦伯理论物理研究所) Department of Physics, University of California, Berkeley, CA, 94720-7300, USA(加州大学伯克利分校物理系) Theoretical Physics Group, Lawrence Berkeley National Laboratory(劳伦斯伯克利国家实验室理论物理组) Department of Radiation Oncology, University of California, San Francisco(加州大学旧金山分校放射肿瘤科) Computational Precision Health, University of California San Francisco(加州大学旧金山分校计算精准健康)

专题命中 多模态生成 :multimodal(title,abstract)

AI总结 多模态扩散模型通过耦合动态机制,揭示生成过程中的时间层次结构及同步间隙现象,为生成模型的稳定性提供理论框架。

Comments 40 pages, 14 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.08145 2026-02-10 cs.LG cs.AI cs.CL cs.CV cs.CY 67%

Reliable and Responsible Foundation Models: A Comprehensive Survey

可靠且负责任的基础模型:全面综述

Xinyu Yang, Junlin Han, Rishi Bommasani, Jinqi Luo, Wenjie Qu, Wangchunshu Zhou, Adel Bibi, Xiyao Wang, Jaehong Yoon, Elias Stengel-Eskin, Shengbang Tong, Lingfeng Shen, Rafael Rafailov, Runjia Li, Zhaoyang Wang, Yiyang Zhou, Chenhang Cui, Yu Wang, Wenhao Zheng, Huichi Zhou, Jindong Gu, Zhaorun Chen, Peng Xia, Tony Lee, Thomas Zollo, Vikash Sehwag, Jixuan Leng, Jiuhai Chen, Yuxin Wen, Huan Zhang, Zhun Deng, Linjun Zhang, Pavel Izmailov, Pang Wei Koh, Yulia Tsvetkov, Andrew Wilson, Jiaheng Zhang, James Zou, Cihang Xie, Hao Wang, Philip Torr, Julian McAuley, David Alvarez-Melis, Florian Tramèr, Kaidi Xu, Suman Jana, Chris Callison-Burch, Rene Vidal, Filippos Kokkinos, Mohit Bansal, Beidi Chen, Huaxiu Yao

专题命中 多模态生成 :multimodal(abstract);分类 cs.CV、cs.CL、cs.AI

AI总结 本文综述了基础模型的可靠和负责任发展,探讨了偏见、安全、不确定性等关键问题,并提出了未来研究方向。

Comments TMLR camera-ready version

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.07983 2026-02-10 cs.AI cs.CL 62%

Accelerating Social Science Research via Agentic Hypothesization and Experimentation

通过代理假设和实验加速社会科学研究

Jishu Sen Gupta, Harini SI, Somesh Kumar Singh, Syed Mohamad Tawseeq, Yaman Kumar Singla, David Doermann, Rajiv Ratn Shah, Balaji Krishnamurthy

机构 * Adobe Media and Data Science Research (MDSR)(Adobe媒体与数据科学研究所)

专题命中 多模态生成 :multimodal(abstract);分类 cs.CL、cs.AI

AI总结 EXPERIGEN通过代理框架实现端到端的科学发现,发现更多显著且预测性强的假设,并通过A/B测试验证其有效性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2310.08785 2026-02-10 cs.CV cs.AI 62%

DeltaSpace: A Semantic-aligned Feature Space for Flexible Text-guided Image Editing

DeltaSpace: 一种语义对齐的特征空间用于灵活的文本引导图像编辑

Yueming Lyu, Kang Zhao, Bo Peng, Huafeng Chen, Yue Jiang, Yingya Zhang, Jing Dong, Caifeng Shan

专题命中 多模态生成 :image-text(abstract);分类 cs.CV、cs.AI

AI总结 DeltaSpace通过语义对齐的特征空间实现文本引导图像编辑的灵活训练和推理,支持零样本推理和无需文本的训练。

Comments 18 pages. arXiv admin note: text overlap with arXiv:2303.06285

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.07176 2026-02-10 cs.CL cs.AI cs.ET cs.HC 62%

Open TutorAI: An Open-source Platform for Personalized and Immersive Learning with Generative AI

Open TutorAI: 一个基于生成AI的开源平台,用于个性化和沉浸式学习

Mohamed El Hajji, Tarek Ait Baha, Aicha Dakir, Hammou Fadili, Youssef Es-Saady

机构 * IRF-SIC Laboratory, Ibnou Zohr University(IRF-SIC实验室,伊本·扎赫尔大学) Regional Center for Education(教育与培训专业地区中心) Polydisciplinary Faculty of Taroudant, Ibnou Zohr University(塔鲁旦多学科学院,伊本·扎赫尔大学) Higher School of Technology of Guelmim, Ibnou Zohr University(盖尔米姆技术高等学校,伊本·扎赫尔大学) Paragraphe laboratory, Paris 8 and CY Cergy Paris Universities(Paragraphe实验室,巴黎8大学和CY塞克巴黎大学)

专题命中 多模态生成 :multimodal(abstract);分类 cs.CL、cs.AI

AI总结 Open TutorAI 是一个基于生成AI的开源平台,通过个性化和沉浸式学习体验提升教育效果。

Comments 19 pages, 15 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.06133 2026-02-10 cs.LG cs.AI cs.RO 57%

A Review of Online Diffusion Policy RL Algorithms for Scalable Robotic Control

在线扩散策略强化学习算法综述:可扩展机器人控制

Wonhyeok Choi, Shutong Ding, Minwoo Choi, Jungwan Woo, Kyumin Hwang, Jaeyeul Kim, Ye Shi, Sunghoon Im

机构 * Daegu Gyeongbuk Institute of Science and Technology(大邱庆尚科学技术院) ShanghaiTech University(上海科技大学)

专题命中 多模态生成 :multimodal(abstract);分类 cs.AI

AI总结 本文综述了在线扩散策略强化学习算法,分析了其在可扩展机器人控制中的性能、挑战及未来发展方向。

详情

展开后加载摘要…

URL PDF HTML 收藏
2409.09777 2026-02-10 cs.CV cs.RO 57%

EgoFSD: Ego-Centric Fully Sparse Paradigm with Uncertainty Denoising and Iterative Refinement for Efficient End-to-End Self-Driving

EgoFSD:面向端到端自动驾驶的以自我为中心的完全稀疏范式,结合不确定性去噪和迭代细化

Haisheng Su, Wei Wu, Zhenjie Yang, Isabel Guan

机构 * School of Computer Science, Shanghai Jiao Tong University(上海交通大学计算机科学学院) SenseAuto The Hong Kong University of Science and Technology(香港理工大学)

专题命中 多模态生成 :multi-modal(abstract);分类 cs.CV

AI总结 EgoFSD通过引入稀疏感知、分层交互和迭代运动规划,提升端到端自动驾驶的效率和性能,减少误差和碰撞,提高训练稳定性。

Comments Accepted to ICRA2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.07564 2026-02-10 cs.CV 57%

SIGMA: Selective-Interleaved Generation with Multi-Attribute Tokens

SIGMA: 带多属性标记的选择性交错生成

Xiaoyan Zhang, Zechen Bai, Haofan Wang, Yiren Song

机构 * Creatly AI University of Michigan(密歇根大学) Lovart AI National University of Singapore(新加坡国立大学)

专题命中 多模态生成 :multimodal(abstract);分类 cs.CV

AI总结 SIGMA通过引入多属性标记实现多条件生成,提升可控性、一致性和视觉质量,优于Bagel在组合任务中的表现。

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.07243 2026-02-10 cs.RO cs.AI cs.GR 57%

Realistic Synthetic Household Data Generation at Scale

大规模生成逼真合成家庭数据

Siddharth Singh, Ifrah Idrees, Abraham Dauhajre

机构 * Siddharth Singh(1 西雅图)

专题命中 多模态生成 :multi-modal(abstract);分类 cs.AI

AI总结 本文提出了一种大规模生成逼真合成家庭数据的方法,通过松耦合生成人机交互与环境数据,实现双向影响建模,提升家庭智能设备的开发与测试能力。

Comments Accepted at Agentic AI Benchmarks and Applications for Enterprise Tasks workshop at AAAI 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.08910 2026-02-10 cond-mat.dis-nn cond-mat.stat-mech q-bio.NC q-bio.PE 50%

Structural coarse-graining enables noise-robust functional connectivity and reveals hidden inter-subject variability

结构粗粒化实现噪声鲁棒的功能连接并揭示隐藏的跨受试者变异性

Izaro Fernandez-Iriondo, Antonio Jimenez-Marin, Jesus Cortes, Pablo Villegas

专题命中 多模态生成 :multimodal(abstract)

AI总结 本文提出一种结合扩散结构粗粒化和频谱噪声过滤的方法,以从有限时间数据中恢复可靠的高维功能网络,揭示隐藏的跨受试者变异。

Comments 10 Pages, 4 Figures and Supplementary Information

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.07735 2026-02-10 cs.LG q-bio.BM 50%

TerraBind: Fast and Accurate Binding Affinity Prediction through Coarse Structural Representations

TerraBind: 通过粗粒结构性表示实现快速且准确的结合亲和力预测

Matteo Rossi, Ryan Pederson, Miles Wang-Henderson, Ben Kaufman, Edward C. Williams, Carl Underkoffler, Owen Lewis Howell, Adrian Layer, Stephan Thaler, Narbe Mardirossian, John Anthony Parkhill

机构 * Terray Therapeutics, Inc.(Terray Therapeutics公司)

专题命中 多模态生成 :multimodal(abstract)

AI总结 TerraBind通过粗粒结构性表示实现了快速且准确的结合亲和力预测,其推理速度比现有方法快26倍,同时在亲和力预测准确性上提升了约20%。

Comments 31 pages, 14 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.06990 2026-02-10 eess.SP 50%

A Pre-trained EEG-to-MEG Generative Framework for Enhancing BCI Decoding

一种预训练的EEG到MEG生成框架用于增强BCI解码

Zhuo Li, Shuqiang Wang

专题命中 多模态生成 :cross-modal(abstract)

AI总结 本文提出了一种基于EEG-MEG时空耦合表示的跨模态生成框架,通过合成MEG信号提升BCI解码性能。

详情

展开后加载摘要…

URL PDF HTML 收藏