arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

多模态大模型

跨文本、图像、视频、音频等模态的大模型与学习方法。

共收录 4965 信号源:cs.CV, cs.CL, cs.AI, cs.MM, eess.AS

1. 多模态生成 4965 篇

2601.03903 2026-01-08 cs.IR 50%

Unleashing the Potential of Neighbors: Diffusion-based Latent Neighbor Generation for Session-based Recommendation

释放邻居的潜力:基于扩散的潜在邻居生成用于基于会话的推荐

Yuhan Yang, Jie Zou, Guojia An, Jiwei Wei, Yang Yang, Heng Tao Shen

专题命中 多模态生成 :multi-modal(abstract)

AI总结 DiffSBR通过基于扩散的潜在邻居生成方法,提升基于会话的推荐性能。

Comments This paper has been accepted by KDD 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.24200 2026-01-01 cs.GR 50%

PartMotionEdit: Fine-Grained Text-Driven 3D Human Motion Editing via Part-Level Modulation

PartMotionEdit: 通过部分级调制实现细粒度文本驱动的3D人体运动编辑

Yujie Yang, Zhichao Zhang, Jiazhou Chen, Zichao Wu

专题命中 多模态生成 :cross-modal(abstract)

AI总结 PartMotionEdit通过部分级调制实现细粒度文本驱动的3D人体运动编辑,提升局部运动的精确控制与语义对齐能力。

Comments 10 pages, 4 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.22279 2025-12-30 cs.LG cs.CE physics.app-ph 50%

Hierarchical Stacking Optimization Using Dirichlet's Process (SoDip): Towards Accelerated Design for Graft Polymerization

基于Dirichlet过程的分层堆叠优化(SoDip):迈向聚合法的加速设计

Amgad Ahmed Ali Ibrahim, Hein Htet, Ryoji Asahi

专题命中 多模态生成 :multimodal(abstract)

AI总结 基于Dirichlet过程的分层堆叠优化框架通过整合Transformer、TabNet、XGBoost和贝叶斯优化,提升聚合法的可重复性和设计效率。

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.06153 2025-12-30 q-bio.NC math.AT stat.ME 50%

Topologically Invariant Permutation Test

拓扑不变排列检验

Sixtus Dakurah

专题命中 多模态生成 :multimodal(abstract)

AI总结 本文提出一种拓扑不变排列检验,用于检测功能脑网络的拓扑不等价性,通过2-瓦瑟斯坦距离和热核扩展实现拓扑结构的稳健比较。

Comments 24 pages, 8 figures, 3 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.21475 2025-12-29 cs.NI 50%

Physics-informed Diffusion Models for Multi-scale Prediction of Reference Signal Received Power in Wireless Networks

基于物理信息的扩散模型用于无线网络中参考信号接收功率的多尺度预测

Xiaoqian Qi, Haoye Chai, Yue Wang, Zhaocheng Wang, Yong Li

专题命中 多模态生成 :multimodal(abstract)

AI总结 本文提出Channel-Diff框架,通过物理信息条件扩散模型实现无线网络中参考信号接收功率的多尺度预测,提升预测精度与模型可迁移性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.04798 2025-12-23 cs.LG 50%

TabRep: Training Tabular Diffusion Models with a Simple and Effective Continuous Representation

TabRep: 通过简单有效的连续表示训练表格扩散模型

Jacob Si, Zijing Ou, Mike Qu, Zhengrui Xiang, Yingzhen Li

机构 * Imperial College London(伦敦帝国学院) Columbia University(哥伦比亚大学)

专题命中 多模态生成 :multi-modal(abstract)

AI总结 TabRep通过统一连续表示训练表格扩散模型,实现高效生成高质量表格数据并保持隐私。

Comments TMLR 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.17482 2025-12-09 cs.RO 50%

Variational Shape Inference for Grasp Diffusion on SE(3)

变分形状推断用于SE(3)上的抓取扩散

S. Talha Bukhari, Kaivalya Agrawal, Zachary Kingston, Aniket Bera

机构 * Department of Computer Science, Purdue University(计算机科学系,普渡大学)

专题命中 多模态生成 :multimodal(abstract)

AI总结 本文提出一种基于变分形状推断的SE(3)抓取扩散框架,通过隐式神经表示训练自编码器并引导扩散模型,实现鲁棒的多模态抓取合成,实验显示在ACRONYM数据集上性能提升6.3%并具备零样本迁移能力。

Comments This work has been submitted to the IEEE for possible publication

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.03874 2025-12-04 cs.RO cs.LG 50%

OmniDexVLG: Learning Dexterous Grasp Generation from Vision Language Model-Guided Grasp Semantics, Taxonomy and Functional Affordance

OmniDexVLG: 从视觉语言模型引导的抓取语义、分类法和功能可及性中学习灵巧抓取生成

Lei Zhang, Diwen Zheng, Kaixin Bai, Zhenshan Bing, Zoltan-Csaba Marton, Zhaopeng Chen, Alois Christian Knoll, Jianwei Zhang

机构 * University of Hamburg(汉堡大学) Agile Robots SE(敏捷机器人公司) Technical University of Munich(慕尼黑技术大学)

专题命中 多模态生成 :multimodal(abstract)

AI总结 OmniDexVLG通过多模态语义推理和统一视觉语言模型,实现从自然语言指令中生成多样且语义连贯的灵巧抓取。

Comments Project Website: https://sites.google.com/view/omnidexvlg, 16 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.02013 2025-12-02 cs.RO 50%

ManualVLA: A Unified VLA Model for Chain-of-Thought Manual Generation and Robotic Manipulation

ManualVLA: 一种统一的VLA模型用于链式思维手动生成和机器人操作

Chenyang Gu, Jiaming Liu, Hao Chen, Runzhong Huang, Qingpo Wuwu, Zhuoyang Liu, Xiaoqi Li, Ying Li, Renrui Zhang, Peng Jia, Pheng-Ann Heng, Shanghang Zhang

机构 * State Key Laboratory of Multimedia Information Processing, School of Computer Science, Peking University(多媒体信息处理国家重点实验室,计算机学院,北京大学) The Chinese University of Hong Kong(香港中文大学) Simplexity Robotics Project(Simplexity机器人项目)

专题命中 多模态生成 :multimodal(abstract)

AI总结 ManualVLA通过融合多模态手册生成与动作执行,提升机器人长周期任务中的规划与操作能力。

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.02058 2025-12-01 q-bio.BM cs.LG 50%

RiboGen: RNA Sequence and Structure Co-Generation with Equivariant MultiFlow

RiboGen:基于等变多流的RNA序列和结构联合生成

Dana Rubin, Allan dos Santos Costa, Manvitha Ponnapati, Joseph Jacobson

机构 * MIT CSAIL(麻省理工学院计算机科学与人工智能实验室) MIT Media Lab(麻省理工学院媒体实验室) Molecular Machines(分子机器) Center for Bits and Atoms(比特与原子中心)

专题命中 多模态生成 :multimodal(abstract)

AI总结 RiboGen通过等变多流方法实现RNA序列和结构的联合生成,为RNA设计提供了新的深度学习解决方案。

Comments 6 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.16900 2025-11-24 eess.SY cs.SY 50%

When Motion Learns to Listen: Diffusion-Prior Lyapunov Actor-Critic Framework with LLM Guidance for Stable and Robust AUV Control in Underwater Tasks

当运动学会倾听:带有LLM引导的扩散先验Lyapunov动作-批评者框架用于水下任务中稳定和鲁棒的AUV控制

Jingzehua Xu, Weiyi Liu, Weihang Zhang, Zhuofan Xi, Guanwen Xie, Shuai Zhang, Yi Li

专题命中 多模态生成 :multimodal(abstract)

AI总结 本文提出一种结合扩散模型、Lyapunov批评者和LLM的框架,用于提升水下机器人控制的稳定性与鲁棒性,通过生成-过滤-优化机制实现高效探索和多目标优化。

Comments This paper is currently under review and does not represent the final version. Jingzehua Xu, Weiyi Liu and Weihang Zhang are co-first authors of this paper, with Zhuofan Xi as the second author

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.15550 2025-11-21 cs.RO 50%

UltraDP: Generalizable Carotid Ultrasound Scanning with Force-Aware Diffusion Policy

UltraDP: 通用的颈动脉超声扫描与力感知扩散策略

Ruoqu Chen, Xiangjie Yan, Kangchen Lv, Gao Huang, Zheng Li, Xiang Li

机构 * Department of Automation, Tsinghua University(自动化系,清华大学) Department of Surgery, Chow Yuk Ho Technology Centre for Innovative Medicine, Li Ka Shing Institute of Health Science and Multi-scale Medical Robotics Center, The Chinese University of Hong Kong, Hong Kong(外科系,创新医学技术中心,利文斯科健康科学研究所及多尺度医学机器人中心,香港中文大学,香港)

专题命中 多模态生成 :multi-modal(abstract)

AI总结 UltraDP通过力感知扩散策略实现通用颈动脉超声扫描,利用多感官输入和混合控制器提高扫描成功率至95%。

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.19646 2025-11-20 cs.LG 50%

Energy-based generator matching: A neural sampler for general state space

Dongyeop Woo, Minsu Kim, Minkyu Kim, Kiyoung Seong, Sungsoo Ahn

机构 * Korea Advanced Institute of Science and Technology (KAIST)(韩国科学技术院)

专题命中 多模态生成 :multimodal(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.15206 2025-11-20 cs.CR cs.IT math.IT 50%

Trustworthy GenAI over 6G: Integrated Applications and Security Frameworks

Bui Duc Son, Trinh Van Chien, Dong In Kim

专题命中 多模态生成 :multimodal(abstract)

Comments 8 pages, 5 figures. Submitted for publication

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.15026 2025-11-20 eess.SP 50%

WiCo-MG: Wireless Channel Foundation Model for Multipath Generation via Synesthesia of Machines

Zengrui Han, Lu Bai, Xuesong Cai, Xiang Cheng

专题命中 多模态生成 :cross-modal(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.14516 2025-11-20 cs.LG 50%

Full-Atom Peptide Design via Riemannian-Euclidean Bayesian Flow Networks

Hao Qian, Shikui Tu, Lei Xu

专题命中 多模态生成 :multimodal(abstract)

Comments AAAI2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.14453 2025-11-19 q-bio.NC 50%

Multi-network Topology Underlying Individual Language Learning Success

Peilun Song, Shuguang Yang, Xiujuan Geng, Zhenzhong Gan, Suiping Wang, Gangyi Feng

专题命中 多模态生成 :multimodal(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.13186 2025-11-18 cs.LG cs.SY eess.SY 50%

DiffFP: Learning Behaviors from Scratch via Diffusion-based Fictitious Play

Akash Karthikeyan, Yash Vardhan Pant

机构 * Department of Electrical and Computer Engineering, University of Waterloo(滑铁卢大学电气与计算机工程系)

专题命中 多模态生成 :multimodal(abstract)

Comments Initial results presented at the IJCAI 2025 Workshop on User-Aligned Assessment of Adaptive AI Systems. Project page: https://aku02.github.io/projects/difffp/

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.03668 2025-11-12 cs.DC cs.LG 50%

Intelligent Orchestration of Distributed Large Foundation Model Inference at the Edge

Fernando Koch, Aladin Djuhera, Alecio Binotto

机构 * Florida Atlantic University, USA(佛罗里达大学) Technical University Munich, Germany(慕尼黑技术大学) Carl Zeiss AG, Germany(蔡司股份公司)

专题命中 多模态生成 :multi-modal(abstract)

Comments 26 pages, 3 figures, 4 tables, 52 references

Journal ref Computer Networks and Communications, 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.06892 2025-11-11 cs.RO 50%

Multi-Agent AI Framework for Road Situation Detection and C-ITS Message Generation

Kailin Tong, Selim Solmaz, Kenan Mujkic, Gottfried Allmer, Bo Leng

机构 * Virtual Vehicle Research GmbH(虚拟车辆研究有限公司) ASFINAG Maut Service GmbH(ASFINAG收费服务有限公司) Tongji University(同济大学)

专题命中 多模态生成 :multimodal(abstract)

Comments submitted to TRA 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.10684 2025-11-11 cs.LG math.OC stat.CO stat.ML 50%

MDNS: Masked Diffusion Neural Sampler via Stochastic Optimal Control

Yuchen Zhu, Wei Guo, Jaemoo Choi, Guan-Horng Liu, Yongxin Chen, Molei Tao

机构 * Georgia Institute of Technology(佐治亚理工学院) FAIR at Meta(Meta的FAIR)

专题命中 多模态生成 :multi-modal(abstract)

Comments NeurIPS 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.02530 2025-11-05 cs.AR 50%

Implementation and Evaluation of Stable Diffusion on a General-Purpose CGLA Accelerator

Takuto Ando, Yu Eto, Yasuhiko Nakashima

专题命中 多模态生成 :multi-modal(abstract)

Comments This paper is accepted at 2025 IEEE 18th International Symposium on Embedded Multicore/Many-core Systems-on-Chip (MCSoC)

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.15069 2025-11-04 stat.CO stat.ME stat.ML 50%

Sampling by averaging: A multiscale approach to score estimation

Paula Cordero-Encinar, Andrew B. Duncan, Sebastian Reich, O. Deniz Akyildiz

专题命中 多模态生成 :multimodal(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.21451 2025-10-31 cs.SE 50%

Scalpel: Automotive Deep Learning Framework Testing via Assembling Model Components

Yinglong Zou, Juan Zhai, Chunrong Fang, An Guo, Jiawei Liu, Zhenyu Chen

专题命中 多模态生成 :multi-modal(abstract)

Comments Accepted by the 48th IEEE/ACM International Conference on Software Engineering (ICSE 2026)

详情

展开后加载摘要…

URL PDF HTML 收藏
2412.20071 2025-10-30 cs.HC 50%

Towards Human-AI Synergy in UI Design: Supporting Iterative Generation with LLMs

Mingyue Yuan, Jieshan Chen, Yongquan Hu, Sidong Feng, Mulong Xie, Gelareh Mohammadi, Zhenchang Xing, Aaron Quigley

专题命中 多模态生成 :multi-modal(abstract)

Comments ACM Transactions on Computer-Human Interaction

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.24083 2025-10-29 math.OC 50%

A Novel Virus Diffusion Optimization (VDO) Algorithm for Global Optimization

Zhaoqi Sun, Qingsong Wang

专题命中 多模态生成 :multimodal(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.10881 2025-10-27 cs.LG 50%

Prior-Guided Diffusion Planning for Offline Reinforcement Learning

Donghyeon Ki, JunHyeok Oh, Seong-Woong Shim, Byung-Jun Lee

机构 * Korea University(韩国大学) Gauss Labs Inc.(Gauss实验室)

专题命中 多模态生成 :multi-modal(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2501.12407 2025-10-23 cs.DC cs.LG 50%

The Streaming Batch Model for Efficient and Fault-Tolerant Heterogeneous Execution

Frank Sifei Luan, Ron Yifeng Wang, Yile Gu, Ziming Mao, Charlotte Lin, Amog Kamsetty, Hao Chen, Cheng Su, Balaji Veeramani, Scott Lee, SangBin Cho, Clark Zinzow, Eric Liang, Ion Stoica, Stephanie Wang

机构 * UC Berkeley(加州大学伯克利分校) University of Washington(华盛顿大学) Anyscale

专题命中 多模态生成 :multimodal(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.14111 2025-10-21 cs.LG 50%

From AI for Science to Agentic Science: A Survey on Autonomous Scientific Discovery

Jiaqi Wei, Yuejin Yang, Xiang Zhang, Yuhan Chen, Xiang Zhuang, Zhangyang Gao, Dongzhan Zhou, Guangshuai Wang, Zhiqiang Gao, Juntai Cao, Zijie Qiu, Ming Hu, Chenglong Ma, Shixiang Tang, Junjun He, Chunfeng Song, Xuming He, Qiang Zhang, Chenyu You, Shuangjia Zheng, Ning Ding, Wanli Ouyang, Nanqing Dong, Yu Cheng, Siqi Sun, Lei Bai, Bowen Zhou

机构 * Shanghai Artificial Intelligence Laboratory(上海人工智能实验室) Zhejiang University(浙江大学) Fudan University(复旦大学) University of British Columbia(不列颠哥伦比亚大学) Tongji University(同济大学) The Chinese University of Hong Kong(香港中文大学) Shanghai Jiaotong University(上海交通大学) Stony Brook University(石溪大学) Lingang Laboratory(临港实验室) Tsinghua University(清华大学)

专题命中 多模态生成 :multimodal(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2411.16039 2025-10-20 physics.med-ph 50%

Label-Free Intraoperative Imaging of Hemodynamics using Deep Learning

Yan Shi, Denghui Zhao, Jingyi Yu, Wei Ni, Pengcheng Li, Yun Gu, Peng Miao, Shanbao Tong

专题命中 多模态生成 :cross-modal(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏