arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

多模态大模型

跨文本、图像、视频、音频等模态的大模型与学习方法。

共收录 4975 信号源:cs.CV, cs.CL, cs.AI, cs.MM, eess.AS

1. 多模态生成 4975 篇

2508.10684 2025-11-11 cs.LG math.OC stat.CO stat.ML 50%

MDNS: Masked Diffusion Neural Sampler via Stochastic Optimal Control

Yuchen Zhu, Wei Guo, Jaemoo Choi, Guan-Horng Liu, Yongxin Chen, Molei Tao

机构 * Georgia Institute of Technology(佐治亚理工学院) FAIR at Meta(Meta的FAIR)

专题命中 多模态生成 :multi-modal(abstract)

Comments NeurIPS 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.02530 2025-11-05 cs.AR 50%

Implementation and Evaluation of Stable Diffusion on a General-Purpose CGLA Accelerator

Takuto Ando, Yu Eto, Yasuhiko Nakashima

专题命中 多模态生成 :multi-modal(abstract)

Comments This paper is accepted at 2025 IEEE 18th International Symposium on Embedded Multicore/Many-core Systems-on-Chip (MCSoC)

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.15069 2025-11-04 stat.CO stat.ME stat.ML 50%

Sampling by averaging: A multiscale approach to score estimation

Paula Cordero-Encinar, Andrew B. Duncan, Sebastian Reich, O. Deniz Akyildiz

专题命中 多模态生成 :multimodal(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.21451 2025-10-31 cs.SE 50%

Scalpel: Automotive Deep Learning Framework Testing via Assembling Model Components

Yinglong Zou, Juan Zhai, Chunrong Fang, An Guo, Jiawei Liu, Zhenyu Chen

专题命中 多模态生成 :multi-modal(abstract)

Comments Accepted by the 48th IEEE/ACM International Conference on Software Engineering (ICSE 2026)

详情

展开后加载摘要…

URL PDF HTML 收藏
2412.20071 2025-10-30 cs.HC 50%

Towards Human-AI Synergy in UI Design: Supporting Iterative Generation with LLMs

Mingyue Yuan, Jieshan Chen, Yongquan Hu, Sidong Feng, Mulong Xie, Gelareh Mohammadi, Zhenchang Xing, Aaron Quigley

专题命中 多模态生成 :multi-modal(abstract)

Comments ACM Transactions on Computer-Human Interaction

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.24083 2025-10-29 math.OC 50%

A Novel Virus Diffusion Optimization (VDO) Algorithm for Global Optimization

Zhaoqi Sun, Qingsong Wang

专题命中 多模态生成 :multimodal(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.10881 2025-10-27 cs.LG 50%

Prior-Guided Diffusion Planning for Offline Reinforcement Learning

Donghyeon Ki, JunHyeok Oh, Seong-Woong Shim, Byung-Jun Lee

机构 * Korea University(韩国大学) Gauss Labs Inc.(Gauss实验室)

专题命中 多模态生成 :multi-modal(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2501.12407 2025-10-23 cs.DC cs.LG 50%

The Streaming Batch Model for Efficient and Fault-Tolerant Heterogeneous Execution

Frank Sifei Luan, Ron Yifeng Wang, Yile Gu, Ziming Mao, Charlotte Lin, Amog Kamsetty, Hao Chen, Cheng Su, Balaji Veeramani, Scott Lee, SangBin Cho, Clark Zinzow, Eric Liang, Ion Stoica, Stephanie Wang

机构 * UC Berkeley(加州大学伯克利分校) University of Washington(华盛顿大学) Anyscale

专题命中 多模态生成 :multimodal(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.14111 2025-10-21 cs.LG 50%

From AI for Science to Agentic Science: A Survey on Autonomous Scientific Discovery

Jiaqi Wei, Yuejin Yang, Xiang Zhang, Yuhan Chen, Xiang Zhuang, Zhangyang Gao, Dongzhan Zhou, Guangshuai Wang, Zhiqiang Gao, Juntai Cao, Zijie Qiu, Ming Hu, Chenglong Ma, Shixiang Tang, Junjun He, Chunfeng Song, Xuming He, Qiang Zhang, Chenyu You, Shuangjia Zheng, Ning Ding, Wanli Ouyang, Nanqing Dong, Yu Cheng, Siqi Sun, Lei Bai, Bowen Zhou

机构 * Shanghai Artificial Intelligence Laboratory(上海人工智能实验室) Zhejiang University(浙江大学) Fudan University(复旦大学) University of British Columbia(不列颠哥伦比亚大学) Tongji University(同济大学) The Chinese University of Hong Kong(香港中文大学) Shanghai Jiaotong University(上海交通大学) Stony Brook University(石溪大学) Lingang Laboratory(临港实验室) Tsinghua University(清华大学)

专题命中 多模态生成 :multimodal(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2411.16039 2025-10-20 physics.med-ph 50%

Label-Free Intraoperative Imaging of Hemodynamics using Deep Learning

Yan Shi, Denghui Zhao, Jingyi Yu, Wei Ni, Pengcheng Li, Yun Gu, Peng Miao, Shanbao Tong

专题命中 多模态生成 :cross-modal(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.07880 2025-10-15 cs.NI eess.SP 50%

Generative Resource Allocation for 6G O-RAN with Diffusion Policies

Salar Nouri, Mojdeh Karbalaeimotaleb, Vahid Shah-Mansouri, Tarik Taleb

专题命中 多模态生成 :multi-modal(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.11138 2025-10-14 cs.SE 50%

What Slows Down FMware Development? An Empirical Study of Developer Challenges and Resolution Times

Zitao Wang, Zhimin Zhao, Michael W. Godfrey

专题命中 多模态生成 :multi-modal(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.15567 2025-10-14 cs.LG 50%

Towards Unified and Lossless Latent Space for 3D Molecular Latent Diffusion Modeling

Yanchen Luo, Zhiyuan Liu, Yi Zhao, Sihang Li, Hengxing Cai, Kenji Kawaguchi, Tat-Seng Chua, Yang Zhang, Xiang Wang

机构 * University of Science and Technology of China(中国科学技术大学) National University of Singapore(新加坡国立大学) DP Technology(DP技术)

专题命中 多模态生成 :multi-modal(abstract)

Comments NeurIPS 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2405.15775 2025-10-14 cs.RO 50%

Humanoid Robots and Humanoid AI: Review, Perspectives and Directions

Longbing Cao

机构 * Frontier AI Research Centre, Macquarie University(前沿人工智能研究中心,麦考瑞大学)

专题命中 多模态生成 :multimodal(abstract)

Comments 35 pages, 3 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.09134 2025-10-13 cs.SE cs.ET 50%

A Semantic Framework for Patient Digital Twins in Chronic Care

Amal Elgammal, Bernd J. Krämer, Michael P. Papazoglou, Mira Raheem

专题命中 多模态生成 :multimodal(abstract)

Comments This manuscript is currently under review at Software and Systems Modeling (SoSyM)

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.08912 2025-10-13 cs.HC 50%

Beyond Words: Infusing Conversational Agents with Human-like Typing Behaviors

Jijie Zhou, Yuhan Hu

专题命中 多模态生成 :multimodal(abstract)

Comments Author's version of a paper published at CUI '24 (ACM Conversational User Interfaces 2024)

Journal ref CUI '24: Proceedings of the ACM Conversational User Interfaces 2024, July 8-10, 2024, Luxembourg, Luxembourg. ACM, New York, NY, USA, 11 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.22258 2025-10-10 math.OC 50%

Diffusion at Absolute Zero: Langevin Sampling using Successive Moreau Envelopes [journal paper]

Andreas Habring, Alexander Falk, Martin Zach, Thomas Pock

专题命中 多模态生成 :multi-modal(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.07627 2025-10-09 cs.CE q-bio.BM 50%

LSMTCR: A Scalable Multi-Architecture Model for Epitope-Specific T Cell Receptor de novo Design

Ruihao Zhang, Xiao Liu

专题命中 多模态生成 :cross-modal(abstract)

Comments 13 main pages, 5 figures, 2 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.16876 2025-10-07 cs.SE 50%

Revolutionizing Validation and Verification: Explainable Testing Methodologies for Intelligent Automotive Decision-Making Systems

Halit Eris, Stefan Wagner

专题命中 多模态生成 :multimodal(abstract)

Comments Preprint to be published at SE4ADS

Journal ref 2025 IEEE/ACM 1st International Workshop on Software Engineering for Autonomous Driving Systems (SE4ADS), Ottawa, ON, Canada, 2025, pp. 34-37

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.03397 2025-10-07 hep-ph 50%

Foundation models for equation discovery in high energy physics

Manuel Morales-Alvarado

专题命中 多模态生成 :multimodal(abstract)

Comments 6 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.04584 2025-10-03 cs.HC 50%

SlideItRight: Using AI to Find Relevant Slides and Provide Feedback for Open-Ended Questions

Chloe Qianhui Zhao, Jie Cao, Eason Chen, Kenneth R. Koedinger, Jionghao Lin

专题命中 多模态生成 :multimodal(abstract)

Comments 14 pages, to be published at the 26th International Conference on Artificial Intelligence in Education (AIED '25)

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.00351 2025-10-02 cs.LG q-bio.BM 50%

Flow Autoencoders are Effective Protein Tokenizers

Rohit Dilip, Evan Zhang, Ayush Varshney, David Van Valen

机构 * California Institute of Technology(加州理工学院) OpenAI(开放人工智能公司) Howard Hughes Medical Institute(霍华德·休斯医学研究院)

专题命中 多模态生成 :multimodal(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.06266 2025-10-01 cs.RO 50%

ADPro: a Test-time Adaptive Diffusion Policy via Manifold-constrained Denoising and Task-aware Initialization for Robotic Manipulation

Zezeng Li, Rui Yang, Ruochen Chen, ZhongXuan Luo, Liming Chen

机构 * École Centrale de Lyon(里昂中央理工大学) Dalian University of Technology(大连理工大学)

专题命中 多模态生成 :multi-modal(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.02353 2025-10-01 cs.RO cs.LG 50%

Controllable Motion Generation via Diffusion Modal Coupling

Luobin Wang, Hongzhan Yu, Chenning Yu, Sicun Gao, Henrik Christensen

机构 * University of California, San Diego(加州大学圣地亚哥分校)

专题命中 多模态生成 :multi-modal(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.17608 2025-09-30 cs.LG 50%

ChartMaster: Advancing Chart-to-Code Generation with Real-World Charts and Chart Similarity Reinforcement Learning

Wentao Tan, Qiong Cao, Chao Xue, Yibing Zhan, Changxing Ding, Xiaodong He

机构 * South China University of Technology(华南理工大学) JD Future Academy(京东未来学院) Wuhan University(武汉大学)

专题命中 多模态生成 :multimodal(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.20266 2025-09-30 cond-mat.mtrl-sci 50%

Emergent properties and the multiscale characterization challenge in condensed matter, from crystals to complex materials: a Review

Elisabetta Nocerino

专题命中 多模态生成 :multimodal(abstract)

Journal ref J. Phys. D: Appl. Phys. 58 393001 (2025)

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.21664 2025-09-29 cs.RO cs.LG 50%

Generating Stable Placements via Physics-guided Diffusion Models

Philippe Nadeau, Miguel Rogel, Ivan Bilić, Ivan Petrović, Jonathan Kelly

机构 * STARS Laboratory, University of Toronto Institute for Aerospace Studies(多伦多大学航空航天研究所STARS实验室) Laboratory for Autonomous Systems and Mobile Robotics(自主系统与移动机器人实验室) University of Zagreb Faculty of Electrical Engineering and Computing(Zagreb大学电气工程与计算学院)

专题命中 多模态生成 :multi-modal(abstract)

Comments Submitted to the IEEE International Conference on Robotics and Automation 2026, Vienna, Austria, June 1-5, 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.21129 2025-09-26 cs.LG cs.CR 50%

EvoMail: Self-Evolving Cognitive Agents for Adaptive Spam and Phishing Email Defense

Wei Huang, De-Tian Chu, Lin-Yuan Bai, Wei Kang, Hai-Tao Zhang, Bo Li, Zhi-Mo Han, Jing Ge, Hai-Feng Lin

机构 * People’s Liberation Army 77606 Unit(中国人民解放军77606单位) Field Engineering College, Army Engineering University of PLA(陆军工程大学工程学院) College Of Software, ZhengZhou University of Light Industry(郑州轻工业大学软件学院)

专题命中 多模态生成 :multi-modal(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.21022 2025-09-26 cs.LG 50%

Actor-Critic without Actor

Donghyeon Ki, Hee-Jun Ahn, Kyungyoon Kim, Byung-Jun Lee

机构 * Korea University(韩国大学) Gauss Labs Inc.(Gauss实验室)

专题命中 多模态生成 :multi-modal(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.20136 2025-09-25 cs.SE 50%

V-GameGym: Visual Game Generation for Code Large Language Models

Wei Zhang, Jack Yang, Renshuai Tao, Lingzheng Chai, Shawn Guo, Jiajun Wu, Xiaoming Chen, Ganqu Cui, Ning Ding, Xander Xu, Hu Wei, Bowen Zhou

专题命中 多模态生成 :multimodal(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏