arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

AI 大模型

多模态大模型

跨文本、图像、视频、音频等模态的大模型与学习方法。

2025-09-03 至 2025-09-03 共收录 18 信号源:cs.CV, cs.CL, cs.AI, cs.MM, eess.AS

1. 多模态生成 18 篇

2501.09012 2025-09-03 cs.CV cs.AI cs.CL cs.MM 83%

Multimodal LLMs Can Reason about Aesthetics in Zero-Shot

Ruixiang Jiang, Changwen Chen

机构 * The Hong Kong Polytechnic University(香港理工大学)

专题命中 多模态生成 :multimodal(title,abstract);分类 cs.CV、cs.CL、cs.AI

Comments ACM MM 2025 Camera Ready

详情

展开后加载摘要…

URL PDF HTML 收藏
2407.01074 2025-09-03 cs.CV cs.GR 79%

Multimodal Conditional 3D Face Geometry Generation

Christopher Otto, Prashanth Chandran, Sebastian Weiss, Markus Gross, Gaspard Zoss, Derek Bradley

机构 * ETH Zürich(苏黎世联邦理工学院) DisneyResearch | Studios(迪士尼研究室)

专题命中 多模态生成 :multimodal(title,abstract);分类 cs.CV

Comments Added more evaluation since the first version. Accepted to SMI 2025. Computers & Graphics

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.13602 2025-09-03 cs.CV 79%

PersonaVlog: Personalized Multimodal Vlog Generation with Multi-Agent Collaboration and Iterative Self-Correction

Xiaolu Hou, Bing Ma, Jiaxiang Cheng, Xuhua Ren, Kai Yu, Wenyue Li, Tianxiang Zheng, Qinglin Lu

专题命中 多模态生成 :multimodal(title,abstract);分类 cs.CV

Comments Project Page: https://personavlog-paper.github.io/

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.02518 2025-09-03 cs.LG 78%

AnalogCoder-Pro: Unifying Analog Circuit Generation and Optimization via Multi-modal LLMs

Yao Lai, Souradip Poddar, Sungyoung Lee, Guojin Chen, Mengkang Hu, Bei Yu, Ping Luo, David Z. Pan

专题命中 多模态生成 :multi-modal(title);multimodal(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.23149 2025-09-03 eess.IV 71%

Towards Interpretable Counterfactual Generation via Multimodal Autoregression

Chenglong Ma, Yuanfeng Ji, Jin Ye, Lu Zhang, Ying Chen, Tianbin Li, Mingjie Li, Junjun He, Hongming Shan

专题命中 多模态生成 :multimodal(title)

Comments MICCAI'25

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.02278 2025-09-03 cs.GR cs.AI cs.MM 62%

Think2Sing: Orchestrating Structured Motion Subtitles for Singing-Driven 3D Head Animation

Zikai Huang, Yihan Zhou, Xuemiao Xu, Cheng Xu, Xiaofen Xing, Jing Qin, Shengfeng He

机构 * School of Computer Science and Engineering, South China University of Technology(华南理工大学计算机科学与工程学院) South China University of Technology(华南理工大学) Guangdong Engineering Center for Large Model and GenAI Technology(广东省大模型与生成式人工智能技术工程中心) Guangdong Provincial Key Lab of Computational Intelligence and Cyberspace Information(广东省计算智能与网络信息重点实验室) Centre for Smart Health, Hong Kong Polytechnic University(香港理工大学智能健康研究中心) CAS-Hong Kong Joint Laboratory for Multimodal Medical Molecular Imaging(中国科学院-香港联合多模态医学分子成像联合实验室) School of Electronic and Information Engineering, South China University of Technology(华南理工大学电子与信息学院) Pazhou Lab, Guangzhou, Guangdong, China(广州琶洲实验室) School of Computing and Information Systems, Singapore Management University(新加坡国立大学计算机与信息系统学院)

专题命中 多模态生成 :multimodal(abstract);分类 cs.AI、cs.MM

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.01656 2025-09-03 cs.CV cs.CL 62%

Reinforced Visual Perception with Tools

Zetong Zhou, Dongping Chen, Zixian Ma, Zhihan Hu, Mingyang Fu, Sinan Wang, Yao Wan, Zhou Zhao, Ranjay Krishna

机构 * ONE Lab, HUST(华中科技大学 ONE 实验室) ONE Lab, HUST University of Maryland(华中科技大学 与 马里兰大学 ONE 实验室) University of Washington(华盛顿大学) Zhejiang University(浙江大学)

专题命中 多模态生成 :multi-modal(abstract);分类 cs.CV、cs.CL

Comments Technical Report

详情

展开后加载摘要…

URL PDF HTML 收藏
2403.12470 2025-09-03 cs.CV 57%

SC-Diff: 3D Shape Completion with Latent Diffusion Models

Simon Schaefer, Juan D. Galvis, Xingxing Zuo, Stefan Leutengger

机构 * Technical University of Munich(慕尼黑技术大学) Munich Center for Machine Learning (MCML)(慕尼黑机器学习中心) ETH Zurich(苏黎世联邦理工学院) MBZUAI(穆桑比克大学人工智能研究所)

专题命中 多模态生成 :multimodal(abstract);分类 cs.CV

Comments 13 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.00428 2025-09-03 cs.CV 57%

Mixture of Global and Local Experts with Diffusion Transformer for Controllable Face Generation

Xuechao Zou, Shun Zhang, Xing Fu, Yue Li, Kai Li, Yushe Cao, Congyan Lang, Pin Tao, Junliang Xing

机构 * Beijing Jiaotong University(北京交通大学) Ant Group(蚂蚁集团) Qinghai University(青海大学) Tsinghua University(清华大学)

专题命中 多模态生成 :multimodal(abstract);分类 cs.CV

Comments 14 pages, 11 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.18512 2025-09-03 physics.optics cs.CL 57%

Designing across domains with declarative thinking: Insights from the 96-Eyes ptychographic imager project

Antony C Chan

机构 * Consultant, high-throughput microscopy and hardware-accelerated algorithms(咨询顾问,高通量显微镜和硬件加速算法)

专题命中 多模态生成 :multi-modal(abstract);分类 cs.CL

Comments Minor changes: resolve HTML rendering issues of sideways tables; Code listing in dark mode. Cite three more journal articles

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.05304 2025-09-03 cs.LG cs.CV 57%

Gaussian Mixture Flow Matching Models

Hansheng Chen, Kai Zhang, Hao Tan, Zexiang Xu, Fujun Luan, Leonidas Guibas, Gordon Wetzstein, Sai Bi

机构 * Stanford University(斯坦福大学) Adobe Research(Adobe研究)

专题命中 多模态生成 :multi-modal(abstract);分类 cs.CV

Comments ICML 2025. Code: https://github.com/Lakonik/GMFlow

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.02551 2025-09-03 cs.NI cs.LG 50%

On Transferring, Merging, and Splitting Task-Oriented Network Digital Twins

Zifan Zhang, Minghong Fang, Mingzhe Chen, Yuchen Liu

机构 * Department of Computer Science, North Carolina State University(计算机科学系,北卡罗来纳州立大学) Department of Computer Science and Engineering, University of Louisville(计算机科学与工程系,路易斯维尔大学) Department of Electrical and Computer Engineering and Frost Institute for Data Science and Computing, University of Miami(电气与计算机工程系及弗罗斯特数据科学与计算研究所,迈阿密大学)

专题命中 多模态生成 :multi-modal(abstract)

Comments Accepted by IEEE MobiWac 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.01942 2025-09-03 stat.AP gr-qc 50%

Efficient Bayesian Sampling with Langevin Birth-Death Dynamics

Alex Leviyev, Francesco Iacovelli, Aaron Zimmerman

专题命中 多模态生成 :multi-modal(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.01819 2025-09-03 cs.RO 50%

ManiFlow: A General Robot Manipulation Policy via Consistency Flow Training

Ge Yan, Jiyue Zhu, Yuquan Deng, Shiqi Yang, Ri-Zhao Qiu, Xuxin Cheng, Marius Memmel, Ranjay Krishna, Ankit Goyal, Xiaolong Wang, Dieter Fox

机构 * University of Washington(华盛顿大学) UC San Diego(圣地亚哥大学) Nvidia(英伟达)

专题命中 多模态生成 :multi-modal(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.00753 2025-09-03 stat.ME cs.LG stat.AP stat.CO stat.ML 50%

FBMS: An R Package for Flexible Bayesian Model Selection and Model Averaging

Florian Frommlet, Jon Lachmann, Geir Storvik, Aliaksandr Hubin

专题命中 多模态生成 :multi-modal(abstract)

Comments 69 pages, 5 tables, 5 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.00098 2025-09-03 physics.ins-det cond-mat.mtrl-sci 50%

Operating advanced scientific instruments with AI agents that learn on the job

Aikaterini Vriza, Michael H. Prince, Tao Zhou, Henry Chan, Mathew J. Cherukara

专题命中 多模态生成 :multimodal(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.21660 2025-09-03 cs.LG 50%

PreGenie: An Agentic Framework for High-quality Visual Presentation Generation

Xiaojie Xu, Xinli Xu, Sirui Chen, Haoyu Chen, Fan Zhang, Ying-Cong Chen

机构 * The Hong Kong University of Science and Technology(Guangzhou)(香港科技大学(广州)) The Hong Kong University of Science and Technology(香港科技大学) Shanghai AI Laboratory(上海人工智能实验室)

专题命中 多模态生成 :multimodal(abstract)

Comments Accepted at EMNLP 2025, Findings

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.08034 2025-09-03 cs.OH 50%

Opportunities and Applications of GenAI in Smart Cities: A User-Centric Survey

Ankit Shetgaonkar, Dipen Pradhan, Lakshit Arora, Sanjay Surendranath Girija, Shashank Kapoor, Aman Raj

专题命中 多模态生成 :multimodal(abstract)

Comments Accepted in IEEE COINS 2025

Journal ref 2025 IEEE International Conference on Omni-layer Intelligent Systems (COINS)

详情

展开后加载摘要…

URL PDF HTML 收藏