arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

多模态大模型

跨文本、图像、视频、音频等模态的大模型与学习方法。

共收录 4959 信号源:cs.CV, cs.CL, cs.AI, cs.MM, eess.AS

1. 多模态生成 4959 篇

2411.06481 2025-04-17 cs.CV 57%

KMM: Key Frame Mask Mamba for Extended Motion Generation

Zeyu Zhang, Hang Gao, Akide Liu, Qi Chen, Feng Chen, Yiran Wang, Danning Li, Rui Zhao, Zhenming Li, Zhongwen Zhou, Hao Tang, Bohan Zhuang

专题命中 多模态生成 :multimodal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.11457 2025-04-16 cs.CV 57%

Aligning Generative Denoising with Discriminative Objectives Unleashes Diffusion for Visual Perception

Ziqi Pang, Xin Xu, Yu-Xiong Wang

专题命中 多模态生成 :multi-modal(abstract);分类 cs.CV

Comments ICLR 2025

Journal ref ICLR 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.10659 2025-04-16 cs.CV 57%

Relation-Rich Visual Document Generator for Visual Information Extraction

Zi-Han Jiang, Chien-Wei Lin, Wei-Hua Li, Hsuan-Tung Liu, Yi-Ren Yeh, Chu-Song Chen

专题命中 多模态生成 :multimodal(abstract);分类 cs.CV

Comments CVPR 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.10041 2025-04-15 cs.RO cs.CV 57%

Prior Does Matter: Visual Navigation via Denoising Diffusion Bridge Models

Hao Ren, Yiming Zeng, Zetong Bi, Zhaoliang Wan, Junlong Huang, Hui Cheng

专题命中 多模态生成 :multimodal(abstract);分类 cs.CV

Journal ref The IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.09219 2025-04-15 cs.SD eess.AS 57%

Generation of Musical Timbres using a Text-Guided Diffusion Model

Weixuan Yuan, Qadeer Khan, Vladimir Golkov

专题命中 多模态生成 :multi-modal(abstract);分类 eess.AS

Comments 10 pages, 5 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.08762 2025-04-15 cs.IR cs.AI 57%

InteractiveSurvey: An LLM-based Personalized and Interactive Survey Paper Generation System

Zhiyuan Wen, Jiannong Cao, Zian Wang, Beichen Guo, Ruosong Yang, Shuaiqi Liu

专题命中 多模态生成 :multi-modal(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.08003 2025-04-14 cs.CV 57%

Have we unified image generation and understanding yet? An empirical study of GPT-4o's image generation ability

Ning Li, Jingran Zhang, Justin Cui

专题命中 多模态生成 :multimodal(abstract);分类 cs.CV

Comments Early work, technical report

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.05979 2025-04-14 cs.CV 57%

An Empirical Study of GPT-4o Image Generation Capabilities

Sixiang Chen, Jinbin Bai, Zhuoran Zhao, Tian Ye, Qingyu Shi, Donghao Zhou, Wenhao Chai, Xin Lin, Jianzong Wu, Chao Tang, Shilin Xu, Tao Zhang, Haobo Yuan, Yikang Zhou, Wei Chow, Linfeng Li, Xiangtai Li, Lei Zhu, Lu Qi

专题命中 多模态生成 :multimodal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.07083 2025-04-11 cs.CV 57%

GenDoP: Auto-regressive Camera Trajectory Generation as a Director of Photography

Mengchen Zhang, Tong Wu, Jing Tan, Ziwei Liu, Gordon Wetzstein, Dahua Lin

专题命中 多模态生成 :multi-modal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.07424 2025-04-11 cs.AI 57%

Routing to the Right Expertise: A Trustworthy Judge for Instruction-based Image Editing

Chenxi Sun, Hongzhi Zhang, Qi Wang, Fuzheng Zhang

专题命中 多模态生成 :multimodal(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2411.19050 2025-04-11 cs.CV 57%

I Dream My Painting: Connecting MLLMs and Diffusion Models via Prompt Generation for Text-Guided Multi-Mask Inpainting

Nicola Fanelli, Gennaro Vessio, Giovanna Castellano

专题命中 多模态生成 :multimodal(abstract);分类 cs.CV

Comments Accepted at WACV 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.06841 2025-04-10 cs.CV 57%

Classifying the Unknown: In-Context Learning for Open-Vocabulary Text and Symbol Recognition

Tom Simon, William Mocaer, Pierrick Tranouez, Clement Chatelain, Thierry Paquet

专题命中 多模态生成 :multimodal(abstract);分类 cs.CV

Comments Submitted to ICDAR 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2409.16294 2025-04-10 cs.CV cs.GR cs.LG 57%

GenCAD: Image-Conditioned Computer-Aided Design Generation with Transformer-Based Contrastive Representation and Diffusion Priors

Md Ferdous Alam, Faez Ahmed

专题命中 多模态生成 :multi-modal(abstract);分类 cs.CV

Comments 24 pages, 13 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.05496 2025-04-09 cs.CL 57%

A Survey on Hypothesis Generation for Scientific Discovery in the Era of Large Language Models

Atilla Kaan Alkan, Shashwat Sourav, Maja Jablonska, Simone Astarita, Rishabh Chakrabarty, Nikhil Garuda, Pranav Khetarpal, Maciej Pióro, Dimitrios Tanoglidis, Kartheik G. Iyer, Mugdha S. Polimera, Michael J. Smith, Tirthankar Ghosal, Marc Huertas-Company, Sandor Kruk, Kevin Schawinski, Ioana Ciucă

专题命中 多模态生成 :multimodal(abstract);分类 cs.CL

Comments 9 pages (+2 pages of references), 2 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2410.20812 2025-04-09 cs.CV cs.LG eess.IV 57%

Fidelity-Imposed Displacement Editing for the Learn2Reg 2024 SHG-BF Challenge

Jiacheng Wang, Xiang Chen, Renjiu Hu, Rongguang Wang, Jiazheng Wang, Min Liu, Yaonan Wang, Hang Zhang

专题命中 多模态生成 :multi-modal(abstract);分类 cs.CV

Comments Accepted at IEEE ISBI 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2408.02349 2025-04-09 cs.LG cs.AI 57%

Toward Cost-efficient Adaptive Clinical Trials in Knee Osteoarthritis with Reinforcement Learning

Khanh Nguyen, Huy Hoang Nguyen, Egor Panfilov, Aleksei Tiulpin

专题命中 多模态生成 :multimodal(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.04842 2025-04-08 cs.CV 57%

FantasyTalking: Realistic Talking Portrait Generation via Coherent Motion Synthesis

Mengchao Wang, Qiang Wang, Fan Jiang, Yaqi Fan, Yunpeng Zhang, Yonggang Qi, Kun Zhao, Mu Xu

专题命中 多模态生成 :audio-visual(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.04185 2025-04-08 cs.CV 57%

SDEIT: Semantic-Driven Electrical Impedance Tomography

Dong Liu, Yuanchao Wu, Bowen Tong, Jiansong Deng

专题命中 多模态生成 :multimodal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.04010 2025-04-08 cs.CV cs.LG 57%

DiTaiListener: Controllable High Fidelity Listener Video Generation with Diffusion

Maksim Siniukov, Di Chang, Minh Tran, Hongkun Gong, Ashutosh Chaubey, Mohammad Soleymani

专题命中 多模态生成 :multimodal(abstract);分类 cs.CV

Comments Project page: https://havent-invented.github.io/DiTaiListener

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.02436 2025-04-04 cs.CV 57%

SkyReels-A2: Compose Anything in Video Diffusion Transformers

Zhengcong Fei, Debang Li, Di Qiu, Jiahua Wang, Yikun Dou, Rui Wang, Jingtao Xu, Mingyuan Fan, Guibin Chen, Yang Li, Yahui Zhou

专题命中 多模态生成 :image-text(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.01358 2025-04-04 math.OC cs.CV cs.NA math.NA 57%

Diffusion at Absolute Zero: Langevin Sampling Using Successive Moreau Envelopes [conference paper]

Andreas Habring, Alexander Falk, Thomas Pock

专题命中 多模态生成 :multi-modal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2412.17811 2025-04-04 cs.CV 57%

ChatGarment: Garment Estimation, Generation and Editing via Large Language Models

Siyuan Bian, Chenghao Xu, Yuliang Xiu, Artur Grigorev, Zhen Liu, Cewu Lu, Michael J. Black, Yao Feng

专题命中 多模态生成 :multimodal(abstract);分类 cs.CV

Comments CVPR 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.02160 2025-04-04 cs.CV cs.LG 57%

Less-to-More Generalization: Unlocking More Controllability by In-Context Generation

Shaojin Wu, Mengqi Huang, Wenxu Wu, Yufeng Cheng, Fei Ding, Qian He

专题命中 多模态生成 :cross-modal(abstract);分类 cs.CV

Comments Project page: https://bytedance.github.io/UNO Code and model: https://github.com/bytedance/UNO

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.01472 2025-04-03 cs.CV 57%

ANNEXE: Unified Analyzing, Answering, and Pixel Grounding for Egocentric Interaction

Yuejiao Su, Yi Wang, Qiongyang Hu, Chuang Yang, Lap-Pui Chau

专题命中 多模态生成 :multimodal(abstract);分类 cs.CV

Comments Computer Vision and Pattern Recognition

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.00394 2025-04-02 cs.CV 57%

AP-CAP: Advancing High-Quality Data Synthesis for Animal Pose Estimation via a Controllable Image Generation Pipeline

Lei Wang, Yujie Zhong, Xiaopeng Sun, Jingchun Cheng, Chengjian Feng, Qiong Cao, Lin Ma, Zhaoxin Fan

专题命中 多模态生成 :multi-modal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2410.04738 2025-04-02 cs.CV 57%

Diffusion Models in 3D Vision: A Survey

Zhen Wang, Dongyuan Li, Yaozu Wu, Tianyu He, Jiang Bian, Renhe Jiang

专题命中 多模态生成 :multimodal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2407.15240 2025-04-02 cs.CV 57%

BIGbench: A Unified Benchmark for Evaluating Multi-dimensional Social Biases in Text-to-Image Models

Hanjun Luo, Haoyu Huang, Ziye Deng, Xinfeng Li, Hewei Wang, Yingbin Jin, Yang Liu, Wenyuan Xu, Zuozhu Liu

专题命中 多模态生成 :multi-modal(abstract);分类 cs.CV

Comments arXiv admin note: substantial text overlap with arXiv:2405.17814

详情

展开后加载摘要…

URL PDF HTML 收藏
2311.16137 2025-04-02 cs.RO cs.AI 57%

A Graph-to-Text Approach to Knowledge-Grounded Response Generation in Human-Robot Interaction

Nicholas Thomas Walker, Stefan Ultes, Pierre Lison

专题命中 多模态生成 :multimodal(abstract);分类 cs.AI

Comments Submitted to Dialogue & Discourse 2023

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.23271 2025-04-01 cs.RO cs.AI 57%

Learning Coordinated Bimanual Manipulation Policies using State Diffusion and Inverse Dynamics Models

Haonan Chen, Jiaming Xu, Lily Sheng, Tianchen Ji, Shuijing Liu, Yunzhu Li, Katherine Driggs-Campbell

专题命中 多模态生成 :multimodal(abstract);分类 cs.AI

Comments Project Page: https://haonan16.github.io/coord_bimanual_page/. 12 pages, 12 figures, Accepted at ICRA 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.19839 2025-04-01 cs.CV 57%

FireEdit: Fine-grained Instruction-based Image Editing via Region-aware Vision Language Model

Jun Zhou, Jiahao Li, Zunnan Xu, Hanhui Li, Yiji Cheng, Fa-Ting Hong, Qin Lin, Qinglin Lu, Xiaodan Liang

专题命中 多模态生成 :cross-modal(abstract);分类 cs.CV

Comments Accepted to CVPR 2025

详情

展开后加载摘要…

URL PDF HTML 收藏