arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

多模态大模型

跨文本、图像、视频、音频等模态的大模型与学习方法。

共收录 4975 信号源:cs.CV, cs.CL, cs.AI, cs.MM, eess.AS

1. 多模态生成 4975 篇

2501.17726 2025-06-26 cs.CV cs.CL 62%

VICCA: Visual Interpretation and Comprehension of Chest X-ray Anomalies in Generated Report Without Human Feedback

Sayeh Gholipour Picha, Dawood Al Chanti, Alice Caplier

机构 * Univ. Grenoble Alpes(格勒诺布尔阿尔卑斯大学) CNRS(法国国家科学研究中心) Grenoble INP(格勒诺布尔研究所) GIPSA-lab(GIPSA实验室)

专题命中 多模态生成 :multimodal(abstract);分类 cs.CV、cs.CL

Journal ref Machine Learning with Applications, Volume 21, 2025, 100684, ISSN 2666-8270

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.18946 2025-06-25 cs.CV cs.AI 62%

DiffRIS: Enhancing Referring Remote Sensing Image Segmentation with Pre-trained Text-to-Image Diffusion Models

Zhe Dong, Yuzhe Sun, Tianzhu Liu, Yanfeng Gu

机构 * School of Electronics and Information Engineering, Harbin Institute of Technology(电子信息工程学院,哈尔滨工业大学)

专题命中 多模态生成 :cross-modal(abstract);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2407.09174 2025-06-24 cs.CV cs.AI 62%

DART: An Automated End-to-End Object Detection Pipeline with Data Diversification, Open-Vocabulary Bounding Box Annotation, Pseudo-Label Review, and Model Training

Chen Xin, Andreas Hartel, Enkelejda Kasneci

专题命中 多模态生成 :multimodal(abstract);分类 cs.CV、cs.AI

Comments Corrected minor typos; no changes to results or conclusions

Journal ref Expert Systems with Applications 258 (2024): 125124

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.17346 2025-06-24 cs.CV cs.AI 62%

A Novel Multi-layer Task-centric and Data Quality Framework for Autonomous Driving

Yuhan Zhou, Haihua Chen, Kewei Sha

机构 * Dept. of Information Science University of North Texas Denton, Texas, USA(信息科学系 诺克斯维尔大学 德顿 Texas USA)

专题命中 多模态生成 :multimodal(abstract);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2410.21264 2025-06-18 cs.CV cs.AI 62%

LARP: Tokenizing Videos with a Learned Autoregressive Generative Prior

Hanyu Wang, Saksham Suri, Yixuan Ren, Hao Chen, Abhinav Shrivastava

机构 * University of Maryland, College Park(马里兰大学 College Park 分校)

专题命中 多模态生成 :multimodal(abstract);分类 cs.CV、cs.AI

Comments ICLR 2025. Project page: https://hywang66.github.io/larp/

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.08106 2025-06-17 cs.LG cs.AI cs.CV stat.ML 62%

PoGDiff: Product-of-Gaussians Diffusion Models for Imbalanced Text-to-Image Generation

Ziyan Wang, Sizhe Wei, Xiaoming Huo, Hao Wang

机构 * Georgia Institute of Technology(佐治亚理工学院) Rutgers University(罗格斯大学)

专题命中 多模态生成 :image-text(abstract);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.10955 2025-06-13 cs.LG cs.AI cs.CV 62%

ReGuidance: A Simple Diffusion Wrapper for Boosting Sample Quality on Hard Inverse Problems

Aayush Karan, Kulin Shah, Sitan Chen

机构 * Harvard SEAS(哈佛大学SEAS) UT Austin(得克萨斯大学奥斯汀分校)

专题命中 多模态生成 :multimodal(abstract);分类 cs.CV、cs.AI

Comments 38 pages, 14 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.03603 2025-06-13 cs.CV cs.MM 62%

A Unit Enhancement and Guidance Framework for Audio-Driven Avatar Video Generation

S. Z. Zhou, Y. B. Wang, J. F. Wu, T. Hu, J. N. Zhang

机构 * Zhejiang University(浙江大学) Fudan University(复旦大学) Shanghai Jiao Tong University(上海交通大学)

专题命中 多模态生成 :audio-visual(abstract);分类 cs.CV、cs.MM

Comments revised

详情

展开后加载摘要…

URL PDF HTML 收藏
2407.17152 2025-06-12 cs.CV cs.AI 62%

XMeCap: Meme Caption Generation with Sub-Image Adaptability

Yuyan Chen, Songzhou Yan, Zhihong Zhu, Zhixu Li, Yanghua Xiao

机构 * Shanghai Key Laboratory of Data Science, School of Computer Science, Fudan University(复旦大学计算机学院数据科学实验室) Peking University(北京大学)

专题命中 多模态生成 :multi-modal(abstract);分类 cs.CV、cs.AI

Comments Accepted to ACM Multimedia 2024

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.08189 2025-06-11 cs.CV cs.CL 62%

Open World Scene Graph Generation using Vision Language Models

Amartya Dutta, Kazi Sajeed Mehrab, Medha Sawhney, Abhilash Neog, Mridul Khurana, Sepideh Fatemi, Aanish Pradhan, M. Maruf, Ismini Lourentzou, Arka Daw, Anuj Karpatne

专题命中 多模态生成 :multimodal(abstract);分类 cs.CV、cs.CL

Comments Accepted in CVPR 2025 Workshop (CVinW)

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.07235 2025-06-10 cs.CV cs.CL 62%

Multi-Step Visual Reasoning with Visual Tokens Scaling and Verification

Tianyi Bai, Zengjie Hu, Fupeng Sun, Jiantao Qiu, Yizhen Jiang, Guangxin He, Bohan Zeng, Conghui He, Binhang Yuan, Wentao Zhang

机构 * HKUST(香港科技大学) Peking University(北京大学) Shanghai AI Lab(上海人工智能实验室) Imperial College London(伦敦帝国理工学院)

专题命中 多模态生成 :multi-modal(abstract);分类 cs.CV、cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2406.05113 2025-06-09 cs.CV cs.AI cs.LG 62%

LlavaGuard: An Open VLM-based Framework for Safeguarding Vision Datasets and Models

Lukas Helff, Felix Friedrich, Manuel Brack, Kristian Kersting, Patrick Schramowski

专题命中 多模态生成 :multimodal(abstract);分类 cs.CV、cs.AI

Comments In Proceedings of the 42st International Conference on Machine Learning (ICML 2025), Project page at https://ml-research.github.io/human-centered-genai/projects/llavaguard/index.html

详情

展开后加载摘要…

URL PDF HTML 收藏
2412.01271 2025-06-06 cs.CL cs.AI 62%

MuLan: Adapting Multilingual Diffusion Models for Hundreds of Languages with Negligible Cost

Sen Xing, Muyan Zhong, Zeqiang Lai, Liangchen Li, Jiawen Liu, Yaohui Wang, Jifeng Dai, Wenhai Wang

专题命中 多模态生成 :image-text(abstract);分类 cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.03735 2025-06-05 cs.CL cs.AI cs.HC 62%

Generating Pedagogically Meaningful Visuals for Math Word Problems: A New Benchmark and Analysis of Text-to-Image Models

Junling Wang, Anna Rutkiewicz, April Yi Wang, Mrinmaya Sachan

专题命中 多模态生成 :multimodal(abstract);分类 cs.CL、cs.AI

Comments Findings of the Association for Computational Linguistics: ACL 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.03139 2025-06-04 cs.CV cs.AI 62%

SVGenius: Benchmarking LLMs in SVG Understanding, Editing and Generation

Siqi Chen, Xinyu Dong, Haolei Xu, Xingyu Wu, Fei Tang, Hang Zhang, Yuchen Yan, Linjuan Wu, Wenqi Zhang, Guiyang Hou, Yongliang Shen, Weiming Lu, Yueting Zhuang

机构 * Zhejiang University(浙江大学)

专题命中 多模态生成 :multimodal(abstract);分类 cs.CV、cs.AI

Comments 19 pages,4 figures, Project page: https://zju-real.github.io/SVGenius, Code: https://github.com/ZJU-REAL/SVGenius-Bench

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.02661 2025-06-04 cs.SD cs.CV cs.GR eess.AS 62%

MotionRAG-Diff: A Retrieval-Augmented Diffusion Framework for Long-Term Music-to-Dance Generation

Mingyang Huang, Peng Zhang, Bang Zhang

机构 * Tongyi Lab, Alibaba Group(通义实验室,阿里巴巴集团)

专题命中 多模态生成 :cross-modal(abstract);分类 cs.CV、eess.AS

Comments 12 pages, 5 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.01955 2025-06-03 cs.CV cs.CL cs.LG 62%

Dual-Process Image Generation

Grace Luo, Jonathan Granskog, Aleksander Holynski, Trevor Darrell

机构 * UC Berkeley(伯克利大学) Runway

专题命中 多模态生成 :multimodal(abstract);分类 cs.CV、cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.01689 2025-06-03 cs.AI cs.CL 62%

Respond Beyond Language: A Benchmark for Video Generation in Response to Realistic User Intents

Shuting Wang, Yunqi Liu, Zixin Yang, Ning Hu, Zhicheng Dou, Chenyan Xiong

机构 * School of Computer Science, Carnegie Mellon University(卡内基梅隆大学计算机科学学院) Gaoling School of Artificial Intelligence, Renmin University of China(中国人民大学北京校区人工智能学院) Serendipity One Inc.(Serendipity One公司)

专题命中 多模态生成 :multimodal(abstract);分类 cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.14318 2025-06-03 cs.CV cs.CL 62%

RADAR: Enhancing Radiology Report Generation with Supplementary Knowledge Injection

Wenjun Hou, Yi Cheng, Kaishuai Xu, Heng Li, Yan Hu, Wenjie Li, Jiang Liu

机构 * Department of Computing, The Hong Kong Polytechnic University(香港理工大学计算机系) Research Institute of Trustworthy Autonomous Systems and Department of Computer Science and Engineering, Southern University of Science and Technology(南方科技大学可信自主系统研究院和计算机科学与工程系) School of Computer Science, University of Nottingham Ningbo China(宁波大学计算机学院)

专题命中 多模态生成 :multimodal(abstract);分类 cs.CV、cs.CL

Comments Accepted to ACL 2025 main

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.07089 2025-06-03 cs.CV cs.CL 62%

OmniCaptioner: One Captioner to Rule Them All

Yiting Lu, Jiakang Yuan, Zhen Li, Shitian Zhao, Qi Qin, Xinyue Li, Le Zhuo, Licheng Wen, Dongyang Liu, Yuewen Cao, Xiangchao Yan, Xin Li, Tianshuo Peng, Shufei Zhang, Botian Shi, Tao Chen, Zhibo Chen, Lei Bai, Peng Gao, Bo Zhang

专题命中 多模态生成 :multimodal(abstract);分类 cs.CV、cs.CL

Comments More visualizations on Homepage: https://alpha-innovator.github.io/OmniCaptioner-project-page and Official code: https://github.com/Alpha-Innovator/OmniCaptioner

详情

展开后加载摘要…

URL PDF HTML 收藏
2411.05902 2025-06-03 cs.CV cs.CL 62%

Autoregressive Models in Vision: A Survey

Jing Xiong, Gongye Liu, Lun Huang, Chengyue Wu, Taiqiang Wu, Yao Mu, Yuan Yao, Hui Shen, Zhongwei Wan, Jinfa Huang, Chaofan Tao, Shen Yan, Huaxiu Yao, Lingpeng Kong, Hongxia Yang, Mi Zhang, Guillermo Sapiro, Jiebo Luo, Ping Luo, Ngai Wong

机构 * The University of Hong Kong(香港大学) Tsinghua University(清华大学) Duke University(杜克大学) University of Rochester(罗切斯特大学) The Ohio State University(俄亥俄州立大学) Bytedance(字节跳动) The University of North Carolina at Chapel Hill(北卡罗来纳大学教堂山分校) Apple(苹果公司) The Hong Kong Polytechnic University(香港理工大学) Princeton University(普林斯顿大学)

专题命中 多模态生成 :multimodal(abstract);分类 cs.CV、cs.CL

Comments The paper is accepted by TMLR

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.24245 2025-06-02 cs.CV cs.AI 62%

LTM3D: Bridging Token Spaces for Conditional 3D Generation with Auto-Regressive Diffusion Framework

Xin Kang, Zihan Zheng, Lei Chu, Yue Gao, Jiahao Li, Hao Pan, Xuejin Chen, Yan Lu

机构 * University of Science and Technology of China(中国科学技术大学) Microsoft Research Asia(微软亚洲研究院) Tsinghua University(清华大学)

专题命中 多模态生成 :multi-modal(abstract);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.22705 2025-05-30 cs.CV cs.MM 62%

HiDream-I1: A High-Efficient Image Generative Foundation Model with Sparse Diffusion Transformer

Qi Cai, Jingwen Chen, Yang Chen, Yehao Li, Fuchen Long, Yingwei Pan, Zhaofan Qiu, Yiheng Zhang, Fengbin Gao, Peihan Xu, Yimeng Wang, Kai Yu, Wenxuan Chen, Ziwei Feng, Zijian Gong, Jianzhuang Pan, Yi Peng, Rui Tian, Siyu Wang, Bo Zhao, Ting Yao, Tao Mei

专题命中 多模态生成 :multi-modal(abstract);分类 cs.CV、cs.MM

Comments Source codes and models are available at https://github.com/HiDream-ai/HiDream-I1 and https://github.com/HiDream-ai/HiDream-E1

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.22126 2025-05-29 cs.CV cs.AI 62%

SridBench: Benchmark of Scientific Research Illustration Drawing of Image Generation Model

Yifan Chang, Yukang Feng, Jianwen Sun, Jiaxin Ai, Chuanhao Li, S. Kevin Zhou, Kaipeng Zhang

机构 * University of Science and Technology of China(中国科学技术大学) Shanghai Innovation Institute(上海创新研究院) Nankai University(南开大学) Wuhan University(武汉大学) Shanghai AI Laboratory(上海人工智能实验室)

专题命中 多模态生成 :multimodal(abstract);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.19354 2025-05-27 cs.CL cs.CV 62%

GC-KBVQA: A New Four-Stage Framework for Enhancing Knowledge Based Visual Question Answering Performance

Mohammad Mahdi Moradi, Sudhir Mudur

机构 * Concordia University(康科迪亚大学)

专题命中 多模态生成 :multimodal(abstract);分类 cs.CV、cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2411.13981 2025-05-27 cs.CV cs.AI 62%

On the Fairness, Diversity and Reliability of Text-to-Image Generative Models

Jordan Vice, Naveed Akhtar, Leonid Sigal, Richard Hartley, Ajmal Mian

机构 * University of Western Australia(西澳大学) University of Melbourne(墨尔本大学) University of British Columbia(不列颠哥伦比亚大学) Australian National University(澳大利亚国立大学)

专题命中 多模态生成 :multimodal(abstract);分类 cs.CV、cs.AI

Comments This research is supported by the NISDRG project #20100007, funded by the Australian Government

详情

展开后加载摘要…

URL PDF HTML 收藏
2408.14744 2025-05-27 cs.CV cs.AI 62%

RSTeller: Scaling Up Visual Language Modeling in Remote Sensing with Rich Linguistic Semantics from Openly Available Data and Large Language Models

Junyao Ge, Xu Zhang, Yang Zheng, Kaitai Guo, Jimin Liang

机构 * School of Electronic Engineering, Xidian University(电子工程学院,西安电子科技大学)

专题命中 多模态生成 :multimodal(abstract);分类 cs.CV、cs.AI

Comments Published on ISPRS, minor typos corrected

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.17897 2025-05-26 cs.AI cs.CL 62%

T2I-Eval-R1: Reinforcement Learning-Driven Reasoning for Interpretable Text-to-Image Evaluation

Zi-Ao Ma, Tian Lan, Rong-Cheng Tu, Shu-Hang Liu, Heyan Huang, Zhijing Wu, Chen Xu, Xian-Ling Mao

专题命中 多模态生成 :multimodal(abstract);分类 cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.16976 2025-05-23 cs.CV cs.MM 62%

Creatively Upscaling Images with Global-Regional Priors

Yurui Qian, Qi Cai, Yingwei Pan, Ting Yao, Tao Mei

专题命中 多模态生成 :multimodal(abstract);分类 cs.CV、cs.MM

Comments International Journal of Computer Vision (IJCV) 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.16425 2025-05-23 cs.CL cs.AI 62%

$I^2G$: Generating Instructional Illustrations via Text-Conditioned Diffusion

Jing Bi, Pinxin Liu, Ali Vosoughi, Jiarui Wu, Jinxi He, Chenliang Xu

机构 * University of Rochester(罗切斯特大学)

专题命中 多模态生成 :multimodal(abstract);分类 cs.CL、cs.AI

Comments 13 pages, 5 figures, under review

详情

展开后加载摘要…

URL PDF HTML 收藏