arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

多模态大模型

跨文本、图像、视频、音频等模态的大模型与学习方法。

共收录 4975 信号源:cs.CV, cs.CL, cs.AI, cs.MM, eess.AS

1. 多模态生成 4975 篇

2310.08541 2024-08-15 cs.CV 70%

Idea2Img: Iterative Self-Refinement with GPT-4V(ision) for Automatic Image Design and Generation

Zhengyuan Yang, Jianfeng Wang, Linjie Li, Kevin Lin, Chung-Ching Lin, Zicheng Liu, Lijuan Wang

专题命中 多模态生成 :multimodal(abstract);image-text(abstract);分类 cs.CV

Comments ECCV 2024; Project page at https://idea2img.github.io/

详情

展开后加载摘要…

URL PDF HTML 收藏
2407.12435 2024-07-18 cs.CV 70%

F-HOI: Toward Fine-grained Semantic-Aligned 3D Human-Object Interactions

Jie Yang, Xuesong Niu, Nan Jiang, Ruimao Zhang, Siyuan Huang

专题命中 多模态生成 :multimodal(abstract);multi-modal(abstract);分类 cs.CV

Comments ECCV24

详情

展开后加载摘要…

URL PDF HTML 收藏
2407.07614 2024-07-12 cs.CV 70%

MARS: Mixture of Auto-Regressive Models for Fine-grained Text-to-image Synthesis

Wanggui He, Siming Fu, Mushui Liu, Xierui Wang, Wenyi Xiao, Fangxun Shu, Yi Wang, Lei Zhang, Zhelun Yu, Haoyuan Li, Ziwei Huang, LeiLei Gan, Hao Jiang

专题命中 多模态生成 :image-text(abstract);any-to-any(abstract);分类 cs.CV

Comments 14 pages, 9 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2403.06801 2024-07-08 eess.IV cs.CV 70%

CT2Rep: Automated Radiology Report Generation for 3D Medical Imaging

Ibrahim Ethem Hamamci, Sezgin Er, Bjoern Menze

专题命中 多模态生成 :multimodal(abstract);multi-modal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2405.04940 2024-07-02 cs.CV 70%

Harnessing the Power of MLLMs for Transferable Text-to-Image Person ReID

Wentao Tan, Changxing Ding, Jiayu Jiang, Fei Wang, Yibing Zhan, Dapeng Tao

专题命中 多模态生成 :multi-modal(abstract);MLLM(abstract);分类 cs.CV

Comments CVPR 2024

详情

展开后加载摘要…

URL PDF HTML 收藏
2404.01291 2024-06-19 cs.CV cs.AI cs.CL cs.LG cs.MM 70%

Evaluating Text-to-Visual Generation with Image-to-Text Generation

Zhiqiu Lin, Deepak Pathak, Baiqi Li, Jiayao Li, Xide Xia, Graham Neubig, Pengchuan Zhang, Deva Ramanan

专题命中 多模态生成 :image-text(abstract);分类 cs.CV、cs.CL、cs.AI

Comments We open-source our data, model, and code at: https://github.com/linzhiqiu/t2v_metrics ; Project page: https://linzhiqiu.github.io/papers/vqascore

详情

展开后加载摘要…

URL PDF HTML 收藏
2401.14688 2024-06-19 cs.CL 70%

Taiyi-Diffusion-XL: Advancing Bilingual Text-to-Image Generation with Large Vision-Language Model Support

Xiaojun Wu, Dixiang Zhang, Ruyi Gan, Junyu Lu, Ziwei Wu, Renliang Sun, Jiaxing Zhang, Pingjian Zhang, Yan Song

专题命中 多模态生成 :multimodal(abstract);image-text(abstract);分类 cs.CL

Comments Taiyi-Diffusion-XL Tech Report

详情

展开后加载摘要…

URL PDF HTML 收藏
2406.06525 2024-06-11 cs.CV 70%

Autoregressive Model Beats Diffusion: Llama for Scalable Image Generation

Peize Sun, Yi Jiang, Shoufa Chen, Shilong Zhang, Bingyue Peng, Ping Luo, Zehuan Yuan

专题命中 多模态生成 :multimodal(abstract);multimodal foundation model(abstract);分类 cs.CV

Comments Codes and models: \url{https://github.com/FoundationVision/LlamaGen}

详情

展开后加载摘要…

URL PDF HTML 收藏
2405.04007 2024-05-08 cs.CV 70%

SEED-Data-Edit Technical Report: A Hybrid Dataset for Instructional Image Editing

Yuying Ge, Sijie Zhao, Chen Li, Yixiao Ge, Ying Shan

专题命中 多模态生成 :multimodal(abstract);MLLM(abstract);分类 cs.CV

Comments Technical Report; Dataset released in https://huggingface.co/datasets/AILab-CVC/SEED-Data-Edit

详情

展开后加载摘要…

URL PDF HTML 收藏
2404.13766 2024-04-23 cs.CV 70%

Object-Attribute Binding in Text-to-Image Generation: Evaluation and Control

Maria Mihaela Trusca, Wolf Nuyts, Jonathan Thomm, Robert Honig, Thomas Hofmann, Tinne Tuytelaars, Marie-Francine Moens

专题命中 多模态生成 :multimodal(abstract);image-text(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2404.09516 2024-04-16 cs.LG cs.AI cs.CL cs.CV cs.MM 70%

State Space Model for New-Generation Network Alternative to Transformers: A Survey

Xiao Wang, Shiao Wang, Yuhe Ding, Yuehang Li, Wentao Wu, Yao Rong, Weizhe Kong, Ju Huang, Shihao Li, Haoxiang Yang, Ziwen Wang, Bo Jiang, Chenglong Li, Yaowei Wang, Yonghong Tian, Jin Tang

专题命中 多模态生成 :multi-modal(abstract);分类 cs.CV、cs.CL、cs.AI

Comments The First review of State Space Model (SSM)/Mamba and their applications in artificial intelligence, 33 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2312.13537 2023-12-22 cs.CV 70%

HyperEditor: Achieving Both Authenticity and Cross-Domain Capability in Image Editing via Hypernetworks

Hai Zhang, Chunwei Wu, Guitao Cao, Hailing Wang, Wenming Cao

专题命中 多模态生成 :cross-modal(abstract);image-text(abstract);分类 cs.CV

Comments Accepted by AAAI2024

详情

展开后加载摘要…

URL PDF HTML 收藏
2312.09251 2023-12-15 cs.CV 70%

VL-GPT: A Generative Pre-trained Transformer for Vision and Language Understanding and Generation

Jinguo Zhu, Xiaohan Ding, Yixiao Ge, Yuying Ge, Sijie Zhao, Hengshuang Zhao, Xiaohua Wang, Ying Shan

专题命中 多模态生成 :multimodal(abstract);image-text(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2311.00571 2023-11-02 cs.CV cs.AI cs.CL cs.HC cs.MM 70%

LLaVA-Interactive: An All-in-One Demo for Image Chat, Segmentation, Generation and Editing

Wei-Ge Chen, Irina Spiridonova, Jianwei Yang, Jianfeng Gao, Chunyuan Li

专题命中 多模态生成 :multimodal(abstract);分类 cs.CV、cs.CL、cs.AI

Comments 31 pages, 22 figures, 30M PDF file size; Project Page: https://llava-vl.github.io/llava-interactive/

详情

展开后加载摘要…

URL PDF HTML 收藏
2310.11513 2023-10-19 cs.CV cs.LG 70%

GenEval: An Object-Focused Framework for Evaluating Text-to-Image Alignment

Dhruba Ghosh, Hanna Hajishirzi, Ludwig Schmidt

专题命中 多模态生成 :multimodal(abstract);image-text(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2309.16206 2023-10-06 eess.IV cs.CV 70%

Alzheimer's Disease Prediction via Brain Structural-Functional Deep Fusing Network

Qiankun Zuo, Junren Pan, Shuqiang Wang

专题命中 多模态生成 :multimodal(abstract);cross-modal(abstract);分类 cs.CV

Comments 10 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2309.00615 2023-09-04 cs.CV cs.AI cs.CL cs.LG cs.MM 70%

Point-Bind & Point-LLM: Aligning Point Cloud with Multi-modality for 3D Understanding, Generation, and Instruction Following

Ziyu Guo, Renrui Zhang, Xiangyang Zhu, Yiwen Tang, Xianzheng Ma, Jiaming Han, Kexin Chen, Peng Gao, Xianzhi Li, Hongsheng Li, Pheng-Ann Heng

专题命中 多模态生成 :multi-modal(abstract);分类 cs.CV、cs.CL、cs.AI

Comments Work in progress. Code is available at https://github.com/ZiyuGuo99/Point-Bind_Point-LLM

详情

展开后加载摘要…

URL PDF HTML 收藏
2307.08041 2023-08-15 cs.CV 70%

Planting a SEED of Vision in Large Language Model

Yuying Ge, Yixiao Ge, Ziyun Zeng, Xintao Wang, Ying Shan

专题命中 多模态生成 :multimodal(abstract);image-text(abstract);分类 cs.CV

Comments Technical Report; Project released at: https://github.com/AILab-CVC/SEED

详情

展开后加载摘要…

URL PDF HTML 收藏
2204.03039 2022-07-22 cs.CV 70%

DSGN++: Exploiting Visual-Spatial Relation for Stereo-based 3D Detectors

Yilun Chen, Shijia Huang, Shu Liu, Bei Yu, Jiaya Jia

专题命中 多模态生成 :multi-modal(abstract);cross-modal(abstract);分类 cs.CV

Comments 13 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2204.02035 2022-04-06 cs.CV 70%

DT2I: Dense Text-to-Image Generation from Region Descriptions

Stanislav Frolov, Prateek Bansal, Jörn Hees, Andreas Dengel

专题命中 多模态生成 :multi-modal(abstract);image-text(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2111.13792 2022-03-25 cs.CV cs.LG 70%

LAFITE: Towards Language-Free Training for Text-to-Image Generation

Yufan Zhou, Ruiyi Zhang, Changyou Chen, Chunyuan Li, Chris Tensmeyer, Tong Yu, Jiuxiang Gu, Jinhui Xu, Tong Sun

专题命中 多模态生成 :multi-modal(abstract);image-text(abstract);分类 cs.CV

Comments Accepted by CVPR 2022, https://github.com/drboog/Lafite

详情

展开后加载摘要…

URL PDF HTML 收藏
2001.08779 2020-01-27 cs.CV cs.AI cs.CL cs.LG cs.MM 70%

Deep Bayesian Network for Visual Question Generation

Badri N. Patro, Vinod K. Kurmi, Sandeep Kumar, Vinay P. Namboodiri

专题命中 多模态生成 :multimodal(abstract);分类 cs.CV、cs.CL、cs.AI

Comments WACV-2020 (Accepted)

详情

展开后加载摘要…

URL PDF HTML 收藏
1711.10485 2017-11-30 cs.CV 70%

AttnGAN: Fine-Grained Text to Image Generation with Attentional Generative Adversarial Networks

Tao Xu, Pengchuan Zhang, Qiuyuan Huang, Han Zhang, Zhe Gan, Xiaolei Huang, Xiaodong He

专题命中 多模态生成 :multimodal(abstract);image-text(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
1307.6549 2013-07-25 cs.CV cs.GR math.SP 70%

Making Laplacians commute

Michael M. Bronstein, Klaus Glashoff, Terry A. Loring

专题命中 多模态生成 :multimodal(abstract);multi-modal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2305.14882 2024-04-16 cs.CL cs.AI cs.CV 69%

Dynamic Clue Bottlenecks: Towards Interpretable-by-Design Visual Question Answering

Xingyu Fu, Ben Zhou, Sihao Chen, Mark Yatskar, Dan Roth

专题命中 多模态生成 :multimodal(abstract,comments);分类 cs.CV、cs.CL、cs.AI

Comments Multimodal, Visual Question Answering, Vision and Language

详情

展开后加载摘要…

URL PDF HTML 收藏
2012.03308 2021-03-30 cs.CV cs.AI cs.MM 69%

TediGAN: Text-Guided Diverse Face Image Generation and Manipulation

Weihao Xia, Yujiu Yang, Jing-Hao Xue, Baoyuan Wu

专题命中 多模态生成 :multi-modal(abstract,comments);分类 cs.CV、cs.AI、cs.MM

Comments CVPR 2021. Code: https://github.com/weihaox/TediGAN Data: https://github.com/weihaox/Multi-Modal-CelebA-HQ Video: https://youtu.be/L8Na2f5viAM

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.20331 2026-08-21 cs.CL cs.AI cs.CV 新提交 67%

G-CARL: Grounded Checklist-Aligned Reward Learning for Patient-Oriented Medical Report Interpretation

G-CARL:面向患者的医学报告解读的基于 grounded 核对清单对齐的奖励学习

Shiao Xie, Siyu Chen, Jianwei Lv, Bo Yuan, Yujin Wang, Xiandong Li

专题命中 多模态生成 :multimodal(abstract);分类 cs.CV、cs.CL、cs.AI

AI总结 本文针对现有医学视觉-语言任务无法兼顾医学事实性与患者语境沟通的问题,提出PMRI任务及G-CARL框架,构建MMedReport基准,实验证实其解读更贴合患者需求。

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.18279 2026-08-20 physics.optics cs.LG 新提交 67%

A Comprehensive Review of Large Language Models for Nanophotonics: From Surrogate Modeling to Autonomous Design

面向纳米光子学的大语言模型综合综述:从代理建模到自主设计

Huanshu Zhang, Kegeng Tang, Lei Kang, Sawyer D. Campbell, Zihao Wang, Douglas H. Werner

专题命中 多模态生成 :multimodal(abstract);multimodal foundation model(abstract)

AI总结 该综述探讨大语言模型(LLMs)如何通过语义接口、代码生成及工具编排改进纳米光子学工作流,梳理相关方法的两类模式及跨学科应用,展望具备物理感知的多模态基础模型,推动AI从被动工具向主动科研合作者转变。

Comments Accepted for publication in Advanced Photonics

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.13560 2026-08-14 cs.CV cs.AI cs.CL 新提交 67%

AutoDesign: Meta-Harness Optimization for Long-Horizon Agentic Design

AutoDesign:面向长视距智能体设计的元工具优化

Yaxin Luo, Haobin Jiang, Jialv Zou, Xu Huang, Wenhao Yan, Haodong Li, Zhengrong Yue, Jing Li, Xiaofu Chen, Xiaohan Zhao, Jiacheng Liu, Jiacheng Cui, Zhiqiang Shen, Xiaotong Li

机构 * Meituan(美团) MBZUAI(Mohamed bin Zayed University of Artificial Intelligence) Huazhong University of Science and Technology(华中科技大学) Peking University(北京大学) Tsinghua University(清华大学) The Chinese University of Hong Kong(香港中文大学) Shanghai Jiao Tong University(上海交通大学)

专题命中 多模态生成 :multimodal(abstract);分类 cs.CV、cs.CL、cs.AI

AI总结 AutoDesign是符合人类设计先验的元工具优化框架,以论文转海报生成任务为实例,在PosterBench上性能优于Claude Design,集成其学习的DesignHarness可提升代码智能体性能,且获人类最高偏好。

Comments Tech Report. Code at: https://github.com/Yaxin9Luo/AutoDesign

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.02833 2026-08-05 cs.CV cs.AI cs.CL 新提交 67%

CURV: Enhancing Chart Understanding Through Curriculum Visual Grounded Reasoning

CURV:通过课程可视化接地推理增强图表理解

Xuehang Guo, Pingyue Zhang, Ruiyi Zhang, Zhenhailong Wang, Hanrui Lyu, Heng Ji, Tong Sun, Qingyun Wang, Manling Li

专题命中 多模态生成 :multimodal(abstract);分类 cs.CV、cs.CL、cs.AI

AI总结 针对多模态大语言模型视觉接地与推理不足的问题,提出CURV课程学习框架,结合CCQA数据集,在图表问答任务中实现显著性能提升并具备良好泛化性。

详情

展开后加载摘要…

URL PDF HTML 收藏