arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

AI 大模型

多模态大模型

跨文本、图像、视频、音频等模态的大模型与学习方法。

2025-09-30 至 2025-09-30 共收录 15 信号源:cs.CV, cs.CL, cs.AI, cs.MM, eess.AS

1. 多模态生成 15 篇

2509.22930 2025-09-30 cs.CV 79%

FishAI 2.0: Marine Fish Image Classification with Multi-modal Few-shot Learning

Chenghan Yang, Peng Zhou, Dong-Sheng Zhang, Yueyun Wang, Hong-Bin Shen, Xiaoyong Pan

专题命中 多模态生成 :multi-modal(title);multimodal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.23760 2025-09-30 cs.CV 77%

UniAlignment: Semantic Alignment for Unified Image Generation, Understanding, Manipulation and Perception

Xinyang Song, Libin Wang, Weining Wang, Shaozhen Liu, Dandan Zheng, Jingdong Chen, Qi Li, Zhenan Sun

机构 * Ant Group(蚂蚁集团)

专题命中 多模态生成 :multimodal(abstract);multi-modal(abstract);cross-modal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.23635 2025-09-30 cs.CV 74%

MotionVerse: A Unified Multimodal Framework for Motion Comprehension, Generation and Editing

Ruibing Hou, Mingshuang Luo, Hongyu Pan, Hong Chang, Shiguang Shan

机构 * Key Laboratory of Intelligent Information Processing, Institute of Computing Technology (ICT), Chinese Academy of Sciences (CAS)(智能信息处理重点实验室,计算技术研究所(ICT),中国科学院(CAS)) University of the Chinese Academy of Sciences(中国科学院大学)

专题命中 多模态生成 :multimodal(title);分类 cs.CV

Comments 17 pages, 6 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.25047 2025-09-30 cs.AI 70%

Scaling Synthetic Task Generation for Agents via Exploration

Ram Ramrakhya, Andrew Szot, Omar Attia, Yuhao Yang, Anh Nguyen, Bogdan Mazoure, Zhe Gan, Harsh Agrawal, Alexander Toshev

机构 * Apple(苹果公司)

专题命中 多模态生成 :multimodal(abstract);MLLM(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.24427 2025-09-30 cs.CV 70%

UI2V-Bench: An Understanding-based Image-to-video Generation Benchmark

Ailing Zhang, Lina Lei, Dehong Kong, Zhixin Wang, Jiaqi Xu, Fenglong Song, Chun-Le Guo, Chang Liu, Fan Li, Jie Chen

机构 * Peking University(北京大学) Huawei Noah’s Ark Lab(华为诺亚实验室) Nankai University(南开大学) Shenzhen Campus of Sun Yat-sen University(中山大学深圳校区) Tsinghua University(清华大学)

专题命中 多模态生成 :multimodal(abstract);MLLM(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.23625 2025-09-30 cs.CV cs.AI cs.CL cs.LG 67%

RIV: Recursive Introspection Mask Diffusion Vision Language Model

YuQian Li, Limeng Qiao, Lin Ma

机构 * Meituan(美团)

专题命中 多模态生成 :multimodal(abstract);分类 cs.CV、cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.22940 2025-09-30 cs.CL cs.CV 62%

LLMs Behind the Scenes: Enabling Narrative Scene Illustration

Melissa Roemmele, John Joon Young Chung, Taewook Kim, Yuqian Sun, Alex Calderwood, Max Kreminski

机构 * Midjourney Northwestern University(西北大学) University of California, Santa Cruz(加州大学圣克鲁兹分校)

专题命中 多模态生成 :cross-modal(abstract);分类 cs.CV、cs.CL

Comments Accepted at EMNLP 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.24903 2025-09-30 cs.RO cs.CV eess.IV 57%

DRCP: Diffusion on Reinforced Cooperative Perception for Perceiving Beyond Limits

Lantao Li, Kang Yang, Rui Song, Chen Sun

机构 * Sony (China) Limited(索尼(中国)有限公司) Renmin University of China(中国人民大学) Fraunhofer IVI(弗劳恩霍夫IVI研究所)

专题命中 多模态生成 :cross-modal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.24875 2025-09-30 cs.CV cs.LG 57%

Environment-Aware Satellite Image Generation with Diffusion Models

Nikos Kostagiolas, Pantelis Georgiades, Yannis Panagakis, Mihalis A. Nicolaou

专题命中 多模态生成 :multimodal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.24514 2025-09-30 cs.CV 57%

Instruction Guided Multi Object Image Editing with Quantity and Layout Consistency

Jiaqi Tan, Fangyu Li, Yang Liu

机构 * Beijing University of Posts(北京邮电大学) Telecommunications School of Digital Media(电信数字媒体学院)

专题命中 多模态生成 :cross-modal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.24299 2025-09-30 cs.CV 57%

SVGThinker: Instruction-Aligned and Reasoning-Driven Text-to-SVG Generation

Hanqi Chen, Zhongyin Zhao, Ye Chen, Zhujin Liang, Bingbing Ni

机构 * Shanghai Jiao Tong University(上海交通大学) SJTU Paris Elite Institute of Technology(上海交通大学巴黎精英理工学院)

专题命中 多模态生成 :multimodal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.23955 2025-09-30 cs.CV 57%

ColLab: A Collaborative Spatial Progressive Data Engine for Referring Expression Comprehension and Generation

Shilan Zhang, Jirui Huang, Ruilin Yao, Cong Wang, Yaxiong Chen, Peng Xu, Shengwu Xiong

机构 * Wuhan University of Technology(武汉理工大学) Northwestern Polytechnical University(西北工业大学) Tsinghua University(清华大学)

专题命中 多模态生成 :multimodal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.10036 2025-09-30 cs.CL 57%

DataPuzzle: Breaking Free from the Hallucinated Promise of LLMs in Data Analysis

Zhengxuan Zhang, Zhuowen Liang, Yin Wu, Teng Lin, Yuyu Luo, Nan Tang

机构 * The Hong Kong University of Science and Technology (Guangzhou)(香港科技大学(广州))

专题命中 多模态生成 :multi-modal(abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.17608 2025-09-30 cs.LG 50%

ChartMaster: Advancing Chart-to-Code Generation with Real-World Charts and Chart Similarity Reinforcement Learning

Wentao Tan, Qiong Cao, Chao Xue, Yibing Zhan, Changxing Ding, Xiaodong He

机构 * South China University of Technology(华南理工大学) JD Future Academy(京东未来学院) Wuhan University(武汉大学)

专题命中 多模态生成 :multimodal(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.20266 2025-09-30 cond-mat.mtrl-sci 50%

Emergent properties and the multiscale characterization challenge in condensed matter, from crystals to complex materials: a Review

Elisabetta Nocerino

专题命中 多模态生成 :multimodal(abstract)

Journal ref J. Phys. D: Appl. Phys. 58 393001 (2025)

详情

展开后加载摘要…

URL PDF HTML 收藏