arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

多模态大模型

跨文本、图像、视频、音频等模态的大模型与学习方法。

2025-11-12 至 2025-11-12 共收录 8 信号源:cs.CV, cs.CL, cs.AI, cs.MM, eess.AS

1. 多模态生成 8 篇

2510.14631 2025-11-12 cs.DB 78%

Towards a Multimodal Stream Processing System

Uélison Jean Lopes dos Santos, Alessandro Ferri, Szilard Nistor, Riccardo Tommasini, Carsten Binnig, Manisha Luthra

专题命中 多模态生成 :multimodal(title,abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.07934 2025-11-12 cs.CV 74%

Laytrol: Preserving Pretrained Knowledge in Layout Control for Multimodal Diffusion Transformers

Sida Huang, Siqi Huang, Ping Luo, Hongyuan Zhang

专题命中 多模态生成 :multimodal(title);分类 cs.CV

Comments Accepted by AAAI 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.07816 2025-11-12 cs.CV 74%

Cancer-Net PCa-MultiSeg: Multimodal Enhancement of Prostate Cancer Lesion Segmentation Using Synthetic Correlated Diffusion Imaging

Jarett Dewbury, Chi-en Amy Tai, Alexander Wong

机构 * Systems Design Engineering University of Waterloo(水力工程系统设计系大学)

专题命中 多模态生成 :multimodal(title);分类 cs.CV

Comments Accepted at ML4H 2025 Findings

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.08536 2025-11-12 cs.CV 57%

3D4D: An Interactive, Editable, 4D World Model via 3D Video Generation

Yunhong He, Zhengqing Yuan, Zhengzhong Tu, Yanfang Ye, Lichao Sun

专题命中 多模态生成 :multi-modal(abstract);分类 cs.CV

Comments Accepted by AAAI 2026 Demo Track

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.07877 2025-11-12 cs.CV 57%

Visual Bridge: Universal Visual Perception Representations Generating

Yilin Gao, Shuguang Dou, Junzhou Li, Zhiheng Yu, Yin Li, Dongsheng Jiang, Shugong Xu

专题命中 多模态生成 :image-text(abstract);分类 cs.CV

Comments Accepted by AAAI2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.07744 2025-11-12 cs.CV 57%

VectorSynth: Fine-Grained Satellite Image Synthesis with Structured Semantics

Daniel Cher, Brian Wei, Srikumar Sastry, Nathan Jacobs

机构 * Washington University in St. Louis(华盛顿大学圣路易斯分校)

专题命中 多模态生成 :cross-modal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.07690 2025-11-12 cs.AI 57%

Towards AI-Assisted Generation of Military Training Scenarios

Soham Hans, Volkan Ustun, Benjamin Nye, James Sterrett, Matthew Green

专题命中 多模态生成 :multi-modal(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.03668 2025-11-12 cs.DC cs.LG 50%

Intelligent Orchestration of Distributed Large Foundation Model Inference at the Edge

Fernando Koch, Aladin Djuhera, Alecio Binotto

机构 * Florida Atlantic University, USA(佛罗里达大学) Technical University Munich, Germany(慕尼黑技术大学) Carl Zeiss AG, Germany(蔡司股份公司)

专题命中 多模态生成 :multi-modal(abstract)

Comments 26 pages, 3 figures, 4 tables, 52 references

Journal ref Computer Networks and Communications, 2025

详情

展开后加载摘要…

URL PDF HTML 收藏