arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

多模态大模型

跨文本、图像、视频、音频等模态的大模型与学习方法。

共收录 4965 信号源:cs.CV, cs.CL, cs.AI, cs.MM, eess.AS

1. 多模态生成 4965 篇

1808.01357 2018-09-06 cs.CV cs.LG stat.ML 57%

A recurrent multi-scale approach to RBG-D Object Recognition

Mirco Planamente, Mohammad Reza Loghmani, Barbara Caputo

专题命中 多模态生成 :multi-modal(abstract);分类 cs.CV

Comments Master thesis extracted from the paper arXiv:1806.01673 submitted to accv 2018

详情

展开后加载摘要…

URL PDF HTML 收藏
1805.03356 2018-08-08 cs.CV 57%

SPG-Net: Segmentation Prediction and Guidance Network for Image Inpainting

Yuhang Song, Chao Yang, Yeji Shen, Peng Wang, Qin Huang, C. -C. Jay Kuo

专题命中 多模态生成 :multi-modal(abstract);分类 cs.CV

Comments BMVC 2018 camera ready

详情

展开后加载摘要…

URL PDF HTML 收藏
1712.02294 2018-07-17 cs.CV 57%

Joint 3D Proposal Generation and Object Detection from View Aggregation

Jason Ku, Melissa Mozifian, Jungwook Lee, Ali Harakeh, Steven Waslander

专题命中 多模态生成 :multimodal(abstract);分类 cs.CV

Comments For any inquiries contact aharakeh(at)uwaterloo(dot)ca

详情

展开后加载摘要…

URL PDF HTML 收藏
1803.10404 2018-05-23 cs.CV 57%

Lip Movements Generation at a Glance

Lele Chen, Zhiheng Li, Ross K. Maddox, Zhiyao Duan, Chenliang Xu

专题命中 多模态生成 :audio-visual(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
1805.03508 2018-05-10 cs.CV 57%

Rethinking Diversified and Discriminative Proposal Generation for Visual Grounding

Zhou Yu, Jun Yu, Chenchao Xiang, Zhou Zhao, Qi Tian, Dacheng Tao

专题命中 多模态生成 :multi-modal(abstract);分类 cs.CV

Comments Accepted in IJCAI 2018

详情

展开后加载摘要…

URL PDF HTML 收藏
1708.05824 2018-02-14 cs.AI 57%

Applying Deep Bidirectional LSTM and Mixture Density Network for Basketball Trajectory Prediction

Yu Zhao, Rennong Yang, Guillaume Chevalier, Rajiv Shah, Rob Romijnders

专题命中 多模态生成 :multi-modal(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
1801.08468 2018-01-26 cs.CV 57%

Convolutional Invasion and Expansion Networks for Tumor Growth Prediction

Ling Zhang, Le Lu, Ronald M. Summers, Electron Kebebew, Jianhua Yao

专题命中 多模态生成 :multimodal(abstract);分类 cs.CV

Journal ref IEEE Transactions on Medical Imaging, 15 November 2017, Volume:PP Issue: 99

详情

展开后加载摘要…

URL PDF HTML 收藏
1801.00356 2018-01-03 cs.CY cs.AI 57%

How will the Internet of Things enable Augmented Personalized Health?

Amit Sheth, Utkarshani Jaimini, Hong Yung Yip

专题命中 多模态生成 :multimodal(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
1706.07613 2017-11-23 cs.SD cs.IR cs.MM 57%

Toward Faultless Content-Based Playlists Generation for Instrumentals

Yann Bayle, Matthias Robine, Pierre Hanna

专题命中 多模态生成 :multi-modal(abstract);分类 cs.MM

Comments single-column 20 pages, 3 figures, 6 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
1707.04935 2017-09-29 cs.HC cs.CL cs.CY 57%

Automatized Generation of Alphabets of Symbols

Serhii Hamotskyi, Anis Rojbi, Sergii Stirenko, Yuri Gordienko

专题命中 多模态生成 :multimodal(abstract);分类 cs.CL

Comments 4 pages, 3 figures; Federated Conference on Computer Science and Information Systems, Prague (FedCSIS-2017) (Prague, Czech Republic)

Journal ref Proceedings of the 2017 Federated Conference on Computer Science and Information Systems (FedCSIS-2017), p.639-642, Prague, Czech Republic, September 3-6, 2017

详情

展开后加载摘要…

URL PDF HTML 收藏
1707.06267 2017-07-21 cs.CV cs.GR 57%

Shape Generation using Spatially Partitioned Point Clouds

Matheus Gadelha, Subhransu Maji, Rui Wang

专题命中 多模态生成 :multi-modal(abstract);分类 cs.CV

Comments To appear at BMVC 2017

详情

展开后加载摘要…

URL PDF HTML 收藏
1704.06567 2017-04-24 cs.CL cs.NE 57%

Attention Strategies for Multi-Source Sequence-to-Sequence Learning

Jindřich Libovický, Jindřich Helcl

专题命中 多模态生成 :multimodal(abstract);分类 cs.CL

Comments 7 pages; Accepted to ACL 2017

详情

展开后加载摘要…

URL PDF HTML 收藏
1204.1177 2012-04-06 cs.CV 57%

Principal Component Analysis-Linear Discriminant Analysis Feature Extractor for Pattern Recognition

Aamir Khan, Hasan Farooq

专题命中 多模态生成 :multimodal(abstract);分类 cs.CV

Journal ref IJCSI International Journal of Computer Science Issues, Vol. 8, Issue 6, No 2, November 2011

详情

展开后加载摘要…

URL PDF HTML 收藏
0806.0784 2009-12-01 cs.AI cs.HC cs.MA 57%

Collaborative model of interaction and Unmanned Vehicle Systems' interface

Sylvie Saget, Francois Legras, Gilles Coppin

专题命中 多模态生成 :multi-modal(abstract);分类 cs.AI

Journal ref Dans HCP workshop on "Supervisory Control in Critical Systems Management" - 3rd International Conference on Human Centered Processes (HCP-2008), Delft : Pays-Bas (2008)

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.08954 2026-01-15 cs.HC 56%

Leveraging learning analytics to enhance immersive teacher simulations: Challenges and opportunities

利用学习分析增强沉浸式教师模拟:挑战与机遇

Sumin Hong, Jewoong Moon, Taeyeon Eom, Juno Hwang, Jibeom Seo

专题命中 多模态生成 :multimodal(abstract,comments)

AI总结 本研究探讨如何利用学习分析增强沉浸式教师模拟,分析其在教师专业发展中的挑战与机遇。

Comments 30 pages, 10 figures. This chapter examines immersive teacher simulations and multimodal analytics. Project website: https://teachergenai.github.io/

详情

展开后加载摘要…

URL PDF HTML 收藏
2405.18406 2025-10-06 cs.CV cs.AI cs.CL 56%

RACCooN: A Versatile Instructional Video Editing Framework with Auto-Generated Narratives

Jaehong Yoon, Shoubin Yu, Mohit Bansal

机构 * UNC Chapel Hill(北卡罗来纳大学教堂山分校) Nanyang Technological University(南洋理工大学)

专题命中 多模态生成 :分类 cs.CV、cs.CL、cs.AI;MLLM(comments)

Comments EMNLP 2025 main; The first two authors contribute equally. Project Page: https://raccoon-mllm-gen.github.io/

详情

展开后加载摘要…

URL PDF HTML 收藏
2404.10054 2024-04-17 cs.CV cs.AI cs.CL cs.RO 56%

AIGeN: An Adversarial Approach for Instruction Generation in VLN

Niyati Rawal, Roberto Bigazzi, Lorenzo Baraldi, Rita Cucchiara

专题命中 多模态生成 :分类 cs.CV、cs.CL、cs.AI;multimodal(comments)

Comments Accepted to 7th Multimodal Learning and Applications Workshop (MULA 2024) at the IEEE/CVF Conference on Computer Vision and Pattern Recognition 2024

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.11346 2025-10-14 cs.CV cs.AI 54%

Uncertainty-Aware ControlNet: Bridging Domain Gaps with Synthetic Image Generation

Joshua Niemeijer, Jan Ehrhardt, Heinz Handels, Hristina Uzunova

机构 * German Aerospace Center (DLR)(德国航空航天中心) University of Lübeck(吕贝克大学) German Research Center for Artificial Intelligence (DFKI)(德国人工智能研究中心)

专题命中 多模态生成 :分类 cs.CV、cs.AI;multimodal(comments);multimodal foundation model(comments)

Comments Accepted for presentation at ICCV Workshops 2025, "The 4th Workshop on What is Next in Multimodal Foundation Models?" (MMFM)

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.20525 2026-08-24 cs.DB 新提交 50%

Bolo: Verified Model Hub for Next-Generation AI Databases

Bolo:面向下一代AI数据库的经验证的模型中心

Yunqi Li, Ila Petrovic, Yongjoo Park

专题命中 多模态生成 :multi-modal(abstract)

AI总结 该研究针对现有模型平台无法满足AI数据库需求的问题,提出基于多阶段智能体系统的Bolo模型平台,可将不可用模型权重转化为经验证的可用推理流水线,在实验中取得良好效果。

Comments 5 Pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.19590 2026-08-21 eess.IV 新提交 50%

Loss-Resilient Semantic Communication over Packet-Loss Networks at Extreme-Low Bandwidth

极端低带宽丢包网络下的抗丢包语义通信

Shengshi Yao, Jincheng Dai, Sixian Wang, Guo Lu, Kai Niu, Wenjun Xu, Wenjun Zhang, Ping Zhang

专题命中 多模态生成 :multi-modal(abstract)

AI总结 针对极端低带宽丢包网络的语义通信问题,提出ResiGLC框架,结合语言模型上下文预测与掩码学习,通过渐进式解码提升抗丢包性,可在低带宽下改善通信的感知质量。

Comments Accepted to appear in IEEE TMC. 14 pages, 15 figures, and 5 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.19809 2026-08-21 cs.SI 新提交 50%

From Latent Influence to Language: Diffusion-Oriented Content Generation via Audience-Susceptible Features

从潜在影响到语言:面向扩散的受众易感特征内容生成

Jiaying Lei, Shengqi Dang, Runqian Bai, Ziqing Qian, Nan Cao

专题命中 多模态生成 :multimodal(abstract)

AI总结 针对现有面向扩散的内容生成方法难以转化扩散影响信号的问题,提出三阶段框架DOCG-AS,实验显示其在预测扩散影响上优于现有基线。

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.19709 2026-08-21 cs.NI 新提交 50%

RFWM: Physics-Guided World Model for Dynamic Wireless Radiance Field Generation

RFWM:用于生成动态无线辐射场的物理引导世界模型

Zijiu Yang, Qianqian Yang

专题命中 多模态生成 :multimodal(abstract)

AI总结 本研究针对动态未知环境下RF场建模泛化难的问题,提出物理引导的RF世界模型RFWM,采用两阶段训练策略,在新构建的基准上实现了优于现有最优方法的RF场生成性能。

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.22950 2026-08-21 stat.ML cs.LG math.ST stat.ME stat.TH 版本更新 50%

Diffusion-based Denoising Beats Vanilla Score Matching in Parameter Estimation: A Theoretical Explanation

基于扩散的去噪在参数估计中优于普通得分匹配:一个理论解释

Benedikt Lütke Schwienhorst, Nadja Klein, Johannes Lederer

专题命中 多模态生成 :multimodal(abstract)

AI总结 本文通过理论分析证明,对于具有分离模态的多峰分布,基于扩散的去噪得分匹配估计器(DDSME)相比普通得分匹配估计器(SME)具有更优的误差界,并解释了扩散方法优越性的原因。

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.19067 2026-08-20 stat.ML cs.LG math.ST stat.TH 新提交 50%

Diffusion Models for High-Dimensional Clustered Data: Intrinsic-Dimension Adaptivity via Bayesian Classification

用于高维聚类数据的扩散模型:通过贝叶斯分类实现本征维度适应性

Yuga Iguchi, Paul Fearnhead

专题命中 多模态生成 :multimodal(abstract)

AI总结 该研究将扩散模型理论与多模态高维聚类数据几何结合,通过K混合高斯框架证明其去噪可作为动态贝叶斯分类器,KL误差界线性依赖聚类最大本征维度,实现了本征维度适应性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.17769 2026-08-19 eess.SP 新提交 50%

Electromagnetic World Model for 6G: A Unified Framework for Joint Environment Reconstruction and Channel Prediction

面向6G的电磁世界模型:环境重建与信道预测的统一框架

Yizhu Zhao, Li Yu, Jianhua Zhang, Yuxiang Zhang, Zhen Zhang, Guangyi Liu

专题命中 多模态生成 :multi-modal(abstract)

AI总结 该研究针对6G智能终端需同时实现环境感知与通信的需求,提出EMWM统一框架,融合多模态信息完成信道预测与环境重建,性能优于基线方法,具备鲁棒性与零样本泛化能力。

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.00490 2026-08-19 eess.SY cs.SY 版本更新 50%

Cooperative Safety Intelligence over V2X Networks: A Survey

车联网中协同安全智能:一项综述

Jiaxun Zhang, Qian Xu, Zhenning Li, Yuan Wu, Chengzhong Xu, Keqiang Li

专题命中 多模态生成 :multi-modal(abstract)

AI总结 本文综述了V2X环境下协同安全智能的研究进展,围绕SPD框架分析了协同感知、多模态预测和风险感知规划的发展,提出基于SPD的设计原则和评估实践,以提升交通安全水平。

Comments Published in IEEE Communications Surveys & Tutorials (Early Access). DOI: 10.1109/COMST.2026.3723946

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.16229 2026-08-18 cs.RO 新提交 50%

Planner-Conditioned Diffusion for Coordinated Multi-Agent Exploration

用于协调多智能体探索的规划器条件扩散模型

Marcus Yu Siong Teo, Jeric Lew, Tanishq Duhan, Guillaume Sartoretti

机构 * National University of Singapore(新加坡国立大学)

专题命中 多模态生成 :multimodal(abstract)

AI总结 提出规划器条件扩散策略PCDP,通过规划器身份作为条件输入训练多模态单智能体策略,结合局部重排序实现协调,在多智能体探索中提升性能并验证了方法有效性。

Comments Code and models are available at https://github.com/marmotlab/PCDP

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.03820 2026-08-18 stat.ML cs.LG 版本更新 50%

A Quantitative Approximation Framework for Flow Distillation in Diffusion Models

扩散模型中流蒸馏的定量近似框架

Weiguo Gao, Ming Li, Lei Shi, Hanfei Zhou

机构 * School of Mathematical Sciences, Fudan University(复旦大学数学学院) Shanghai Key Laboratory of Contemporary Applied Mathematics, Fudan University(复旦大学当代应用数学重点实验室)

专题命中 多模态生成 :multimodal(abstract)

AI总结 针对扩散模型中的流蒸馏,提出一个定量近似框架,将少步采样视为学习流映射组合下的误差传播,通过理论分析和实验验证了稳定性平衡的非均匀时间网格能显著降低端到端相对MSE。

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.10383 2026-08-14 cs.RO 版本更新 50%

Real-World Cooperative Bimanual Dexterous Grasp of Large Objects from Single-View Observations

基于单视图观测的真实场景中大型物体的双臂协同灵巧抓取

Ziming Li, Mingxuan Wu, Jiaqi Zhang, Hongfei Li, Yan Gan, Deqiang Ouyang, Ning Wang

机构 * The University of Auckland(奥克兰大学) Chongqing University(重庆大学) Southwest University(西南大学)

专题命中 多模态生成 :multimodal(abstract)

AI总结 本文针对机器人双臂抓取大型物体的挑战,提出含多模态数据集、DDPM模块及执行策略的真实场景双臂抓取框架,实验表明其对未见物体抓取成功率高。

Comments Accepted to the 2026 IEEE/RSJ International Conference on Intelligent Robots and Systems (IROS 2026)

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.04235 2026-08-12 cs.ET 版本更新 50%

Scale-CDA: A Scalable Retrofit Platform for Cooperative Driving Automation in Production Vehicles

Scale-CDA:一款用于量产车的、可扩展的原型,旨在普及AI辅助的协同驾驶自动化(CDA)

Hao Zhou, Shengming Yuan, Yuhang Wang, Alina Hagen, Haibin Wen

专题命中 多模态生成 :MLLM(abstract_cn)

AI总结 本研究提出Scale-CDA这一开源工具链,基于OpenDBC与Openpilot构建,成本低于1000美元,可实现AI辅助CDA的即插即用改装,通过多车辆测试验证了其低时延与数据隐私保护能力,为普及协同自动驾驶提供了实用方案。

详情

展开后加载摘要…

URL PDF HTML 收藏