arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

多模态大模型

跨文本、图像、视频、音频等模态的大模型与学习方法。

2025-11-24 至 2025-11-24 共收录 4 信号源:cs.CV, cs.CL, cs.AI, cs.MM, eess.AS

1. 多模态生成 4 篇

2511.16917 2025-11-24 cs.CV 83%

UniModel: A Visual-Only Framework for Unified Multimodal Understanding and Generation

UniModel: 一种仅依赖视觉的统一多模态理解和生成框架

Chi Zhang, Jiepeng Wang, Youming Wang, Yuanzhi Liang, Xiaoyan Yang, Zuoxin Li, Haibin Huang, Xuelong Li

机构 * Institute of Artificial Intelligence (TeleAI)(人工智能研究所)

专题命中 多模态生成 :multimodal(title,abstract);cross-modal(abstract);分类 cs.CV

AI总结 UniModel通过统一模型、任务和表示,在单一视觉空间中实现多模态理解和生成的统一框架。

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.17501 2025-11-24 cs.CV cs.GR 57%

Native 3D Editing with Full Attention

原生3D编辑与全关注

Weiwei Cai, Shuangkang Fang, Weicai Ye, Xin Dong, Yunhan Yang, Xuanyang Zhang, Wei Cheng, Yanpei Cao, Gang Yu, Tao Chen

机构 * Fudan University(复旦大学) StepFun, Inc.(StepFun公司) Zhejiang University(浙江大学) Tsinghua University(清华大学) VAST

专题命中 多模态生成 :multi-modal(abstract);分类 cs.CV

AI总结 本文提出一种高效的原生3D编辑框架,通过多模态数据集和3D标记连接方法,提升3D编辑的效率和一致性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.00737 2025-11-24 cs.HC cs.AI 57%

How LLMs are Shaping the Future of Virtual Reality

大型语言模型如何塑造虚拟现实的未来

Süeda Özkaya, Santiago Berrezueta-Guzman, Stefan Wagner

机构 * Technical University of Munich(慕尼黑技术大学)

专题命中 多模态生成 :multimodal(abstract);分类 cs.AI

AI总结 本文探讨了大型语言模型如何通过增强叙事生成、NPC互动和个性化来塑造虚拟现实的未来,并提出了多模态AI和伦理保障等未来研究方向。

Comments Pre-print

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.16900 2025-11-24 eess.SY cs.SY 50%

When Motion Learns to Listen: Diffusion-Prior Lyapunov Actor-Critic Framework with LLM Guidance for Stable and Robust AUV Control in Underwater Tasks

当运动学会倾听:带有LLM引导的扩散先验Lyapunov动作-批评者框架用于水下任务中稳定和鲁棒的AUV控制

Jingzehua Xu, Weiyi Liu, Weihang Zhang, Zhuofan Xi, Guanwen Xie, Shuai Zhang, Yi Li

专题命中 多模态生成 :multimodal(abstract)

AI总结 本文提出一种结合扩散模型、Lyapunov批评者和LLM的框架,用于提升水下机器人控制的稳定性与鲁棒性,通过生成-过滤-优化机制实现高效探索和多目标优化。

Comments This paper is currently under review and does not represent the final version. Jingzehua Xu, Weiyi Liu and Weihang Zhang are co-first authors of this paper, with Zhuofan Xi as the second author

详情

展开后加载摘要…

URL PDF HTML 收藏