arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

多模态大模型

跨文本、图像、视频、音频等模态的大模型与学习方法。

2026-02-02 至 2026-02-02 共收录 5 信号源:cs.CV, cs.CL, cs.AI, cs.MM, eess.AS

1. 多模态生成 5 篇

2601.19700 2026-02-02 cs.LG cs.AI 85%

Generalizable Multimodal Large Language Model Editing via Invariant Trajectory Learning

通过不变轨迹学习实现通用的多模态大语言模型编辑

Jiajie Su, Haoyuan Wang, Xiaohua Feng, Yunshan Ma, Xiaobo Xia, Yuyuan Li, Xiaolin Zheng, Jianmao Xiao, Chaochao Chen

机构 * Zhejiang University, China(浙江大学) Singapore Management University, Singapore(新加坡管理学院) National University of Singapore, Singapore(新加坡国立大学) Hangzhou Dianzi University, China(杭州电子科技大学) Jiangxi Normal University, China(江西师范大学)

专题命中 多模态生成 :multimodal(title,abstract);MLLM(abstract);cross-modal(abstract);分类 cs.AI

AI总结 本文提出ODEit框架,通过不变轨迹学习提升多模态大语言模型的编辑可靠性、局部性和泛化能力。

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.22407 2026-02-02 cond-mat.mtrl-sci 78%

Nanoscale mapping of phase-transformation pathways in medium-Mn TRIP steel by multimodal STEM

中等锰TRIP钢中相变路径的纳米级映射:多模态STEM

Marc Raventós-Tato, S. Leila Panahi, Núria Bagués, David Frómeta, Oleg Usoltsev, Núria Cuadrado, Joaquín Otón

专题命中 多模态生成 :multimodal(title,abstract)

AI总结 本研究通过多模态STEM技术,实现了中等锰TRIP钢中相变路径的纳米级映射,结合电子衍射与能谱分析,实现了铁、奥氏体和马氏体的相分离与晶格参数细化。

Comments 16 pages, 5 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.23121 2026-02-02 cs.MM 57%

An Automatic Deep Learning Approach for Trailer Generation through Large Language Models

基于大语言模型的自动 trailers 生成方法

Roberto Balestri, Pasquale Cascarano, Mirko Degli Esposti, Guglielmo Pescatore

专题命中 多模态生成 :multimodal(abstract);分类 cs.MM

AI总结 本文提出基于大语言模型的自动化 trailers 生成方法,通过多模态策略提升 trailers 的视觉吸引力和叙事体验。

Comments 2024 9th International Conference on Frontiers of Signal Processing (ICFSP)

Journal ref ICFSP, Paris, France, 2024, pp. 93-100

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.22455 2026-02-02 cs.CV 57%

ScribbleSense: Generative Scribble-Based Texture Editing with Intent Prediction

ScribbleSense: 基于生成的涂鸦纹理编辑与意图预测

Yudi Zhang, Yeming Geng, Lei Zhang

机构 * School of Computer Science, Beijing Institute of Technology(计算机学院,北京理工大学)

专题命中 多模态生成 :multimodal(abstract);分类 cs.CV

AI总结 ScribbleSense通过结合多模态大语言模型和图像生成模型,实现了基于生成的涂鸦纹理编辑与意图预测,提升交互式编辑性能。

Comments Accepted by IEEE TVCG. Personal use of this material is permitted. Permission from IEEE must be obtained for all other uses

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.07222 2026-02-02 cs.CV 57%

Omni-View: Unlocking How Generation Facilitates Understanding in Unified 3D Model based on Multiview images

Omni-View: 通过多视角图像解锁生成如何促进理解的统一3D模型

JiaKui Hu, Shanshan Zhao, Qing-Guo Chen, Xuerui Qiu, Jialun Liu, Zhao Xu, Weihua Luo, Kaifu Zhang, Yanye Lu

机构 * Institute of Medical Technology(医学技术研究所) Alibaba International Digital Commerce Group(阿里巴巴国际数字商务集团) CASIA TeleAI

专题命中 多模态生成 :multimodal(abstract);分类 cs.CV

AI总结 Omni-View通过多视角图像实现3D场景的理解与生成,结合纹理和几何模块,提升3D场景建模性能。

Comments Accepted by ICLR 2026

详情

展开后加载摘要…

URL PDF HTML 收藏