arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

多模态大模型

跨文本、图像、视频、音频等模态的大模型与学习方法。

2025-12-29 至 2025-12-29 共收录 6 信号源:cs.CV, cs.CL, cs.AI, cs.MM, eess.AS

1. 多模态生成 6 篇

2512.21867 2025-12-29 cs.CV 57%

DPAR: Dynamic Patchification for Efficient Autoregressive Visual Generation

DPAR:动态分块化用于高效的自回归视觉生成

Divyansh Srivastava, Akshay Mehra, Pranav Maneriker, Debopam Sanyal, Vishnu Raj, Vijay Kamarshi, Fan Du, Joshua Kimball

机构 * University of California, San Diego(加州大学圣地亚哥分校) Dolby Laboratories(杜比实验室)

专题命中 多模态生成 :multimodal(abstract);分类 cs.CV

AI总结 DPAR通过动态分块化技术提升自回归视觉生成效率,减少计算资源消耗并提高生成质量。

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.21691 2025-12-29 cs.CV 57%

Analyzing the Mechanism of Attention Collapse in VGGT from a Dynamics Perspective

从动力学角度分析VGGT中注意力崩溃的机制

Huan Li, Longjun Luo, Yuling Shi, Xiaodong Gu

机构 * Huazhong University of Science and Technology(华中科技大学) Guangdong University of Technology(广东工业大学) Shanghai Jiao Tong University(上海交通大学)

专题命中 多模态生成 :multi-modal(abstract);分类 cs.CV

AI总结 VGGT中注意力崩溃现象从动力学角度进行数学解释,揭示token特征流收敛规律及token合并方法对延迟崩溃的作用。

Comments 8 pages, 4 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.21560 2025-12-29 cs.CV 57%

Toward Intelligent Scene Augmentation for Context-Aware Object Placement and Sponsor-Logo Integration

迈向智能场景增强以实现情境感知的对象放置和赞助商-标志集成

Unnati Saraswat, Tarun Rao, Namah Gupta, Shweta Swami, Shikhar Sharma, Prateek Narang, Dhruv Kumar

机构 * Birla Institute of Technology and Science(巴里尔理工学院和科学研究院)

专题命中 多模态生成 :multimodal(abstract);分类 cs.CV

AI总结 本文提出两个新任务:情境感知的对象插入和赞助商-产品标志增强,并构建了相应数据集以支持这些任务。

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.17376 2025-12-29 cs.CV 57%

Towards Deeper Emotional Reflection: Crafting Affective Image Filters with Generative Priors

迈向更深层次的情感反思:基于生成先验的富有情感的图像滤镜

Peixuan Zhang, Shuchen Weng, Jiajun Tang, Si Li, Boxin Shi

机构 * School of Artificial Intelligence, Beijing University of Posts and Telecommunications(北京邮电大学人工智能学院) Beijing Academy of Artificial Intelligence(北京人工智能研究院) State Key Laboratory of Multimedia Information Processing and National Engineering Research Center of Visual Technology, School of Computer Science, Peking University(北京大学多媒体信息处理国家重点实验室和视觉技术国家工程研究中心)

专题命中 多模态生成 :multi-modal(abstract);分类 cs.CV

AI总结 本文提出AIF任务,通过生成先验提升图像情感表达,实现更深层次的情感反思。

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.18699 2025-12-29 cs.CV 57%

Affective Image Editing: Shaping Emotional Factors via Text Descriptions

情感图像编辑:通过文本描述塑造情感因素

Peixuan Zhang, Shuchen Weng, Chengxuan Zhu, Binghao Tang, Zijian Jia, Si Li, Boxin Shi

机构 * School of Artificial Intelligence, Beijing University of Posts and Telecommunications, China(北京邮电大学人工智能学院) Beijing Academy of Artificial Intelligence(北京人工智能研究院) State Key Laboratory for Multimedia Information Processing and National Engineering Research Center of Visual Technology, School of Computer Science, Peking University, China(多媒体信息处理国家重点实验室和视觉技术国家工程研究中心,北京大学计算机学院)

专题命中 多模态生成 :MLLM(abstract);分类 cs.CV

AI总结 AIEdiT通过文本描述实现情感图像编辑,利用情感映射器和MLLM生成符合用户情感需求的图像。

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.21475 2025-12-29 cs.NI 50%

Physics-informed Diffusion Models for Multi-scale Prediction of Reference Signal Received Power in Wireless Networks

基于物理信息的扩散模型用于无线网络中参考信号接收功率的多尺度预测

Xiaoqian Qi, Haoye Chai, Yue Wang, Zhaocheng Wang, Yong Li

专题命中 多模态生成 :multimodal(abstract)

AI总结 本文提出Channel-Diff框架,通过物理信息条件扩散模型实现无线网络中参考信号接收功率的多尺度预测,提升预测精度与模型可迁移性。

详情

展开后加载摘要…

URL PDF HTML 收藏