arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

视觉大模型 / VLM

视觉语言模型、视觉推理、视觉问答、图文理解和视觉 grounding。

2025-12-29 至 2025-12-29 共收录 34 信号源:cs.CV, cs.AI, cs.LG

1. 其他VLM 4 篇

2512.21360 2025-12-29 cs.AI cs.SY eess.SY 70%

From Visual Perception to Deep Empathy: An Automated Assessment Framework for House-Tree-Person Drawings Using Multimodal LLMs and Multi-Agent Collaboration

从视觉感知到深度共情:一种利用多模态大语言模型和多智能体协作的房屋-树-人绘画自动评估框架

Shuide Wen, Yu Sun, Beier Ku, Zhi Gao, Lijun Ma, Yang Yang, Can Jiao

专题命中 其他VLM :multimodal large language model(abstract);MLLM(abstract);分类 cs.AI

AI总结 本文提出利用多模态大语言模型和多智能体协作,开发自动评估房屋-树-人绘画测试的框架,以提升投射评估的标准化和效率。

Comments 16 pages, 8 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.21637 2025-12-29 cs.CV 57%

Training-Free Disentangled Text-Guided Image Editing via Sparse Latent Constraints

无需训练的解耦文本引导图像编辑:通过稀疏潜在约束

Mutiara Shabrina, Nova Kurnia Putri, Jefri Satria Ferdiansyah, Sabita Khansa Dewi, Novanto Yudistira

机构 * Department of Informatics Engineering Universitas Brawijaya Malang, Indonesia(信息工程系 乌姆拉大学 马拉邦,印度尼西亚)

专题命中 其他VLM :vision-language model(abstract);分类 cs.CV

AI总结 本文提出一种无需训练的解耦文本引导图像编辑方法,通过引入稀疏潜在约束减少属性纠缠问题,提升编辑的可控性与准确性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.18699 2025-12-29 cs.CV 57%

Affective Image Editing: Shaping Emotional Factors via Text Descriptions

情感图像编辑:通过文本描述塑造情感因素

Peixuan Zhang, Shuchen Weng, Chengxuan Zhu, Binghao Tang, Zijian Jia, Si Li, Boxin Shi

机构 * School of Artificial Intelligence, Beijing University of Posts and Telecommunications, China(北京邮电大学人工智能学院) Beijing Academy of Artificial Intelligence(北京人工智能研究院) State Key Laboratory for Multimedia Information Processing and National Engineering Research Center of Visual Technology, School of Computer Science, Peking University, China(多媒体信息处理国家重点实验室和视觉技术国家工程研究中心,北京大学计算机学院)

专题命中 其他VLM :MLLM(abstract);分类 cs.CV

AI总结 AIEdiT通过文本描述实现情感图像编辑,利用情感映射器和MLLM生成符合用户情感需求的图像。

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.21863 2025-12-29 cs.IR cs.MM 50%

Frozen LVLMs for Micro-Video Recommendation: A Systematic Study of Feature Extraction and Fusion

冻结的大型视频语言模型用于微视频推荐:特征提取与融合的系统研究

Huatuan Sun, Yunshan Ma, Changguang Wu, Yanxin Zhang, Pengfei Wang, Xiaoyu Du

专题命中 其他VLM :vision-language model(abstract)

AI总结 本文通过系统研究冻结LVLMs的特征提取与融合策略,提出DFF框架,证明中间隐藏状态优于标题表示,ID嵌入融合优于替换,并在微视频推荐中取得最佳性能。

Comments 10 pages, 4 figures

详情

展开后加载摘要…

URL PDF HTML 收藏