arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

视觉大模型 / VLM

视觉语言模型、视觉推理、视觉问答、图文理解和视觉 grounding。

2025-08-08 至 2025-08-08 共收录 3 信号源:cs.CV, cs.AI, cs.LG

1. 其他VLM 3 篇

2508.05580 2025-08-08 cs.CV 85%

Follow-Your-Instruction: A Comprehensive MLLM Agent for World Data Synthesis

Kunyu Feng, Yue Ma, Xinhua Zhang, Boshi Liu, Yikuang Yuluo, Yinhan Zhang, Runtao Liu, Hongyu Liu, Zhiyuan Qin, Shanhui Mo, Qifeng Chen, Zeyu Wang

机构 * HKUST(GZ)(香港科技大学(广州)) HKUST(香港科技大学) Tsinghua Univerisity(清华大学) Peking University(北京大学) Chongqing University(重庆大学) Beijing Innovation Center of Humanoid Robotics(北京人形机器人创新中心)

专题命中 其他VLM :MLLM(title,abstract);vision-language model(abstract);multimodal large language model(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.04998 2025-08-08 cs.CV 57%

Attribute Guidance With Inherent Pseudo-label For Occluded Person Re-identification

Rui Zhi, Zhen Yang, Haiyang Zhang

机构 * Beijing University of Post and Telecommunication(北京邮电大学)

专题命中 其他VLM :vision-language model(abstract);分类 cs.CV

Comments 8 pages, 2 supplement pages, 3 figures, ECAI2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.04732 2025-08-08 cs.LG cs.GR 57%

LumiGen: An LVLM-Enhanced Iterative Framework for Fine-Grained Text-to-Image Generation

Xiaoqi Dong, Xiangyu Zhou, Nicholas Evans, Yujia Lin

机构 * Dali University(大理大学) Bandırma Onyedi Eylül University(班迪尔马第十七个九月大学)

专题命中 其他VLM :vision-language model(abstract);分类 cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏