arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

视觉大模型 / VLM

视觉语言模型、视觉推理、视觉问答、图文理解和视觉 grounding。

2025-11-25 至 2025-11-25 共收录 2 信号源:cs.CV, cs.AI, cs.LG

1. 其他VLM 2 篇

2511.17727 2025-11-25 cs.CV 79%

The Potential and Limitations of Vision-Language Models for Human Motion Understanding: A Case Study in Data-Driven Stroke Rehabilitation

视觉-语言模型在人类运动理解中的潜力与局限性:数据驱动中风康复的案例研究

Victor Li, Naveenraj Kamalakannan, Avinash Parnandi, Heidi Schambra, Carlos Fernandez-Granda

机构 * Center for Data Science(数据科学中心) Tandon School of Engineering(工程学院) VitalConnect(VitalConnect公司) Department of Neurology(神经病学部) Department of Rehabilitation Medicine(康复医学部) Courant Institute of Mathematical Sciences(数学科学研究所)

专题命中 其他VLM :vision-language model(title,abstract);分类 cs.CV

AI总结 本文研究了视觉-语言模型在数据驱动中风康复中的应用,发现其在运动识别和剂量估计方面存在局限性,但通过优化提示和后处理可实现中等准确性的活动分类和运动检测。

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.06908 2025-11-25 cs.CV cs.AI 62%

Find Them All: Unveiling MLLMs for Versatile Person Re-identification

找到它们全部:揭示MLLMs用于多功能人物重识别

Jinhao Li, Zijian Chen, Lirong Deng, Guangtao Zhai, Changbo Wang

机构 * School of Computer Science and Technology, East China Normal University(东华大学计算机科学与技术学院) Institute of Image Communication and Information Processing, Shanghai Jiao Tong University(上海交通大学图像通信与信息处理研究所) Macao Polytechnic University(澳门 polytechnic 大学) Shanghai AI Laboratory(上海人工智能实验室)

专题命中 其他VLM :MLLM(abstract);分类 cs.CV、cs.AI

AI总结 本文提出VP-ReID基准,利用MLLMs提升人物重识别的多功能性和有效性,同时揭示其在处理某些模态时的局限性。

详情

展开后加载摘要…

URL PDF HTML 收藏