Enhancing Spatio-Temporal Zero-shot Action Recognition with Language-driven Description Attributes
机构 * Department of Artificial Intelligence, Korea University(韩国大学人工智能系)
专题命中 GUI与屏幕智能体 :vision-language model(abstract);分类 cs.CV
AI 大模型
视觉语言模型、视觉推理、视觉问答、图文理解和视觉 grounding。
机构 * Department of Artificial Intelligence, Korea University(韩国大学人工智能系)
专题命中 GUI与屏幕智能体 :vision-language model(abstract);分类 cs.CV
机构 * Australian Institute for Machine Learning, the University of Adelaide(澳大利亚机器学习研究所、阿德莱德大学) ; CREATE Lab, Swiss Federal Institute of Technology Lausanne (EPFL)(洛桑联邦理工学院CREATE实验室)
专题命中 GUI与屏幕智能体 :multimodal large language model(abstract);分类 cs.CV
机构 * Fudan University(复旦大学) ; Shanghai Innovation Institute(上海创新研究院) ; National University of Singapore(新加坡国立大学)
专题命中 GUI与屏幕智能体 :multimodal large language model(abstract);分类 cs.CV