Zero-shot HOI Detection with MLLM-based Detector-agnostic Interaction Recognition
基于多模态大语言模型的零样本人-物交互检测
机构 * School of Computer Science and Engineering, Nanjing University of Science and Technology(南京理工大学计算机科学与工程学院) ; School of Computing and Artificial Intelligence, Southwestern University of Finance and Economics(西南财经大学计算机与人工智能学院) ; Nanjing Forestry University(南京林业大学)
专题命中 视觉定位与Grounding :MLLM(title);vision-language model(abstract);VLM(abstract);visual question answering(abstract)
AI总结 本文提出基于多模态大语言模型的零样本人-物交互检测框架,通过解耦检测与识别任务,结合确定性生成方法和空间感知模块,实现高效准确的零样本交互识别。
Comments ICLR 2026