发表机构
Beijing Institute of Technology; Shandong University(北京理工大学; 山东大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
提出MOVE框架,通过多模态证据验证与选择性扩展类别空间,解决多模态图学习中开放世界新类别识别与扩展问题,平均提升11.87%。
AI 中文摘要
多模态图学习面临一个根本性挑战:模型在固定标签空间下训练,而部署后可能出现新类别。现有方法通常检测未知节点,并利用大语言模型(LLM)生成候选类别描述,但并未判断现有类别是否不足以覆盖这些节点,或生成的类别是否足够可靠以扩展类别空间。我们的实证研究揭示了三个挑战:识别未知节点需要超越单一模态的多模态信息,LLM生成的类别描述可能无法完全捕捉多模态类别特征,直接添加候选类别可能引入冗余类别。基于这些观察,我们提出MOVE,一个多模态开放世界类别验证与扩展框架。MOVE通过联合考虑视觉标记、文本属性和图上下文来识别无法归入现有类别的节点,利用多模态LLM生成候选类别,并且仅在候选类别得到多模态证据一致支持且不引入不必要类别时,有选择地扩展类别空间。实验表明,MOVE在未知识别、开放域标注和下游图学习任务上平均提升11.87%。
英文摘要
Multimodal graph learning faces a fundamental challenge: new classes may emerge after deployment, while models are trained with a fixed label space. Existing approaches typically detect unknown nodes and use LLMs to generate candidate class descriptions, but they do not determine whether existing classes are insufficient to cover these nodes or whether a generated class is reliable enough to expand the class space. Our empirical study reveals three challenges: multimodal information beyond individual modalities is required for unknown-node identification, LLM-generated class descriptions may not fully capture multimodal class characteristics, and directly adding candidate classes can introduce redundant categories. Based on these observations, we propose MOVE, a multimodal open-world class verification and expansion framework. MOVE identifies nodes that cannot be assigned to existing classes by jointly considering visual tokens, textual attributes, and graph context, leverages a multimodal LLM to generate candidate classes, and selectively expands the class space only when candidates are consistently supported by multimodal evidence without introducing unnecessary categories. Experiments demonstrate that MOVE achieves an average improvement of 11.87\% across unknown recognition, open-domain annotation, and downstream graph learning tasks.