诊断大视觉语言模型中的密集同类属性绑定错误
Diagnosing Dense Same-Class Attribute Misbinding in Large Vision-Language Models
浏览论文内容
中文总结 AI 辅助
本研究提出InstaBind-Lite基准,量化大视觉语言模型的密集同类属性绑定错误(DSCAM),发现其错误率被总准确率掩盖,多数转移来自相邻实例,该基准可评估模型对属性所属实例的认知。
中文摘要 AI 辅助
大型视觉语言模型能够识别拥挤场景中的物体和属性,但会将属性错误分配给同类的不同实例。通用视觉问答准确率会将此类响应标记为错误,而物体幻觉指标可能会认为物体和属性均得到图像支持;两者均未揭示这种属性转移(即错误绑定)。本研究将这一盲区正式定义为密集同类属性绑定错误(Dense Same-Class Attribute Misbinding, DSCAM),并提出了InstaBind-Lite,这是一个受控基准,可直接对该问题进行测量。该基准包含524张图像,其中有529组精心策划的3-6个同类实体、1773个带框实例、有序邻居、可区分的类属性,以及四个互补的问题级别,共产生9580个确定性评估问题。与现有协议不同,源实例注释将不支持的生成和识别失败与从另一个可见实体复制的属性区分开来。绑定特定指标进一步量化了转移频率、邻接性、序数距离和干预效果。在五个开源模型和两个商业/API模型中,开源系统的平均绑定错误率为19.84%,API系统为7.55%;这些错误被总准确率所掩盖。在可识别的转移中,开源模型和API模型分别有80.70%和81.51%来自相邻实例。定位和实例优先干预对部分模型有帮助,但并非通用解决方案。因此,InstaBind-Lite将先前无法区分的错误答案转化为可识别源的失败类别,并测试了传统基准无法确定的可靠性维度:模型不仅知道可见内容,还知道每个属性属于哪个实例。
英文摘要
Large vision-language models can recognize the objects and attributes in a crowded scene yet assign an attribute to the wrong same-class instance. Generic visual-question-answering accuracy marks the response as wrong, while object-hallucination metrics may regard both the object and attribute as image-supported; neither reveals the transfer. This study formalizes this blind spot as Dense Same-Class Attribute Misbinding (DSCAM) and presents InstaBind-Lite, a controlled benchmark that makes it directly measurable. Its 524 images contain 529 curated groups of 3-6 same-class entities, 1773 boxed instances, ordered neighbors, distinguishable color-like attributes, and four complementary question levels, yielding 9580 deterministically evaluated questions. Unlike existing protocols, source-instance annotations separate unsupported generation and recognition failure from an attribute copied from another visible entity. Binding-specific metrics further quantify transfer frequency, adjacency, ordinal distance, and intervention effects. Across five open-source and two commercial/API models, the open-source systems average 19.84% Misbinding Rate and the API systems 7.55%; these errors are hidden by aggregate accuracy. Among identifiable transfers, 80.70% and 81.51%, respectively, originate from adjacent instances. Localization and instance-first interventions help selected models but are not universal remedies. InstaBind-Lite therefore turns previously undifferentiated wrong answers into source-identifiable failure categories and tests a reliability dimension that conventional benchmarks cannot determine: whether a model knows not only what is visible, but which instance owns each attribute.
发表机构
- Qilu University of Technology (Shandong Academy of Sciences)(齐鲁工业大学(山东省科学院))
- China Telecom Digital Intelligence Technology Co., Ltd.(中国电信数字智能科技有限公司)
- Shenyang Aerospace University(沈阳航空航天大学)
- University of Nottingham Ningbo China(宁波诺丁汉大学)
机构由 AI 辅助整理,请以论文原文为准。