发表机构
Technion University(以色列理工学院)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本文提出动态语义框架,通过显式约定状态绑定集合与感知对齐流程,解决视觉-语言模型在词汇同步任务中的不足,在斯坦福重复参考游戏语料库上达到83.56%的top-5准确率,核心是透明符号层与可检查感知通道的结合。
AI 中文摘要
人类通过反复互动,针对难以描述的新物体达成共同命名,这一过程被心理语言学家称为词汇同步(lexical entrainment)。主流视觉-语言模型在这一任务上表现不佳:近期实证研究表明,它们无法缩短指代表达、复用成功的表达式,也无法在多轮对话中维持稳定的约定状态。本文提出了一个框架以填补这一空白,该框架将约定状态外部化为三个明确、可检查的指称对象绑定集合(Γ, Ξ, Ω),并通过动态语义语境变化规则进行更新。该框架的符号层构建于轻量级感知对齐流程之上,该流程通过SIFT单应性和通用质量指数(Universal Quality Index),将嘈杂的人类指代表达式在众包图像中建立基础。在斯坦福重复参考游戏语料库(包含针对抽象七巧板刺激的15000余条指挥者-匹配者话语)上进行评估时,该框架仅通过单条指挥者话语就能在其前5个假设集中正确放置目标的比例达83.56%;同一语料库中人类匹配者的前1准确率约为77%-80%。本文还报告了保留条件下的结果,即从检索集中排除明显与七巧板相关的图像,这为基础信号提供了更保守的测量。消融实验分离了每个组件的贡献:SIFT对齐、UQI、查询预处理和图像增强。核心贡献在于二者的结合:一个透明、可审计的符号层,逐轮恢复词汇同步的结构,搭配一个可通过消融实验逐一检查其行为的感知通道。本文还详细讨论了该框架未完成的任务:它不具备交互性,无法与指挥者形成闭环,且其检索驱动的感知通道易受一类泄漏效应的影响,本文对这类效应进行了量化和界定,而非回避。
英文摘要
Humans converge on shared names for novel, hard-to-describe objects through repeated interaction, a process psycholinguists call lexical entrainment. Leading vision-language models fail at this: recent empirical work documents that they do not shorten references, reuse successful expressions, or maintain stable pact state across turns. We present a framework that addresses the gap by externalizing pact state into three explicit, inspectable sets of referent-object bindings ($Γ, Ξ, Ω$), updated by a dynamic-semantics context-change rule. The symbolic layer sits on top of a lightweight perceptual-alignment pipeline that grounds noisy human referring expressions in crowd-sourced imagery via SIFT homographies and the Universal Quality Index. Evaluated on the Stanford Repeated Reference Game corpus (over 15{,}000 director-matcher utterances on abstract tangram stimuli), the framework places the correct target in its top-5 hypothesis set 83.56% of the time from a single director utterance. Human matcher top-1 accuracy on the same corpus is approximately 77-80%. We also report results on a held-out condition in which obvious tangram-adjacent images are excluded from the retrieved set, which provides a more conservative measurement of the grounding signal. Ablations isolate the contribution of each component: SIFT alignment, UQI, query preprocessing, and image augmentation. The central contribution is the combination: a transparent, auditable symbolic layer that recovers the structure of lexical entrainment turn by turn, paired with a perceptual channel whose behavior can be examined ablation by ablation. We also discuss in detail what the framework does not do. It is not interactive, it does not close the loop with the director, and its retrieval-driven perceptual channel is vulnerable to a class of leakage effects that we quantify and bound rather than wave away.
Comments23 pages, 5 figures