发表机构
University of California, Berkeley(加州大学伯克利分校)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本研究通过将听者注视扫描路径转化为学习信号微调视觉语言模型,生成更优指称表达,显著缩短序列并提高指称成功率。
AI 中文摘要
我们提出通过将增量式听者理解观察(以注视扫描路径形式)转化为学习信号,来微调视觉语言模型,以生成更具语用最优性的指称表达。在训练过程中,指称表达从被优化的说话者策略中采样,并以图像和目标指称物为条件;随后,一个估计人类注视行为的神经听者将图像和采样的指称表达映射为扫描路径,每条扫描路径由一系列注视点表示,每个注视点对应指称表达中的一个词。我们实验了几种方法,将注视点序列和目标指称物转换为词级和序列级奖励,用于优化策略参数。通过人类听者评估,我们发现使用注视估计听者训练的说话者策略比基础模型产生显著更具语用最优性的指称,将序列长度从15.4个词减少到4.0个词,同时将指称成功率从75.2%提高到80.0%。我们的工作展示了通过基于语言的交互学习生成话语的有前景的机会,不仅利用交际成功的显式信号,还利用听者理解过程中隐式可用的观察。
英文摘要
We propose to finetune vision-language models to generate more pragmatically optimal referring expressions by transforming observations of incremental listener comprehension, in the form of gaze scanpaths, into learning signals. During training, referring expressions are sampled from the speaker policy being optimized, conditioned on images and target referents; then, a neural listener estimating human gaze behavior maps from images and sampled referring expressions to scanpaths, each represented by a sequence of fixations, with each fixation corresponding to a word in the referring expression. We experiment with several approaches to convert fixation sequences and target referents into token- and sequence-level rewards, which are used to optimize policy parameters. Through evaluation with human listeners, we find that speaker policies trained with gaze-estimating listeners result in significantly more pragmatically-optimal references than base models, reducing sequence length from 15.4 down to 4.0 words while increasing referential success from 75.2 up to 80.0%. Our work demonstrates a promising opportunity for learning to generate utterances through language-based interaction, not only from the explicit signal of communicative success, but also from implicitly-available observations of a listener's process of comprehension.
CommentsAccepted at COLM 2026