发表机构
Stony Brook University; Australian Institute for Machine Learning; Adelaide University(纽约州立大学石溪分校; 澳大利亚机器学习研究所; 阿德莱德大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
该研究提出将凝视解码为自然语言描述人类目标的新问题,构建Gazette框架,基于多模态大语言模型,利用新颖策略合成出声思考转录本并指令调优,使其在多任务凝视解码中达最优性能,能推断多样场景下人类目标意图。
AI 中文摘要
我们引入了一个新的学习问题:在各种视觉任务中将凝视解码为关于人类目标的自然语言描述。与之前将凝视解码视为对预定义类别进行判别任务的工作不同,我们将其表述为一个生成学习问题:训练一个模型来生成自由形式的描述,以捕捉人类意图丰富的细微差别和开放性,而不仅仅是固定标签。为此,我们引入了Gazette,首个凝视到文本的解码框架。基于多模态大语言模型,Gazette学习将凝视扫描路径解码为自然语言,用于可能超出类别标签且需要用自然语言表达的目标。为帮助Gazette滤除凝视行为中的个体差异并学习对生成准确自然语言目标描述至关重要的特定目标时空动态,我们提出一种新颖策略,利用大语言模型的百科知识和推理能力来合成目标导向注意力行为的自然语言解释,即出声思考转录本。对这些合成叙述进行指令调优,使Gazette在多个任务的凝视解码中取得了当前最优性能,证明了其通用性和多功能性,从而使凝视能够作为一种强大的、非侵入性线索,用于推断不同场景中的人类目标和意图。
英文摘要
We introduce a novel learning problem: decoding gaze into natural language descriptions of human goals across diverse visual tasks. Unlike prior work, which frames gaze decoding as a discriminative task over predefined categories, we formulate it as a generative learning problem: training a model to produce free-form descriptions that capture the rich nuances and open-ended nature of human intentions beyond fixed labels. To this end, we introduce Gazette, the first gaze-to-text decoding framework. Based on multimodal large language models (MLLMs), Gazette learns to decode gaze scanpaths into natural language for goals that may extend beyond categorical labels and require articulation in natural language. To help Gazette filter out individual differences in gaze behavior and learn the goal-specific spatiotemporal dynamics crucial for generating accurate natural language goal descriptions, we propose a novel strategy that leverages the encyclopedic knowledge and reasoning abilities of a large language model to synthesize natural language explanations of goal-directed attentional behavior called think-aloud transcripts. Instruction tuning on these synthetic narratives allows Gazette to achieve state-of-the-art performance in gaze decoding across multiple tasks, demonstrating its generalizability and versatility, thereby enabling gaze to serve as a powerful, non-intrusive cue for inferring human goals and intentions in diverse scenarios.
CommentsTo appear in European Conference on Computer Vision (ECCV) 2026