FRAME:面向对象中心场景记忆的属性读出因子化检索
FRAME: Factored Retrieval via Attribute Readouts for Object-Centric Scene Memory
- Seoul National University(首尔大学)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
针对语言引导机器人场景记忆中的多属性组合检索问题,提出FRAME方法,将语言转化为属性权重,利用学习读出器聚合对象嵌入证据进行排序,在保留场景上超越基线并实现高效计算。
AI中文摘要:
语言引导的机器人需要持久的场景记忆,以便遵循指令、重新访问物体并解析对随时间遇到的物体的引用。尽管许多语言引导的场景记忆检索工作强调空间或关系引用,但许多日常物体引用通过多个持久属性(如类别、材质、尺寸或表面外观)来指定物体。我们将此问题形式化为属性组合检索,其中固定的对象中心场景记忆被自然语言查询,以检索满足所请求属性的物体。为了直接研究这种能力,我们引入了一个受控评估协议,该协议使用固定的场景记忆和属性定义的目标,将检索与感知和注释歧义分开。然后,我们提出了FRAME,它将语言转化为查询相关的属性权重,使用学习到的读出器从对象嵌入中估计每个属性的证据,并根据查询聚合这些证据对对象进行排序。在保留的场景和对象资产上,FRAME优于代表性的场景记忆检索基线,同时将分解后的对象评分简化为轻量级的矩阵-向量计算。这些结果将属性组合检索定位为语言引导机器人的一种互补场景记忆能力,表明持久对象属性可以作为可组合的证据暴露出来,以实现准确且高效的多属性检索。
英文摘要:
Language-guided robots need persistent scene memories to follow instructions, revisit objects, and resolve references to objects encountered over time. While much of language-guided scene-memory retrieval has emphasized spatial or relational references, many everyday object references specify objects by multiple persistent attributes, such as category, material, size, or surface appearance. We formalize this problem as attribute-compositional retrieval, where a fixed object-centric scene memory is queried with natural language to retrieve the object satisfying the requested attributes. To investigate this capability directly, we introduce a controlled evaluation protocol with fixed scene memories and attribute-defined targets, separating retrieval from perception and annotation ambiguities. We then propose FRAME, which turns language into query-relevant attribute weights, uses learned readouts to estimate per-attribute evidence from object embeddings, and ranks objects by aggregating this evidence according to the query. Across held-out scenes and object assets, FRAME outperforms representative scene-memory retrieval baselines while reducing post-decomposition object scoring to lightweight matrix-vector computation. These results position attribute-compositional retrieval as a complementary scene-memory capability for language-guided robots, showing that persistent object attributes can be exposed as composable evidence for accurate and efficient multi-attribute retrieval.