arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

EgoAfford:基于自我中心指称分割的面向任务的可供性定位

EgoAfford: Task-Oriented Affordance Grounding via Egocentric Referring Segmentation

Xinyuan Guan, Feifan Chen, Xinyu Zhan, Fu-Cheng Zhang, Cewu Lu, Lixin Yang

arXiv 2608.04533首次发表:更新:

AI 中文总结

针对部件级可供性定位难以适配复杂多步桌面任务的问题,提出基准EgoAfford及参考模型EgoLens,验证了下一步推理与动作角色条件部件定位的互补挑战,为感知与规划联合研究提供基础。

AI 中文摘要

部件级可供性定位已在与基本动作相关的功能性物体区域定位方面取得进展。将该能力扩展至复杂任务,需将参与物体的语义角色与任务状态对齐的视觉观测及多步规划建立关联。我们提出EgoAfford,一个用于连接这三个方面的基准。给定一个自我中心观测和一个高级桌面任务,模型必须生成剩余计划,并分割出下一步动作最多三个组成部分的功能性区域:直接对象、工具和目标区域。EgoAfford包含来自2000个生成的多步场景的约15500张经人工验证的图像,这些图像被组织成语义对齐、任务完整的图像序列,还有EgoAfford-Real,即覆盖26个任务的102张人工采集图像。我们进一步提出EgoLens,一个具有角色特定掩码解码器的30亿参数多模态大语言模型,作为该联合任务的域内参考模型。对近期指称分割多模态大语言模型、商业VLM-SAM2管道及EgoLens的评估,凸显了下一步推理和动作角色条件部件定位的互补挑战。EgoLens在生成的和人工采集的观测上均取得了优异的参考性能。综上,EgoAfford和EgoLens为联合研究多步桌面任务中的感知与规划提供了基础。本项目页面可访问:this https URL

英文摘要

Part-level affordance grounding has advanced the localization of functional object regions associated with elemental actions. Extending this capability to complex tasks calls for connecting the semantic roles of participating objects with task-state-aligned visual observations and multi-step planning. We introduce EgoAfford, a benchmark designed to connect these three aspects. Given an egocentric observation and a high-level tabletop task, a model must generate the remaining plan and segment the functional regions of up to three components of the next action: the direct object, instrument, and destination. EgoAfford comprises approximately 15.5k human-verified images from 2,000 generated multi-step scenes, organized as semantically aligned, task-complete image series, together with EgoAfford-Real, 102 manually captured images spanning 26 tasks. We further present EgoLens, a 3B multimodal large language model with role-specific mask decoders, as an in-domain reference model for this joint task. Evaluations of recent referring-segmentation MLLMs, commercial-VLM--SAM2 pipelines, and EgoLens highlight the complementary challenges of next-step inference and action-role-conditioned part grounding. EgoLens establishes strong reference performance on both generated and manually captured observations. Together, EgoAfford and EgoLens provide a foundation for jointly studying perception and planning in multi-step tabletop tasks. Our project page is available at: https://egoafford.github.io

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑