发表机构
University of California, Santa Cruz(加州大学圣克鲁兹分校)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
InsightSeg是一种情景记忆机制,将成功修正片段转化为可复用视觉见解,用于多智能体精修系统,在Waymo和Cityscapes数据集上提升了符合准则的分割性能并减少了精修步骤。
AI 中文摘要
符合准则的语义分割不仅需要类别识别,因为现实世界的标注策略要求细粒度、特定任务的决策。近期的多智能体精修系统通过检测和修正错误来提升对这类文本准则的遵循度,但它们是无状态的:批评智能体的反馈会被丢弃,导致数据集上反复重新发现并修正相同的准则特定错误,额外精修成本高。我们提出InsightSeg,一种将成功修正片段转化为可复用、视觉接地见解的情景记忆机制。元分析器将每个合格片段提炼为指令性自然语言见解,并利用补丁级视觉概念向量将其锚定到引发错误的局部图像区域。在后续图像上,这些概念与密集补丁嵌入匹配以检索相关见解,在分割智能体首次预测前对其进行条件调整。这将系统从修正重复错误转向预防错误,在任何精修发生前提升分割质量。在Waymo和Cityscapes数据集上,InsightSeg同时提升了首次预测和最终符合准则的分割性能,且所需精修步骤更少,证明多智能体精修可通过利用过往修正经验变得更准确高效。
英文摘要
Guideline-consistent semantic segmentation requires more than category recognition, as real-world labeling policies demand fine-grained, task-specific decisions. Recent multi-agent refinement systems improve compliance with such textual guidelines by detecting and correcting errors. However, they are stateless: feedback from the critiquing agent is discarded, causing the same guideline-specific mistakes to be repeatedly rediscovered and corrected across the dataset at the cost of additional refinement. We introduce InsightSeg, an episodic memory mechanism that converts successful correction episodes into reusable, visually grounded insights. A meta-analyzer distills each qualifying episode into directive natural-language insights and anchors them to the local image regions that caused the error using patch-level visual concept vectors. On subsequent images, these concepts are matched against dense patch embeddings to retrieve relevant insights, which condition the segmenting agent before making its first prediction. This shifts the system from correcting recurring errors to preventing them, improving segmentation quality before any refinement occurs. Across Waymo and Cityscapes, InsightSeg improves both first-pass and final guideline-consistent segmentation performance while requiring fewer refinement steps, demonstrating that multi-agent refinement can become more accurate and efficient by drawing on past correction experience.