arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

MNIST-PRO:MNIST作为AI智能体的部分可观测世界回归

MNIST-PRO: MNIST is Back as a Partially Observable World for AI Agents

Vernon Toh, Navonil Majumder, Zhengyuan Liu, Nancy F. Chen, Soujanya Poria

arXiv 2608.31022首次发表:更新:

发表机构

Nanyang Technological University; Agency for Science, Technology, and Research (A*STAR)(南洋理工大学; 新加坡科学、技术与研究局(A*STAR))

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本研究提出MNIST-PRO基准,将MNIST数字识别转化为带回溯约束的顺序搜索任务,评估多模态模型在部分可观测性下的表现,发现智能体存在感知整合、探索持续性及信念修正三大瓶颈,凸显构建更新可靠感知状态的重要性。

AI 中文摘要

部分可观测环境中的AI智能体需协调主动感知与工作记忆,以维持不断演化的感知状态。然而,现有基准难以分离感知状态构建与解释能力,因为它们引入了物理与控制复杂度。我们通过MNIST-PRO解决该问题,这是一个将MNIST数字识别转换为带回溯约束的、基于 glimpses 的顺序搜索任务的基准,以此分离智能体感知能力。我们在四种记忆表示下评估了十种多模态模型,包括原始视觉历史、文本状态、结构化度量网格地图及整合视觉画布。尽管模型在完全可观测性下表现出色,部分可观测性却暴露出明显的性能差距。我们确定了三个不同瓶颈:其一,感知状态构建与解释存在挑战,智能体难以整合碎片化的 glimpses;其二,智能体常未完成完整序列就停止探索;其三,模型即便面对后续矛盾证据,也常无法修正早期错误信念。这些结果表明,仅获取视觉证据不足够,智能体还必须能够构建并更新可靠的感知状态。

英文摘要

AI agents in partially observable environments need to coordinate active sensing with working memory to maintain an evolving perceptual state. However, existing benchmarks struggle to isolate this perceptual-state construction and interpretation capability because they introduce physical and control complexities. We address this with MNIST-PRO, a benchmark that isolates agentic perception by converting MNIST digit recognition into a sequential, glimpse-based search task with lookback constraints. We evaluate ten multimodal models across four memory representations, including raw visual history, textual states, structured metric grid maps, and a consolidated visual canvas. While models excel under full observability, partial observability exposes a clear performance gap. We identify three distinct bottlenecks. First, perceptual-state construction and interpretation present a challenge, as agents struggle to integrate fragmented glimpses. Second, agents often stop exploring before they see the full sequence. Third, models often fail to revise early, incorrect beliefs even when faced with subsequent contradictory evidence. These results show that simply acquiring visual evidence is not enough. Agents must also be able to build and update a reliable perceptual state.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑