NutriBench-Kitchen:具身人工智能营养管理基准
NutriBench-Kitchen: Benchmarking Embodied AI for Nutrition Management
浏览论文内容
中文总结 AI 辅助
针对具身厨房助手在动态场景中营养状态跟踪与规划能力评估缺失的问题,提出含1500个问答对的基准NutriBench-Kitchen及诊断智能体Nutri-Vgent,揭示现有模型与人类差距并验证结构化记忆的有效性。
中文摘要 AI 辅助
一个具身厨房助手必须做的不仅仅是识别孤立帧中的食物。它必须随时间跟踪食材状态,并将视觉观察与食谱和营养知识相结合,以支持受约束感知的决策。我们将这种能力形式化为“具身营养管理”:感知与营养相关的事件,维护持久的食物状态,并将其用于基于知识的规划。现有基准评估静态食物理解或具身烹饪动作,但不衡量智能体是否能在动态厨房中持续更新和使用与营养相关的状态。为填补这一空白,我们引入了NutriBench-Kitchen,一个包含来自160个烹饪视频的1500个手工验证的问答对的基准。它涵盖五个任务族:食材录入、记忆管理、食谱查询、长期规划和短期规划,跨越食物状态构建、维护、知识检索以及不同规划视野下的决策。对专有和开源大型视觉语言模型的评估显示,与人类表现存在显著差距,尤其是在定量食材估计、长期状态跟踪以及交互约束下的推理方面。我们进一步引入了Nutri-Vgent,一个具有独立情景、食物状态和食谱记忆的诊断性长视频智能体。其持续改进证明了显式状态表示和结构化记忆在营养管理中的价值。NutriBench-Kitchen和Nutri-Vgent共同为研究动态厨房中的持久状态跟踪和基于知识的推理提供了一个测试平台。代码可在https://此URL获取。
英文摘要
An embodied kitchen assistant must do more than recognize food in isolated frames. It must track ingredient states over time and integrate visual observations with recipe and nutritional knowledge to support constraint-aware decision-making. We formalize this capability as \emph{Embodied Nutrition Management}: perceiving nutrition-relevant events, maintaining a persistent food state, and using it for knowledge-grounded planning. Existing benchmarks evaluate static food understanding or embodied cooking actions, but do not measure whether an agent can continuously update and use nutrition-relevant states in dynamic kitchens. To fill this gap, we introduce \textbf{NutriBench-Kitchen}, a benchmark containing 1,500 manually verified question--answer pairs from 160 cooking videos. It covers five task families: Ingredient Entry, Memory Management, Recipe Query, Long-Term Planning, and Short-Term Planning, spanning food-state construction, maintenance, knowledge retrieval, and decision-making across different planning horizons. Evaluations of proprietary and open-source large vision-language models reveal a substantial gap from human performance, particularly in quantitative ingredient estimation, long-term state tracking, and reasoning under interacting constraints. We further introduce \textbf{Nutri-Vgent}, a diagnostic long-video agent with separate episodic, food-state, and recipe memories. Its consistent improvements demonstrate the value of explicit state representations and structured memory for nutrition management. Together, NutriBench-Kitchen and Nutri-Vgent provide a testbed for studying persistent state tracking and knowledge-grounded reasoning in dynamic kitchens. Code is available at https://github.com/V1ol1n/NutriBench-Kitchen.
发表机构
- Georgia Institute of Technology(佐治亚理工学院)
- Southern University of Science and Technology(南方科技大学)
- SpatialTemporal AI(时空人工智能公司)
机构由 AI 辅助整理,请以论文原文为准。