发表机构
Purdue University(普渡大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
研究探讨多模态大语言模型用于膳食评估时检索增强基础方法是否仍有价值。提出无需训练、可本地部署的Open-KNEAD框架,通过营养感知检索关联食物项与FNDDS代码,提高份量估计,还能恢复非美国烹饪风格估计偏差,优势显著并开源相关框架与知识库。
AI 中文摘要
多模态大语言模型(MLLMs)越来越多地用于根据膳食图像进行饮食评估,检索增强的基础方法被证明可以提高营养估计的准确性。然而,研究发现当前的MLLMs这一前提不再成立,现代MLLM的直接估计现在已经匹配或超过了完整的检索流程。提出问题:如果检索不再能提高整体估计,它能否仍然提供临床医生重视的两件事,即准确的份量和可追溯的逐项记录?在保留临床应用重要因素的同时进行研究,引入了Open-KNEAD,一个无需训练且可本地部署的基于知识的智能膳食营养估计框架。通过选择性的、营养感知检索将每个分解的食物项与饮食研究的食品和营养数据库(FNDDS)代码相关联,组成可审计的逐项记录。在两个开放的MLLM家族和三种烹饪风格中,Open-KNEAD在大多数骨干数据集设置中比先前的基础方法和直接估计都提高了份量估计。智能内部食谱先验步骤进一步恢复了使非美国烹饪风格估计产生偏差的无形烹饪添加能量。在营养师验证的ACETADA数据集上优势最大,本地开放智能体比两个前沿封闭模型的直接份量估计分别高出约30%和53%,同时将所有膳食图像保留在本地硬件上。还发布了Open-KNEAD框架及其智能体就绪的FNDDS知识库。
英文摘要
Multimodal Large Language Models (MLLMs) are increasingly used for dietary assessment from meal images, where retrieval-augmented grounding was shown to sharpen nutrition estimates. However, we find this premise no longer holds for current MLLMs. A modern MLLM's direct estimate now matches or surpasses the full retrieval pipeline. This raises a question: if retrieval no longer improves the overall estimate, can it still deliver the two things clinicians value, accurate portions and a traceable, item-by-item record? We pursue this while preserving what matters for clinical adoption: minimal user burden (a single, unannotated meal image), explainability (an auditable record), and privacy (locally hosted inference). We introduce Open-KNEAD, a knowledge-grounded agentic framework for meal nutrition estimation that is training-free and locally deployable. Each decomposed food item is grounded to a Food and Nutrient Database for Dietary Studies (FNDDS) code via selective, nutrient-aware retrieval, composing an auditable per-item record. Across two open MLLM families and three cuisines, Open-KNEAD improves portion estimates over both prior grounding methods and direct estimation in most backbone-dataset settings. An agent-internal recipe-prior step further recovers the invisible cooking-added energy that biases estimates on non-US cuisine. The advantage is largest on the dietitian-verified ACETADA dataset, where the local open agent surpasses the direct portion estimates of two frontier closed models by roughly $30\%$ and $53\%$, all while keeping every meal image on local hardware. We release the Open-KNEAD framework and its agent-ready FNDDS knowledge base.
Comments10 pages main paper, 5 pages supplementary