发表机构
Elmore Family School of Electrical and Computer Engineering, Purdue University(普渡大学埃尔莫尔家族电气与计算机工程学院)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
该研究针对单图像营养估计易遗漏食物或区域无法支持分量估计的问题,提出基于多模态大语言模型的框架,通过食物项验证与恢复提升营养估计精度,且无需任务特定微调。
AI 中文摘要
单图像营养估计在遗漏可见食物时可能静默失效;即使食物被正确识别,其提议区域也可能无法支持分量估计。本文提出一种框架,利用多模态大语言模型(MLLM)对可见食物进行清点,并分别验证食物身份及每个提议的2D区域是否支持分量估计。一次全图审查会利用这些验证结果识别未解决的缺口与遗漏食物,触发至多一次针对性恢复过程。恢复的区域在无法访问恢复提示的情况下被重新验证,随后被整合为用于营养估计的最终食物项集合。该框架无需特定任务的微调。对常见有效输出样本的匹配评估显示,项级定位在所有测试设置中提升了质量精度,且相较于适配的检索基线提升了能量精度,同时项身份的精确率与召回率也有所提升;恢复后的视觉覆盖度在推理时无需真值标注即可单独评估。
英文摘要
Single-image nutrition estimation can fail silently when visible foods are missed. Even when a food is correctly identified, its proposed region may not support portion estimation. We propose a framework that uses multimodal large language models (MLLMs) to inventory visible foods and separately verify food identity and whether each proposed 2D region supports portion estimation. One whole-image review uses these verification results to identify unresolved gaps and omitted foods, triggering at most one targeted recovery pass. Recovered regions are re-verified without access to the recovery prompt, then reconciled into a final item set for nutrition estimation. The framework requires no task-specific fine-tuning. Matched evaluation on common valid-output samples shows that item-level grounding improves mass accuracy across all tested settings and energy accuracy relative to an adapted retrieval baseline, with item-identity precision and recall also improving, while post-recovery visual coverage is assessed separately at inference time without ground-truth annotations.
Comments5 pages, 2 figures, 3 tables. Submitted to IEEE ICASSP 2027