AI 中文总结
NutriVision提出端到端框架,融合视觉几何与食材语义,从单张RGB图像估计营养,在Nutrition5k上平均PMAE达13.60%,优于现有方法,无需专用深度硬件。
AI 中文摘要
营养估计是消费者饮食追踪、临床营养学、慢性病管理、运动与医院营养以及更广泛的食品计算系统中的一项基本任务。现有方法主要沿着两条大致独立的路线发展:依赖校准的RGB-深度图像的视觉模型,以及利用文本线索但多模态融合有限的食材感知方法。我们提出了NutriVision,一个端到端框架,利用视觉几何和食材语义从单张RGB图像和可选的食材列表中估计卡路里、质量、脂肪含量、碳水化合物和蛋白质。它使用DepthAnything-V3获取不可用的深度模态,并使用CLIP编码食材描述。它整合了三种互补机制:(1)食材条件频率对齐融合模块(IC-FAFM),利用文本引导对RGB-深度频率分量进行重新加权和对齐;(2)食材感知掩码预测头(IA-MPH),其门控和通道掩码以食物身份为条件;(3)模态特定的内部语义建模(ISM)块。在Nutrition5k数据集上,NutriVision实现了平均PMAE为13.60±0.10%,优于我们实现的IGSMNet 0.89个百分点,优于OmniFood8k 2.90个百分点(两者p<0.001)。模块级消融实验确定食材感知预测头为主要架构贡献者,将平均PMAE改善了1.50±0.17个百分点(p<0.001)。这些结果表明,食材条件预测和频率感知的RGB-深度融合为单图像营养估计提供了可衡量的改进。更广泛地说,NutriVision为利用几何和语义线索而不需要专门深度传感硬件的营养评估系统提供了一条实用途径。
英文摘要
Nutrition estimation is a fundamental task in consumer diet tracking, clinical dietetics, chronic disease management, sports and hospital nutrition, and broader food computing systems. The existing approaches have progressed along two largely separate axes, vision models that rely on calibrated RGB-depth captures and ingredient-aware methods that use textual cues but use limited multimodal fusion. We introduce NutriVision, an end-to-end framework that leverages visual geometry and ingredient semantics to estimate calories, mass, fat content, carbohydrates, and protein from a single RGB image and an optional ingredient list. It obtains the unavailable depth modality using DepthAnything-V3 and encodes ingredient descriptions using CLIP. It integrates three complementary mechanisms: (1) an \emph{Ingredient-Conditioned Frequency-Aligned Fusion Module (IC-FAFM)}, which uses textual guidance to reweight and align RGB-depth frequency components; (2) an \emph{Ingredient-Aware Mask-based Prediction Head (IA-MPH)}, whose gating and channel masks are conditioned on food identity; and (3) modality-specific \emph{Internal Semantic Modeling (ISM)} blocks. On the Nutrition5k dataset, NutriVision achieves a mean PMAE of $\mathbf{13.60\pm0.10\%}$, outperforming our IGSMNet implementation by $0.89$ percentage points and OmniFood8k by $2.90$ percentage points (both $p<0.001$). The module-level ablations identify the ingredient-aware prediction head as the primary architectural contributor, improving mean PMAE by $1.50\pm0.17$ percentage points ($p<0.001$). These results demonstrate that ingredient-conditioned prediction and frequency-aware RGB-depth fusion provide measurable gains for single-image nutrient estimation. More broadly, NutriVision offers a practical route toward nutrition-assessment systems that exploit geometric and semantic cues without requiring specialized depth-sensing hardware