AI 中文总结
介绍OmniFood-Bench基准评估用于营养推理和个性化健康建议的VLMs,基于MM-Food-100K数据集,评估其三种能力,实验发现模型在菜品命名与质量估计等方面存在“语义-物理差距”,为公共卫生自主智能体建立可信度标准。
AI 中文摘要
大型视觉语言模型(VLMs)快速融入关键基础设施有望变革个性化医疗保健和饮食管理。但在食品系统领域,自主智能体面临视觉外观与内在营养成分间的“系统性信息不对称”挑战。现有基准主要关注粗粒度分类任务,无法评估现实饮食管理所需的复杂推理链。本文引入基于MM-Food-100K数据集构建的OmniFood-Bench综合基准,评估VLMs的三种递进能力:基本感知、定量推理和安全关键建议。对六个先进VLMs的评估发现了惊人的“语义-物理差距”。该工作为用于公共卫生的自主智能体建立了严格的可信度标准。
英文摘要
The rapid integration of Large Vision-Language Models (VLMs) into critical infrastructure promises to revolutionize personalized healthcare and dietary management. However, in the domain of food systems, autonomous agents face a unique and persistent challenge: the "Systemic Information Asymmetry" between visual appearance and intrinsic nutritional composition. Existing benchmarks primarily focus on coarse-grained classification tasks, such as food category recognition, which fail to evaluate the intricate reasoning chain required for real-world dietary management -- specifically, the ability to traverse from identifying hidden ingredients to estimating physical mass, and finally synthesizing safety-critical medical advice. In this paper, we introduce OmniFood-Bench, a comprehensive benchmark constructed from the MM-Food-100K dataset. Unlike previous works, OmniFood-Bench evaluates VLMs across three progressive capabilities: Basic Perception (Ingredients & Cooking Methods), Quantitative Reasoning (Portion Size & Nutritional Profiling), and Safety-Critical Advisory (Disease-Specific Recommendations). We evaluate six state-of-the-art VLMs, including gpt-5.1, gemini-3-flash, and qwen3-vl-8B. Our extensive experiments reveal a startling "Semantic-Physical Gap": while models achieve near-human accuracy in naming dishes, they exhibit catastrophic failure in mass estimation and frequently hallucinate benign advice for high-risk diabetic profiles. This work establishes a rigorous standard for trustworthiness in autonomous agents deployed for public health. The code and datasets are available in: https://anonymous.4open.science/r/OmniFood-Bench-7D0B