CulturalMenuBench:探究多模态烹饪推理中的知识应用差距
CulturalMenuBench: Probing the Knowledge-Application Gap in Multimodal Culinary Reasoning
浏览论文内容
中文总结 AI 辅助
该研究推出多模态烹饪推理基准CulturalMenuBench,发现模型在食物识别任务中存在知识应用差距,无法通过视觉输入激活文化知识,推动了关联感知、过程与文化背景的训练。
中文摘要 AI 辅助
多模态语言模型在食物识别基准测试中取得接近天花板的分数,但目前尚不清楚这种成功是反映了真正的文化理解,还是仅仅是视觉匹配。为探究这种区别,我们推出了CulturalMenuBench,这是一个涵盖18个地区、10种语言、共4870个条目的基准;其10项任务将成品菜肴和分步烹饪图像与食材、步骤文本及地区标签配对,范围从基础识别到基于过程的文化归因。对12个模型的评估揭示了显著的知识应用差距:在标准多项选择任务中超过94%的模型,在将菜肴归因于中国地方菜系时,准确率降至至多56%,尽管格式相同(四选一)。诊断分析解释了原因:错误模式与随机猜测一致,准确率追踪的是视觉独特性而非文化结构,且模型仅从菜肴名称分类菜系的准确率比从图像分类高7-18个百分点。因此知识是存在的,但无法通过视觉输入激活。一项 ablation 研究证实这些任务确实需要过程证据:移除连续烹饪图像会选择性降低基于过程的任务的性能,而其他任务保持稳定。总体而言,CulturalMenuBench表明近乎完美的识别可能掩盖了应用文化知识的能力缺失,推动了将感知、过程与文化背景明确关联的训练。代码和数据公开可用。
英文摘要
Multimodal language models achieve near-ceiling scores on food recognition benchmarks, yet it remains unclear whether this success reflects genuine cultural understanding or mere visual matching. To probe this distinction, we introduce CulturalMenuBench, a benchmark of 4,870 items in 10 languages across 18 regions; its 10 tasks pair final-dish and step-by-step cooking images with ingredients, procedural text, and regional labels, spanning basic recognition to process-grounded cultural attribution. Evaluating 12 models exposes a substantial knowledge-application gap: models exceeding 94% on standard multiple-choice tasks drop to at most 56% when attributing dishes to Chinese regional cuisines, despite an identical four-way format. Diagnostic analyses explain why: error patterns are consistent with random guessing, accuracy tracks visual distinctiveness rather than cultural structure, and models classify cuisines more accurately from dish names alone than from images (+7-18 points). The knowledge is thus present but cannot be activated through visual input. An ablation confirms these tasks genuinely require procedural evidence: removing sequential cooking images selectively degrades process-grounded tasks while others remain stable. Overall, CulturalMenuBench shows that near-perfect recognition can conceal an inability to apply cultural knowledge, motivating training that explicitly connects perception, procedure, and cultural context. Code and data are publicly available.
发表机构
- Alibaba Group(阿里巴巴集团)
- Xiamen University(厦门大学)
- University of Hamburg(汉堡大学)
机构由 AI 辅助整理,请以论文原文为准。