arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.03526cs.AI

CulturalMenuBench:探究多模态烹饪推理中的知识应用差距

CulturalMenuBench: Probing the Knowledge-Application Gap in Multimodal Culinary Reasoning

Bo Zeng, Linfeng Gao, Peiqin Lin, Yu Zhao, Mingyan Zeng, Yu Tong, Xintong Wang, Linlong Xu, Longyue Wang, Weihua Luo, Qinggang Zhang, Jinsong Su

首次发表
浏览论文内容

中文总结 AI 辅助

该研究推出多模态烹饪推理基准CulturalMenuBench,发现模型在食物识别任务中存在知识应用差距,无法通过视觉输入激活文化知识,推动了关联感知、过程与文化背景的训练。

中文摘要 AI 辅助

多模态语言模型在食物识别基准测试中取得接近天花板的分数,但目前尚不清楚这种成功是反映了真正的文化理解,还是仅仅是视觉匹配。为探究这种区别,我们推出了CulturalMenuBench,这是一个涵盖18个地区、10种语言、共4870个条目的基准;其10项任务将成品菜肴和分步烹饪图像与食材、步骤文本及地区标签配对,范围从基础识别到基于过程的文化归因。对12个模型的评估揭示了显著的知识应用差距:在标准多项选择任务中超过94%的模型,在将菜肴归因于中国地方菜系时,准确率降至至多56%,尽管格式相同(四选一)。诊断分析解释了原因:错误模式与随机猜测一致,准确率追踪的是视觉独特性而非文化结构,且模型仅从菜肴名称分类菜系的准确率比从图像分类高7-18个百分点。因此知识是存在的,但无法通过视觉输入激活。一项 ablation 研究证实这些任务确实需要过程证据:移除连续烹饪图像会选择性降低基于过程的任务的性能,而其他任务保持稳定。总体而言,CulturalMenuBench表明近乎完美的识别可能掩盖了应用文化知识的能力缺失,推动了将感知、过程与文化背景明确关联的训练。代码和数据公开可用。

英文摘要

Multimodal language models achieve near-ceiling scores on food recognition benchmarks, yet it remains unclear whether this success reflects genuine cultural understanding or mere visual matching. To probe this distinction, we introduce CulturalMenuBench, a benchmark of 4,870 items in 10 languages across 18 regions; its 10 tasks pair final-dish and step-by-step cooking images with ingredients, procedural text, and regional labels, spanning basic recognition to process-grounded cultural attribution. Evaluating 12 models exposes a substantial knowledge-application gap: models exceeding 94% on standard multiple-choice tasks drop to at most 56% when attributing dishes to Chinese regional cuisines, despite an identical four-way format. Diagnostic analyses explain why: error patterns are consistent with random guessing, accuracy tracks visual distinctiveness rather than cultural structure, and models classify cuisines more accurately from dish names alone than from images (+7-18 points). The knowledge is thus present but cannot be activated through visual input. An ablation confirms these tasks genuinely require procedural evidence: removing sequential cooking images selectively degrades process-grounded tasks while others remain stable. Overall, CulturalMenuBench shows that near-perfect recognition can conceal an inability to apply cultural knowledge, motivating training that explicitly connects perception, procedure, and cultural context. Code and data are publicly available.

发表机构

  • Alibaba Group(阿里巴巴集团)
  • Xiamen University(厦门大学)
  • University of Hamburg(汉堡大学)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑