arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

OliveGemma:一款用于识别地中海与欧洲饮食的30亿参数视觉语言模型

OliveGemma: A 3 Billion Visual Language Model for Recognising the Mediterranean & European Diet

Dimitrios I. Zaridis, Traianos Tsiokris, Vasileios C. Pezoulas, Daphni Plati, Eugenia Mylona, Eleni Georga, Nikos Tsiknakis, Antonis Sakellarios, Dimitrios I. Fotiadis

arXiv 2608.03428首次发表:更新:

AI 中文总结

本研究构建基于PaliGemma-2-3B的OliveGemma,经LoRA微调后在欧洲饮食识别任务上超越多数前沿模型,仅逊于DenseNet-121,小型VLM经PEFT适配可在专业任务超越大模型。

AI 中文摘要

基于图像的饮食评估为自我报告食物日记提供了可扩展的替代方案,但由于类内高变异性和视觉相似的菜肴,细粒度食物识别仍然具有挑战性。本研究提出了OliveGemma,一款用于识别和推理地中海与欧洲菜肴的视觉语言模型。该模型基于开放权重的PaliGemma-2-3B架构构建,采用LoRA(低秩适配)在统一语料库上进行微调,该语料库包含来自三个欧洲研究项目数据集(MedGR、ODIN和VIPPSTAR)的17340张图像,被整理为包含216种复合菜肴类别的词汇,并配有102642个指令式问答项,涵盖菜肴识别、可能及可见食材、类别边界判别、视觉证据和整体视觉食物理解。在3折交叉验证方案下,OliveGemma的Top-1准确率为92.96%±0.91%,比最强的CNN基线模型DenseNet-121高出7.31%,并在精确指令和限定类别下,超过零样本前沿模型Gemini Flash 3和3.5、GPT-5.4 Mini以及Claude Haiku 4.6,分别提升8%、46%和64%。此外,OliveGemma在Top-3和Top-5准确率上表现具有竞争力,在CNN和前沿模型中排名第二,仅落后于DenseNet-121。另外,OliveGemma在食物类别的可能食材识别上达到90.79%±1.3%的Exact-Set指标。这些结果表明,对小型视觉语言模型(VLM)进行PEFT(参数高效微调)适配,可在专业食物识别任务上超越规模大得多的专有模型。该模型可在指定URL公开获取,实验和结果也可在该URL查看。

英文摘要

Image based dietary assessment offers a scalable alternative to self reported food diaries, yet fine-grained food recognition remains challenging due to high intra-class variability and visually similar dishes. This study presents OliveGemma, a vision language model for recognising and reasoning about Mediterranean and European cuisine. Built on the open-weight PaliGemma-2-3B architecture, OliveGemma is fine-tuned with LoRA on a unified corpus of 17,340 images from three European research project datasets (MedGR, ODIN, and VIPPSTAR), reconciled into a vocabulary of 216 composed dish categories and paired with 102,642 instruction style question-answer items covering dish recognition, likely and visible ingredients, class boundary discrimination, visual evidence and overall visual food understanding. Under a 3-fold cross-validation scheme, OliveGemma achieves a top-1 accuracy of 92.96% +/- 0.91%, exceeding the strongest CNN baseline (DenseNet-121) by 7.31% and outperforming zero-shot frontier models with exact instructions and bounded classes including Gemini Flash 3 and 3.5, GPT-5.4 Mini, and Claude Haiku 4.6 by 8%, 46%, and 64% respectively. Furthermore, OliveGemma demonstrates competitive performance on Top-3 and Top-5 accuracy, being second best across CNNs and frontier models, surpassed only by DenseNet-121. In addition, OliveGemma achieves 90.79% +/- 1.3% Exact-Set on the likely ingredients of the food categories. These results demonstrate that PEFT adaptation of a small VLM can surpass substantially larger proprietary models on specialised food recognition. The model is publicly available at https://huggingface.co/JamesZar/OliveGemma-3B and the experiments and results can be found at https://github.com/tsiokris/OliveGemma.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑