arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

通过目标感知数据对齐实现细粒度食品图像理解

Fine-Grained Food Image Understanding via Target-Aware Data Alignment

Jui-Feng Chi, Wei-Lun Chu, Bruce Coburn, Jinge Ma, Fengqing Zhu

arXiv 2607.25794首次发表:更新:

AI 中文总结

研究针对细粒度食品图像理解,提出以数据为中心的多模态对齐方法,先选视觉相关训练子集,再细化字幕,训练互补检索专家并融合决策,提升了检索性能,完整方法检索分数超纯VLM检索两倍且更高效。

AI 中文摘要

细粒度食品视觉语义理解要求模型捕捉食材、烹饪方法、熟度、颜色、质地和摆盘构成等方面的细微差别。尽管CLIP风格的视觉语言模型为此任务提供了自然框架,但依赖网络收集的异质图像文本对训练时效果有限。我们提出一种以数据为中心的多模态对齐方法,先进行目标感知数据选择,再用基于VLM的字幕细化生成视觉基础的目标风格描述,训练互补的CLIP风格检索专家并通过分层VLM辅助多专家决策级融合策略组合决策。实验表明数据细化策略显著提升检索性能,完整方法检索分数超纯VLM检索两倍且更高效。

英文摘要

Fine-grained food visual--semantic understanding requires models to capture subtle distinctions across ingredients, cooking methods, doneness, color, texture, and plate composition. Although CLIP-style vision-language models provide a natural framework for this task, their effectiveness is limited when training relies on heterogeneous web-collected image--text pairs. Such data often exhibit a web-to-target domain gap and cross-modal misalignment, where images differ from the target distribution and captions are noisy, multilingual, or weakly grounded in visual content. We propose a data-centric multimodal alignment method for fine-grained food description and recognition. Our method first performs target-aware data selection to identify visually relevant training subsets, then applies VLM-based caption refinement to generate visually grounded, target-style descriptions. Using these curated image--caption pairs, we train complementary CLIP-style retrieval experts and further combine their decisions through a hierarchical VLM-assisted multi-expert decision-level fusion strategy that invokes the VLM only when experts disagree. Experiments show that our data refinement strategy significantly improves retrieval performance over naive web supervision, with VLM-based caption refinement alone yielding an average performance gain of approximately 19%. Our full method also achieves more than twice the retrieval score of pure VLM-based retrieval while remaining substantially more efficient.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑