arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2607.25820cs.CV

基于大语言模型衍生成分标签和多模态融合的食品图像分割

Food Image Segmentation with LLM-Derived Ingredient Labels and Multimodal Fusion

  • National Cheng Kung University(国立成功大学)

机构由 AI 辅助整理,请以论文原文为准。

Jui-Feng Chi, Wei-Ta Chu, Sheng-Long Lin

AI总结:

针对现有食品图像分割模型在相似成分和罕见类别上表现不佳的问题,提出两个多模态模块LIM-F和LIM-Q,利用大语言模型衍生标签提升分割性能,在FoodSeg103基准测试中取得领先,且内存消耗增加适度,为细粒度食品理解提供实用方案。

AI中文摘要:

食品图像分割在营养追踪和个性化健康监测等健康相关应用中起着至关重要的作用。然而,现有模型在视觉上相似的成分和罕见食品类别上往往表现不佳。为解决此问题,我们提出两个即插即用的多模态模块,通过利用从食品图像中使用大语言模型推断出的成分标签来提高分割性能。第一个模块LIM-F旨在与产生多层输出的任何图像编码器配对,第二个模块LIM-Q针对基于Mask2Former的Transformer解码器。在FoodSeg103基准测试中,该方法取得了领先性能。将LIM-Q集成到带有Swin-L图像编码器的Mask2Former解码器中,平均交并比(mIoU)达到55.0。LIM-F也表现出强大的泛化能力和竞争力。该方法在训练期间GPU内存消耗仅适度增加(最多3.8GB),为细粒度食品理解提供了实用且可扩展的解决方案。

英文摘要:

Food image segmentation plays a vital role in health-related applications such as nutrition tracking and personalized health monitoring. However, existing models often underperform on visually similar ingredients and rare food categories. To address this issue, we propose two plug-and-play multimodal modules that enhance the segmentation performance by leveraging ingredient labels inferred from food images using large language models (LLMs). The first module, called LIM-F (Language Injection Module for Features), is designed to pair with any image encoder that produces multi-layer outputs (e.g., Swin Transformer), while the second module, LIM-Q (Language Injection Module for Queries), targets Mask2Former-style Transformer-based decoders. Both modules enable training without the need for pre-aligning images with text by directly injecting semantic ingredient information into the visual analysis pipeline. On the FoodSeg103 benchmark, the proposed method achieves state-of-the-art performance. Specifically, integrating LIM-Q into the Mask2Former decoder with a Swin-L image encoder yields a mean Intersection over Union (mIoU) of 55.0. LIM-F also demonstrates strong generalization and competitive performance, reaching an mIoU of 54.4 under the same model (Swin-L+Mask2Former). Furthermore, its applicability extends beyond Transformer-based decoders, as evidenced by an improvement from 47.7 to 49.8 mIoU when integrated into a CNN-based architecture. Notably, the improved segmentation accuracy is achieved with only a moderate (at most 3.8 GB) increase in the GPU memory consumption during training. Thus, the proposed approach offers a practical and scalable solution for fine-grained food understanding.

↑