发表机构
Institute of Computing Technology, CAS; Hangzhou Dianzi University; North China Electric Power University; Lishui Institute of Hangzhou Dianzi University(中国科学院计算技术研究所; 杭州电子科技大学; 华北电力大学; 杭州电子科技大学丽水学院)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对现有食物分割基准测试无法捕捉现实用餐场景复杂性的问题,提出DishSeg24k基准测试,并基于此设计FEAST方法,通过将查询解码视为马尔可夫决策过程及重新设计解码器,提升分割性能,在多个指标上优于此前方法。
AI 中文摘要
食物分割对智能餐饮、饮食评估和推荐等应用至关重要。现有基准测试无法捕捉现实世界用餐场景的复杂性。为填补这一空白,我们引入了DishSeg24k,这是一个大规模的餐盘级分割基准测试,包含真实用餐环境中的24096张图像、112281个实例和278个细粒度类别。在此基础上,我们提出了食物专家自适应分割变换器(FEAST)来应对这些挑战。FEAST将基于查询的解码视为马尔可夫决策过程,重新设计了解码器,采用强化学习引导的专家混合模块,通过双评论家解耦优化方案分离任务导向的查询细化和结构感知的专家路由。实验表明FEAST性能优于之前方法,在DishSeg24k上mIoU提高3.21%、mDice提高3.68%、mAcc提高4.00%,在FoodSeg103上也验证了其有效性,数据集和代码将公开发布。
英文摘要
Food segmentation is essential for applications such as intelligent catering, dietary assessment, and recommendation. However, existing benchmarks fail to capture the complexity of real-world dining scenes. The challenges of dense inter-dish overlap, fine-grained class similarity, and extreme long-tail class distributions exceed the fidelity of current datasets. To fill this gap, we introduce \textbf{DishSeg24k}, a large-scale dish-level segmentation benchmark with 24,096 images, 112,281 instances, and 278 fine-grained categories in real-world dining environments. Based on DishSeg24k, we further propose \textbf{Food Expert-Adaptive Segmentation Transformers (FEAST)} to address these challenges. FEAST models query-based decoding as a Markov Decision Process (MDP), where each decoder layer update is treated as a sequential decision step that explores uncertainty along dish boundaries. We further redesign the decoder with a reinforcement learning (RL)-guided Mixture-of-Experts (MoE) module, in which a dual-critic decoupled optimization scheme separates task-oriented query refinement from structure-aware expert routing. This design promotes expert specialization and prevents expert collapse under long-tail category distributions. Finally, extensive experiments on DishSeg24k demonstrate the state-of-the-art performance of FEAST, which outperforms previous methods by {+3.21\%} mIoU, {+3.68\%} mDice, and {+4.00\%} mAcc, respectively. We further validate the effectiveness of FEAST on FoodSeg103. The dataset and code will be publicly released.
Comments9 pages, 8 figures. This paper has been accepted by ACMMM 2026