arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

多尺度水果胶囊:用于野外可解释水果识别的空洞卷积与动态路由

Multi-Scale Fruit Capsules: Dilated Convolutions and Dynamic Routing for In-the-Wild Explainable Fruit Recognition

Subhankar Chattoraj, Sawon Pratiher, Samiran Das, Hubert Konik

arXiv 2608.21454首次发表:更新:

发表机构

University of Maryland Baltimore County; Indian Institute of Technology Kharagpur; Indian Institute of Science Education and Research Bhopal; UJM-Saint-Etienne; CNRS; Laboratoire Hubert Curien(马里兰大学巴尔的摩分校; 印度理工学院克勒格布尔分校; 印度科学教育与研究学院博帕尔分校; 圣埃蒂安大学; 法国国家科学研究中心; 于贝尔·居里安实验室)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

针对野外自动水果分类的类内宽变异与类间窄变异问题,提出FruitCapsNet胶囊网络,用空洞卷积结合动态路由,在四个数据集上性能优于10个迁移学习骨干网络,且决策归因更合理。

AI 中文摘要

同一水果可能成串、未采摘、去皮、装塑料袋或切片放在盘子中,因此野外自动水果分类(AFCW)必须应对形状、大小、颜色和纹理方面的大类内宽变异和类间窄变异。卷积网络通过池化路由信息,池化会丢弃感兴趣区域的姿态和位置,因此在这些表现形式上泛化能力较差。我们提出FruitCapsNet,这是一种胶囊网络,其水果胶囊用空洞卷积替代标准卷积前端:感受野以恒定参数成本呈指数增长,因此每个胶囊在动态路由解决部分-整体空间一致性之前编码多尺度上下文。超参数(包括膨胀因子)通过贝叶斯优化而非网格搜索选择。在三个公开数据集(SMP、FruitsGB、Fruits-360)和一个新的19类、10639张图像的野外数据集(PD-19)上,FruitCapsNet的深度仅为十分之三,却超过了10个微调的迁移学习骨干网络,在最难的数据集上优势最大(比最接近的竞争对手高2.7%)。从DigitCaps层传播的Grad-CAM显著性显示,改进源于将决策归因于整个水果区域而非物体边缘,这提供了事后证据,表明该改进并非数据集人工产物。

英文摘要

The same fruit appears in a bunch, unpicked, peeled, bagged in plastic, or sliced on a dish, so automated fruit classification in the wild (AFCW) must absorb wide intra- class and narrow inter-class variability in shape, size, colour and texture. Convolutional networks route information through pooling, which discards the pose and location of the region of interest and therefore generalises poorly across these presentations. We propose FruitCapsNet, a capsule network whose Fruit Capsules replace the standard convolutional front end with dilated convolutions: the receptive field grows exponentially at constant parameter cost, so each capsule encodes multi-scale context before dynamic routing resolves part whole spatial agreement. Hyper-parameters, including the dilation factor, are selected by Bayesian optimisation rather than grid search. On three public datasets (SMP, FruitsGB, Fruits-360) and a new 19-class, 10,639-image in-the-wild dataset (PD-19), FruitCapsNet exceeds ten fine-tuned transfer-learning backbones at one-third the depth, with the largest margin (+2.7% over the nearest competitor) on the hardest set. Grad-CAM saliency propagated from the DigitCaps layer shows that the improvement comes from attributing decisions to whole-fruit regions rather than to object edges, giving post-hoc evidence that the gain is not a dataset artefact.

CommentsAccepted in IEEE Industrial Electronics Conference (IECON) 2026

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑