H$^2$EDL:用于分层分类的超证据深度学习
H$^2$EDL: Hyper Evidential Deep Learning for Hierarchical Classification
AI总结:
该研究针对细粒度分层分类的结构化歧义问题,提出H$^2$EDL模型,其在FGVC-Aircraft等数据集上使校准误差较交叉熵基线降低约一半,改进在深层级和大训练预算下更显著。
AI中文摘要:
细粒度识别常涉及分层标签空间,模型可能对粗粒度语义概念有信心,但对其后代类不确定,这种结构化歧义需要能同时捕捉细粒度类和中间概念的不确定性表示。然而现有工具各仅捕捉一半:平面证据分类器在叶框架用单一空虚量化总无知,分层分类器传播点概率却无证据概念;超意见本可统一两者,但其一般形式随标签数量呈指数增长,现有超证据网络要么需训练数据中提供复合标签,要么从非结构化权重模式读取,无原则性概念说明哪些复合应分配质量。我们观察到分类法本身是缺失的超域,其子树和叶单例形成线性大小的焦点族,每个分支节点的局部狄利克雷意见以闭式诱导每个复合质量。所得模型H$^2$EDL可通过同一组参数以两种互补方式解释:从预测角度,它是在标签树不同级别保持一致性的分层分类器;从概率角度,它定义了有效的树结构超意见,每个节点分配的质量代表到达该节点但无足够信心进一步分化为后代的信念。在FGVC-Aircraft和DERM12345数据集上,H$^2$EDL与交叉熵基线相比校准误差降低约一半,在更深的分层级别和更大训练预算下,改进更为显著。
英文摘要:
Fine-grained recognition often involves hierarchical label spaces, where a model may be confident about a coarse semantic concept while remaining uncertain among its descendant classes. Such structured ambiguity requires uncertainty representations that capture both fine-grained classes and intermediate concepts. However, existing tools each capture only half of it: flat evidential classifiers quantify total ignorance with a single vacuity on the leaf frame, and hierarchical classifiers propagate point probabilities with no notion of evidence. Hyper-opinions would unify the two, but their general form is exponential in the label count, and existing hyper-evidential networks either require composite labels to be supplied in the training data or read them off an unstructured weight pattern, with no principled notion of which composites deserve mass. We observe that the taxonomy itself is the missing hyperdomain. Its subtrees and leaf singletons form a linear-size focal family, and one local Dirichlet opinion per branching node induces every composite mass in closed form. The resulting model, H$^2$EDL, can be interpreted in two complementary ways using the same set of parameters. From a prediction perspective, it functions as a hierarchical classifier that preserves consistency across different levels of the label tree. From a probabilistic perspective, it defines a valid tree-structured hyper-opinion, where the mass assigned to each node represents the belief that reaches that node but does not provide sufficient confidence to further specialize into its descendants. On FGVC-Aircraft and DERM12345, H$^2$EDL reduces calibration error by approximately half compared with cross-entropy baselines, with the improvement becoming more pronounced at deeper hierarchy levels and under larger training budgets.