arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2607.04028cs.LGcs.AI

用于分类性能评估的统一代数框架

A Unified Algebraic Framework for Classification Performance Evaluation

Ronaldo C. Prati

首次发表
浏览论文内容

中文总结 AI 辅助

提出统一代数框架评估分类性能,涵盖多种设置。以二值指示矩阵表示标签,结合三种聚合算子,可自动扩展二值性能度量到所有设置,还适应多种情况,建立了理论结果并经实证验证。

中文摘要 AI 辅助

我们提出了一个用于分类性能评估的统一代数框架,该框架在单一形式体系中涵盖了二分类、多分类、多标签、有序、层次、成本敏感和软标签设置。其基础是将实际标签和预测标签表示为二值指示矩阵,并结合三个聚合算子——全局、按列和按行——分别对应于微观、宏观/加权和示例平均。通过替换这些算子,任何以真/正/负计数表示的二值性能度量都能自动扩展到所有设置,无需针对特定度量进行推导即可生成多分类和多标签版本。该框架还通过argmax或阈值处理来适应软分类器输出,通过三角范数来适应软真值,通过隶属函数或累积编码来适应有序分类,以及通过包含MAE和MSE作为特殊情况的成本矩阵来适应成本敏感评估。我们建立了几个理论结果:微观平均等于分母加权宏观平均;乘积t范数是唯一保持混淆矩阵划分的范数;偏斜不变度量被表征为召回率和特异性的函数;在多分类设置中,微观精度、微观召回率和微观F1都等于准确率。在合成数据和真实数据上的实证说明证实了这些理论发现。

英文摘要

We propose a unified algebraic framework for classification performance evaluation covering binary, multiclass, multilabel, ordinal, hierarchical, cost-sensitive, and soft-label settings. Actual and predicted labels are represented as binary indicator matrices, where three aggregation operators (global, column-wise, row-wise) correspond directly to micro, macro/weighted, and exemplar averaging. Any binary measure expressed in terms of the four confusion-matrix counts extends to all these settings by substituting an operator, with no measure-specific derivation. We show that structural properties governing an extension are derivable from the binary formula. Micro-averaging equals denominator-weighted macro-averaging precisely for aggregation-decomposable (linear-fractional) measures, a strict class characterised algebraically. For soft ground truth, we prove from t-norm axioms alone that the product t-norm is the unique choice whose confusion counts preserve marginal memberships. In multiclass settings, micro-precision, micro-recall, and micro-F1 collapse identically onto accuracy. Furthermore, binary skew-invariance transfers unconditionally to multilabel aggregation, but only partially to multiclass problems. For measures with a linear numerator and prediction-independent denominator, the optimal decision threshold is the share of the numerator weight favouring a negative prediction, revealing when standard training targets the measure. Under one-hot multiclass encoding, a measure fails to attain its theoretical minimum whenever its zero-true-positive value still depends on true negatives, establishing non-trivial performance floors even for completely incorrect classifiers (e.g., zero correct predictions on 10 classes yields a label accuracy of 0.8).

发表机构

  • Universidade Federal do ABC (UFABC)(巴西联邦大学ABC分校)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑