发表机构
CNRS, Univ. Grenoble Alpes, Grenoble-INP, GIPSA-lab; Université Sorbonne Paris Nord, LAGA, UMR 7539. IRL CNRS-CRM 3457. Université de Montréal(法国国家科学研究中心、格勒诺布尔阿尔卑斯大学、格勒诺布尔国立综合理工学院、信号与自动化格勒诺布尔实验室; 巴黎索邦大学北校区、巴黎第十三大学、法国国家科学研究中心 - 巴黎第十三大学联合研究单位7539、法国国家科学研究中心 - 蒙特利尔大学联合研究单位3457、蒙特利尔大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
研究基于真值表结构化数据集的二分类器最优线性组合,通过分类校准函数分析凸化经验风险,建立存在唯一性条件,给出三个分类器情况配置,推导最优权重公式,引入φ前沿评估稳定性和数据质量。
AI 中文摘要
本文基于真值表对数据集进行逻辑结构化研究二分类器的最优线性组合。给定分类器将数据划分为等价类,通过分类校准函数的多维泛化对凸化经验风险进行严格分析。我们为任意分类器列表的凸化经验风险最小值(全局)点的存在性和唯一性建立了充分条件。对于三个分类器的情况,分析列出所有导致唯一解、下确界或非唯一最小值点的配置。此外,利用指数(Boost)和逻辑(Logit)损失函数推导最优权重的显式解析公式,绕过迭代优化。通过引入φ前沿概念可评估所得分类器的稳定性和数据质量分析。
英文摘要
This paper studies an optimal linear combination of binary classifiers based on a logical structuration of the dataset via truth tables. The given classifiers partition data into equivalence classes, allowing for a rigorous analysis of the convexified empirical risk through a multidimensional generalization of classification calibrated functions. We establish sufficient conditions for the existence and uniqueness of the (global) point of minimum of the convexified empirical risk for any list of classifiers (when the number of classifiers is large, there frequently could be no point of minimum). In the case of three classifiers, our analysis allows to list all the configurations leading to either a unique solution, infima or non-unique points of minimum. Furthermore, we derive explicit analytical formulae for optimal weights using Exponential (Boost) and Logistic (Logit) loss functions, bypassing iterative optimization. The stability of the resulting classifier and the analysis of data quality can be evaluated through the introduction of the notion of $ϕ$-frontiers.
Comments35 pages