数据包络分析的统一最大间隔框架:理论基础与数值证据
A Unified Maximum-Margin Framework for Data Envelopment Analysis: Theoretical Foundations and Numerical Evidence
AI总结:
针对经典DEA的CCR模型区分能力不足的问题,提出MM-DEA模型,将间隔最大化融入效率框架,可转化为线性规划,在高维数据集上实现了决策单元的有效排序。
AI中文摘要:
数据包络分析(DEA)是一种广泛应用的非参数方法,用于评价决策单元(DMU)的相对效率。然而,经典的Charnes-Cooper-Rhodes(CCR)模型常常存在区分能力不足的问题,尤其是当输入和输出变量的数量相对于决策单元的数量较大时,许多决策单元会被评价为“CCR有效”且得分为1,这使得对表现最优者进行排序或区分变得困难。为解决这一局限,我们提出一种新颖的最大间隔DEA(MM-DEA)模型,将受机器学习中结构风险最小化启发的“间隔最大化”概念直接融入效率评价框架。与传统的超效率模型类似,所提MM-DEA模型在评价目标决策单元时会将其从参考集中排除;但与超效率模型不同,MM-DEA对目标决策单元自身的得分施加了明确的上界,确保所有得分均保持在标准的[0,1]范围内,而非可能超过1。我们从理论上证明,MM-DEA模型是CCR模型的广义扩展,当权衡参数α为0时,该模型会退化为CCR模型。此外,我们表明MM-DEA的分式规划形式可严格转化为线性规划(LP)问题,保证了计算效率。在数值实验中,MM-DEA对一个高维合成数据集生成了唯一的排序,而CCR模型未能区分该数据集中任何决策单元的差异。我们的结果表明,“间隔”是衡量管理韧性和竞争优势的关键指标。
英文摘要:
Data Envelopment Analysis (DEA) is a widely used non-parametric method for evaluating the relative efficiency of decision-making units (DMUs). However, the classic Charnes-Cooper-Rhodes (CCR) model often suffers from a lack of discriminatory power, particularly when the number of input and output variables is large relative to the number of DMUs. Under such conditions, many DMUs are evaluated as "CCR-efficient" with a score of unity, making it difficult to rank or distinguish between top performers. To address this limitation, we propose a novel Maximum-Margin DEA (MM-DEA) model that integrates the concept of "margin maximization" -- inspired by structural risk minimization in machine learning -- directly into the efficiency measurement framework. Like the traditional super-efficiency model, the proposed MM-DEA model excludes the target DMU from the reference set used to evaluate it; unlike the super-efficiency model, however, MM-DEA imposes an explicit upper bound on the target DMU's own score, ensuring that every score remains within the standard $[0,1]$ range rather than potentially exceeding one. We theoretically demonstrate that the MM-DEA model is a generalized extension of the CCR model, reducing to the latter when the trade-off parameter $α$ is zero. Furthermore, we show that the fractional programming formulation of MM-DEA can be rigorously transformed into a linear programming (LP) problem, ensuring computational efficiency. In our numerical experiments, MM-DEA produced a unique ranking for a high-dimensional synthetic dataset on which the CCR model failed to discriminate among any of the DMUs. Our results suggest that the "margin" serves as a critical metric for managerial resilience and competitive advantage.