arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

使用机器学习指标提供对选择模型性能的更深入洞察

Using machine learning metrics to provide deeper insights into the performance of choice models

Lorenzo Muñoz, Stephane Hess, Thomas O. Hancock, Georges Sfeir

arXiv 2609.20655首次发表:更新:

发表机构

University of Leeds(利兹大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本文提出将机器学习中的概率混淆矩阵引入选择模型性能评估,以揭示整体拟合指标无法捕捉的备选方案级预测差异,指导模型设定并提升预测性能。

AI 中文摘要

机器学习(ML)技术日益引起选择建模(CM)领域的兴趣。以往的研究重点主要在于比较这些不同方法的性能,或改进机器学习技术的行为洞察,而非将思想从一个领域转化到另一个领域。在本文中,我们特别关注在模型性能评估背景下,将知识从机器学习转移到选择建模。在选择建模中,模型性能通常使用对数似然及相关指标进行评估,这些是关注整体拟合度的聚合拟合指标。相反,在机器学习中,重点在于备选方案级别的错误分类和正确分类,这提供了对结果更细致的视角。为弥合这些方法,我们探索使用概率版本的混淆矩阵,该矩阵报告模型预测每个备选方案的平均概率,条件是在所有选择任务中观察到选择了哪个备选方案。这使得能够为经典选择模型和机器学习算法计算概率机器学习指标。我们联合分析模型在整体拟合度和备选方案级别预测方面的性能。我们的研究结果表明,具有相似对数似然的模型可能表现出显著不同的混淆矩阵,揭示了聚合指标无法捕捉的不同概率模式。该框架识别出模型系统性地“混淆”备选方案的位置,突出备选方案之间的权衡,并可能指导模型设定。此外,在样本外评估这些矩阵和指标,揭示了显著影响预测性能的备选方案级别预测变化。

英文摘要

Machine learning (ML) techniques are increasingly drawing interest in the choice modelling (CM) field. The focus has primarily been on comparing the performance of these contrasting approaches or on improving behavioural insights for ML techniques, rather than translating ideas from one field into the other. In the present paper, we specifically focus on knowledge transfer from ML into CM in the context of model performance evaluation. In CM, model performance is typically evaluated using log-likelihood and related indicators, which are aggregate fit metrics that focus on overall fit. Conversely, in ML, the focus is on alternative-level misclassifications and correct classifications, which provide a more nuanced view of the results. To bridge these approaches, we explore the use of a probabilistic version of the confusion matrix, which reports the average probability of the model predicting each alternative, conditional on which alternative was observed to be chosen, across all choice tasks. This enables the computation of probabilistic ML metrics for both classic choice models and ML algorithms. We analyse model performance jointly in terms of overall fit and alternative-level predictions. Our findings demonstrate that models with similar log-likelihood can exhibit substantially different confusion matrices, revealing different probability patterns that aggregate metrics cannot capture. This framework identifies where models systematically `confuse' alternatives, highlighting trade-offs between alternatives, and potentially guiding model specification. Furthermore, evaluating these matrices and metrics out-of-sample reveals alternative-level prediction shifts that significantly impact forecasting performance.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑