结合邻近树与集成学习扩展数据无关的关键实例选择模型
Expanding Data-Agnostic Pivotal Instances Selection Models with Proximity Trees and Ensemble Learning
浏览论文内容
中文总结 AI 辅助
该研究提出一种分层可解释关键实例选择模型,结合邻近树、斜树与集成学习扩展数据无关的模型,在多类型数据集上性能优于同类策略,实现了可解释性与预测效果的平衡。
中文摘要 AI 辅助
随着决策过程日益复杂,机器学习工具已成为应对商业与社会挑战的必要手段。然而,许多现有方法依赖的决策过程难以解释。由于人类自然会通过将新案例与少量代表性示例比较来做决策,我们旨在设计一种选择此类关键实例(pivots)以构建可解释预测模型的方法。受决策树启发,我们提出一种分层的、天生可解释的关键实例选择模型,该模型基于关键实例与输入实例的相似性。我们的方法既可用作关键实例选择技术,也可作为独立预测模型。除单个关键实例外,我们还纳入邻近树、斜树及集成所使用的关键实例对,这提升了我们方案的通用性与有效性。此外,我们的方法与数据模态无关,可利用预训练网络进行数据转换。在表格数据、文本、图像、时间序列等各类数据集上开展的实验表明,我们的方法具有有效性,其性能优于其他实例选择策略,且在保持关键实例数量最少的同时,达到了与最先进可解释模型相当的结果。
英文摘要
As decision-making processes grow more complex, machine learning tools have become essential for tackling business and societal challenges. However, many existing methods rely on decision-making procedures that are difficult to interpret. Since humans naturally make decisions by comparing new cases with a few representative examples, we aim to design an approach that selects such pivots to construct an interpretable predictive model. Inspired by decision trees, we propose a hierarchical, interpretable-by-design pivot selection model based on the similarity between pivots and input instances. Our method functions both as a pivot selection technique and a standalone predictive model. Extending beyond single pivots, we incorporate pairs of pivots that are used by proximity and oblique trees, as well as ensembles, which enhance the versatility and effectiveness of our proposal. Additionally, our approach is data modality-agnostic, leveraging pre-trained networks for data transformation. Experiments across diverse datasets, including tabular data, text, images, and time series, demonstrate the effectiveness of our approach, outperforming alternative instance selection strategies and achieving competitive results against state-of-the-art interpretable models while maintaining a minimal number of pivots.