AI 中文总结
本研究探讨不完整训练数据下的主动特征获取问题,分析三种学习方法,发现生成式恢复能有效恢复多步获取的性能损失。
AI 中文摘要
在许多预测任务中,获取所有特征可能代价高昂甚至完全不可能。此外,在许多情况下,静态的特征子集可能不足以在各种实例上充分解决问题。主动特征获取(AFA)通过在顺序特征选择过程中形式化特征成本与预测性能之间的权衡来解决这些问题。然而,先前的AFA工作大多假设可以访问完整的训练数据,这一假设在实践中经常被违反。在此,我们研究不完整训练数据下的AFA(AFA-ITD),表明在完全随机缺失(MCAR)数据下,一步获取值保持不变,而多步获取值可能降低。我们分析了三种从不完整数据中学习的方法:别名化、过滤和生成式恢复。我们表明,过滤可能需要训练实例数量随维度呈指数增长,而生成式恢复则随获取预算呈指数增长。我们在受控实验和常见AFA数据集上实证检验了我们的理论,发现缺失主要损害利用多步获取的方法,并且生成式恢复能够在许多实验中恢复丢失的性能。代码可从此https URL获取。
英文摘要
In many prediction tasks, acquiring all features can be a prohibitively expensive or outright impossible task. Further, in many cases a static subset of features may not be enough to solve the problem sufficiently across various instances. Active Feature Acquisition (AFA) addresses these problems by formalizing the trade-off between feature cost and predictive performance during sequential feature selection. However, prior AFA work largely assumes access to complete training data, an assumption that is often violated in practice. Here we study AFA with Incomplete Training Data (AFA-ITD), showing that under missing completely at random (MCAR) data, one-step acquisition values remain unchanged, whereas multi-step values can decrease. We analyze three approaches to learning from incomplete data: aliasing, filtering, and generative restoration. We show that filtering can require a number of training instances scaling exponentially with the dimension, whereas generative restoration scales exponentially with the acquisition budget. We empirically test our theory on a controlled experiment and across common AFA datasets and find that missingness mainly damages methods that exploit multi-step acquisitions and that generative restoration is able to recover lost performance in many experiments. Code is available at https://github.com/Linusaronsson/AFA-Benchmark/tree/missing-data.
Comments24 pages, 8 figures, 2 tables