发表机构
Columbia University(哥伦比亚大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对活动特征获取策略,提出用后验期望熵替代总预测熵作为评估奖励,以消除离线数据覆盖不平衡导致的认知偏差,实验证明可减少价值估计偏差并改善策略选择。
AI 中文摘要
活动特征获取学习策略,通过顺序获取特征来最大化关于目标变量的信息。我们研究如何利用先验数据拟合网络(PFNs)从有限的离线数据中学习和评估此类策略,PFNs是现成的模型,无需针对特定任务训练即可输出后验预测分布。我们表明,在离线数据覆盖不平衡的情况下,使用总预测熵作为奖励会产生认知偏差,惩罚获取稀疏观测特征。具体而言,这种奖励将认知不确定性(源于离线数据不足)与偶然不确定性(源于无信息特征)混为一谈。为解决此问题,我们针对PFN输出的后验期望(偶然)熵而非总预测熵来评估特征获取。在合成和真实世界数据集上的实证评估表明,我们的方法持续减少价值估计偏差,并产生具有强经验覆盖率的可信区间,这可转化为改进的下游策略选择。
英文摘要
Active feature acquisition learns policies that sequentially acquire features to maximize information about a target variable. We study how to learn and evaluate such policies from finite offline data using prior-data fitted networks (PFNs), which are off-the-shelf models that output posterior predictive distributions without task-specific training. We show that under the imbalanced coverage of offline data, using total predictive entropy as a reward creates an epistemic bias that penalizes acquiring sparsely observed features. Specifically, this reward conflates epistemic uncertainty (arising from lack of offline data) with aleatoric uncertainty (arising from uninformative features). To address this, we target the posterior expected (aleatoric) entropy instead of the total predictive entropy output by a PFN for evaluating feature acquisitions. Empirical evaluations on synthetic and real-world datasets demonstrate that our approach consistently reduces value estimation bias and yields credible intervals with strong empirical coverage, which can translate to improved downstream policy selection.