JevForest:面向预算受限特征获取的路径投票方法
JevForest: Path Voting for Budgeted Feature Acquisition
查看机构详情
- University of Chinese Academy of Sciences(中国科学院大学)
机构由 AI 辅助整理,请以论文原文为准。
浏览论文内容
中文总结 AI 辅助
本文提出JevForest路径投票策略,用于预算受限的特征获取,在AG News、TREC、MiniBooNE等数据集上验证了其效果,发现其性能依赖于任务等因素,且直接Jev分类成本更低。
中文摘要 AI 辅助
选择观测哪些信息是观测预算有限情况下预测的核心问题。本文研究JevForest,这是一种特征获取策略,它聚合自举树的路径依赖提议,通过全局训练信息增益对这些提议加权,并使用共享掩码分类器从获取的特征值中进行预测。其在线实现会查询Jev以获取该策略选择的语义答案。在小型平衡保留样本上,针对AG News数据集(n=48),4个问题的森林获取方法准确率为0.729,而静态增益排序的准确率为0.667,随机排序的准确率为0.583;针对TREC数据集(n=24),排序结果反转:森林方法准确率为0.667,静态增益排序为0.750,随机排序为0.833。一次性批量询问全部8个问题,相比4次顺序森林查询,在更低的测量成本和延迟下获得更高准确率;直接Jev分类与批量准确率匹配,但成本更低。离线MiniBooNE实验在10个特征上的准确率为0.845±0.010,在40个特征上的准确率为0.885±0.008,该结果来自3个联合变化的数据和森林种子(均值±样本标准差)。配套的Newton boosting实现提供了初步的全特征合成结果。这些探索性发现确立了可行的Jev获取工作流程,但不支持路径投票具有普遍优势:其价值取决于任务、预测器,以及问题预算与实际查询成本之间的差异。
英文摘要
Choosing which information to observe is central to prediction under limited observation budgets. We study JevForest, a feature acquisition policy that aggregates path-dependent proposals from bootstrapped trees, weights them by global training information gain, and predicts from the acquired values with a shared masked classifier. An online implementation queries Jev for semantic answers selected by this policy. On small balanced held-out samples, four-question forest acquisition achieves accuracy $0.729$ on AG News ($n=48$), compared with $0.667$ for a static gain ranking and $0.583$ for random ordering. On TREC ($n=24$), the ordering reverses: forest accuracy is $0.667$, compared with $0.750$ and $0.833$. Asking all eight questions in one batch yields higher accuracy at lower measured cost and latency than four sequential forest queries; direct Jev classification matches the batch accuracy while costing less. Offline MiniBooNE experiments yield accuracy $0.845\pm0.010$ at ten features and $0.885\pm0.008$ at forty features over three jointly varying data and forest seeds (mean $\pm$ sample standard deviation). A companion Newton boosting implementation provides preliminary full-feature synthetic results. These exploratory findings establish a working Jev acquisition workflow but do not support a general advantage for path voting: its value depends on the task, predictor, and the distinction between question budgets and actual query costs.