arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

通过离线规划学习源获取策略

Learning Source Acquisition Policies by Offline Planning

Ziqi Zhao, Run Xu, Qingjian Ni

arXiv 2609.14299首次发表:更新:

发表机构

Southeast University(东南大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本文提出O-MPAC方法,通过离线规划学习源获取策略,在预算约束下选择特征组,实验表明其在多个任务上优于现有方法。

AI 中文摘要

在获取预算限制下进行预测,需要选择特征组,其价值可能取决于后续查询。O-MPAC将有限时域的风险-成本目标从完整的训练记录转移到一个共享的源-动作评分器中。在推理时,评分器利用部分观测和源元数据,在每次查询后重新评分,并应用硬成本掩码。我们分析了绑定的教师目标和剩余规划时域如何影响学习到的决策。对绑定最小值进行统一监督,在源重标记下保持了目标分布。在五次种子的路由实验中,它在原始顺序和上下文最后顺序下均达到0.965的准确率。在六个真实任务上,验证在所有三十个分割中均选择了没有动作交叉熵的H1。在五个任务上,与源适应的GDFS、DIME、AACO+NN以及静态策略相比,O-MPAC具有最高的平均预算集成准确率。

英文摘要

Predicting under an acquisition budget requires choosing feature groups whose value can depend on later queries. O-MPAC transfers finite-horizon risk-cost targets from complete training records into a shared source-action scorer. At inference time, the scorer uses partial observations and source metadata, re-scores after each query, and applies a hard cost mask. We analyze how tied teacher targets and the remaining planning horizon affect the learned decisions. Uniform supervision over tied minima preserves the target distribution under source relabeling. In a five-seed routing experiment, it achieves 0.965 accuracy under both original and context-last orders. On six real tasks, validation selects H1 without action cross-entropy in all thirty splits. O-MPAC has the highest mean budget-integrated accuracy on five tasks against source-adapted GDFS, DIME, AACO+NN and a static policy.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑