arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.38096cs.LG

尾部影响采样用于CVaR策略评估

Tail-Influence Sampling for CVaR Policy Evaluation

Pauline Bourigault, Xiaotong Ji, Matthieu Zimmer, Rasul Tutunov, Haitham Bou-Ammar

首次发表
浏览论文内容

中文总结 AI 辅助

针对CVaR策略评估,提出尾部影响采样(TIS)方法,通过分配查询预算至尾部关键核,达到oracle渐近方差,在CliffWalking和语言模型审查任务中显著降低均方误差。

中文摘要 AI 辅助

具有相似平均回报的策略在罕见失败上可能差异巨大,然而准确估计下尾条件风险价值(CVaR)可能需要大量昂贵的rollout。当随机工作流的不同条件组件可以被分别查询时,我们研究如何分配固定的评估预算以最准确地估计固定策略的CVaR。我们为每个可查询的条件律推导出一个尾部影响,它聚合了其不确定性如何通过每次Bellman重用影响CVaR。其方差给出了固定设计效率界限和oracle Neyman分配。尾部影响采样(TIS)从试点模型估计这些影响尺度,并将新查询重新分配给对尾部最重要的核;一种基于访问锚定的变体防止试点分配不足。在固定维度和正分位数边际下,TIS达到oracle渐近方差和包括试点成本在内的一阶均方误差,而锚定变体在oracle的2倍因子内。我们还刻画了一个精确网格机制,其中尾部最优和均值最优分配重合。在CliffWalking上,TIS在相同计费转换预算下,相比学习占用减少41%的MSE,相比完整rollout减少76%。在冻结语言模型审查工作流中,锚定TIS在23/24个MMLU-Pro设置中优于同等正则化的均值影响混合,并在六次调用的FinQA审查中达到比rollout低2.4-3.4倍的MSE。

英文摘要

Policies with similar mean returns can differ sharply in rare failures, yet estimating lower-tail conditional value-at-risk (CVaR) accurately can require many costly rollouts. When different conditional components of a stochastic workflow can be queried separately, we ask how to allocate a fixed evaluation budget to estimate a fixed policy's CVaR most accurately. We derive a tail influence for each queryable conditional law that aggregates how its uncertainty affects CVaR across every Bellman reuse. Its variance yields the fixed-design efficiency bound and the oracle Neyman allocation. Tail-Influence Sampling (TIS) estimates these influence scales from a pilot model and reallocates fresh queries toward kernels that matter most for the tail; a visitation-anchored variant protects against pilot underallocation. Under fixed dimension and a positive quantile margin, TIS attains oracle asymptotic variance and first-order MSE including pilot cost, while the anchored variant is within a factor two of the oracle. We also characterize an exact-grid regime in which tail- and mean-optimal allocations coincide. On CliffWalking, TIS reduces MSE by 41% versus learned occupancy and 76% versus complete rollouts at the same charged transition budget. In frozen language-model review workflows, anchored TIS beats an equally regularized mean-influence blend in 23 of 24 MMLU-Pro settings and reaches 2.4-3.4$\times$ lower MSE than rollouts on six-call FinQA reviews.

发表机构

  • Imperial College London(帝国理工学院)
  • UCL Centre for AI(伦敦大学学院人工智能中心)

机构由 AI 辅助整理,请以论文原文为准。

↑