arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

查询并非承诺:在线延迟决策中学习修正专家答案

A Query Is Not a Commitment: Learning to Correct Expert Answers in Online Deferral

Yannis Montreuil, Axel Carlier, Lai Xing Ng, Wei Tsang Ooi

arXiv 2610.07084首次发表:更新:

发表机构

National University of Singapore; Toulouse INP; Institute for Infocomm Research, A*STAR(新加坡国立大学; 图卢兹国立理工学院; 新加坡科技研究局资讯通信研究院)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

针对在线延迟决策中专家答案可能不准确的问题,提出ORUCB算法,通过共享和专家特定多项式响应及误差校准,实现高概率伪遗憾界,并在实验中降低成本。

AI 中文摘要

一个不准确的专家在修正后仍能提供有用信息。我们研究在线延迟决策(online learning to defer),其中学习者在购买答案前选择一位专家并确定一个修正函数,然后将该函数应用于收到的答案。难点在于观察到的损失既反映专家质量,也反映未完成的修正:早期错误可能阻碍那些在学习后会有价值的查询。我们提出ORUCB,它汇集了共享和专家特定的多项式响应。对累积响应学习误差的界限校准了置信度加权风险回归和探索,使路由器在决定购买哪些答案时能考虑该误差。在残差和有界分歧、最优响应的固定可行模型以及自由和最优查询风险的线性模型下,校准算法在固定问题参数下,在$T$轮内实现了高概率的伪遗憾$O(\sqrt T\log(T+1))$。该保证允许奇异的答案分布和错误指定的共享响应;最优性是相对于有界响应类而言的。在四个测试流上,选定的三次策略比七个直接使用答案的基线具有更低的含费成本。与常见修正学习者的比较考察了路由,而六种价格比较则衡量了成本和查询率。

英文摘要

An inaccurate expert can still provide useful information after correction. We study online learning to defer in which the learner chooses an expert and fixes a correction function before purchasing its answer, then applies that function to the answer received. The difficulty is that observed losses reflect both expert quality and an unfinished correction: early errors can discourage queries that would be valuable after learning. We propose ORUCB, which pools shared and expert-specific polynomial responses. A bound on cumulative response-learning error calibrates confidence-weighted risk regression and exploration, allowing the router to account for this error when deciding which answers to buy. Under bounded residuals and disagreements, a fixed feasible model of optimal responses, and linear models of free and optimal queried risk, the calibrated algorithm achieves high-probability pseudo-regret $O(\sqrt T\log(T+1))$ over $T$ rounds for fixed problem parameters. The guarantee permits singular answer distributions and misspecified shared responses; optimality is relative to the bounded response class. On four test streams, the selected cubic policy has lower fee-inclusive cost than seven baselines that deploy answers unchanged. Comparisons with a common correction learner examine routing, while six-price comparisons measure cost and query rates.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑