arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.29724cs.AI

预算阈值激励请求的通用框架

A General Framework for Budgeted Threshold Incentives on Request

Zhuolin Wu, Chengrui Zhu, Wenhua Nie, Kenny Ye Liang, Junming Lin, Haiyang Li, Zhilin Li, Wenjia Geng, Zeyu Wu, Yinan Wu, Jinghua Hao, Renqing He

首次发表
浏览论文内容

中文总结 AI 辅助

针对按需配送平台预算阈值激励请求,提出一个由四个阶段和七个可替换模块组成的通用框架,通过响应校正和轨迹整合,在保证价值损失有界的同时显著提升响应速度并降低遗憾。

中文摘要 AI 辅助

按需配送平台通过激励活动向骑手支付报酬,其层级是根据具有相似历史的骑手近期完成情况设定的。运营方常常针对节假日或恶劣天气等场景,为不同时期、骑手群体、支付规则和预算请求此类方案,而随机试验在这些场景下稀缺且需要数月才能收集数据。我们提出一个请求驱动的框架,该框架通过七个可替换模块组合四个阶段(条件预测、群体缩减、轨迹整合和预算分配),这些模块交换条件轨迹规律,其奖励概率和奖励标记时刻为任何活动规则提供支付和提升。一个响应校正步骤对来自大量无报价历史的轨迹进行重新加权,以匹配短期试点的矩。我们证明,在固定方案菜单和给定阶段误差的情况下,端到端价值损失受四个阶段项之和约束,并且对于每个阶段,都存在省略该阶段会留下其他阶段无法消除的误差下限的实例。在45个周起始点上的3000名骑手中,一周内所有127个窗口的响应速度提高了11.04倍,且场景相同,分配导致的价值损失最多为0.92%。在24个新的受控响应规律上,使用一周试点的响应校正相对于具有相同名义随机骑手周的试验,遗憾降低了51.2%,而使用精确求和的四周试点与18周试验相比,差距在+0.007以内。在注册研究中,当窗口、群体、规则和约束预算随请求变化时,该框架的遗憾低于具有相同名义骑手周的试验,也低于相同试点数据的剂量插值,并且重用其一次性准备以14.1倍和2.70倍的速度回答了60个请求,且答案相同。与使用框架自身剂量曲线拟合的九报价试验相比,一周遗憾降低了0.055。

英文摘要

On-demand delivery platforms pay riders through incentive activities whose tiers are set from recent completions of riders with a similar history. Operators request such plans for changing periods, rider populations, payment rules and budgets, often for holidays or bad weather, where randomized trials are scarce and take months to collect. We present a request-driven framework that composes four stages (conditional prediction, population reduction, trajectory integration and budget allocation) through seven replaceable modules that exchange conditional trajectory laws, whose award probabilities and award-marked moments give payment and uplift for any activity rule. A response-correction step reweights trajectories from abundant no-offer history to match the moments of a short pilot. We prove that, on a fixed plan menu and given the stage errors, the end-to-end value loss is bounded by the sum of four stage terms, and that for every stage there are instances on which omitting it leaves an error floor the others cannot remove. On 3,000 riders over 45 weekly origins, all 127 windows of a week are answered 11.04x faster with identical scenarios and at most 0.92% value lost by the allocation. On 24 new controlled response laws, the response correction with a one-week pilot lowers regret by 51.2% relative to a trial with the same nominal randomized rider-weeks, and a four-week pilot with exact summation comes within +0.007 of an 18-week trial. In registered studies where windows, populations, rules and binding budgets change from request to request, the framework's regret is below that of a trial with the same nominal rider-weeks and below dose interpolation of the same pilot data, and reusing its one-off preparation answers 60 requests 14.1x and 2.70x faster with identical answers. Against a nine-offer trial fitted with the framework's own dose curve, one-week regret is 0.055 lower.

发表机构

  • Meituan(美团)
  • National Taiwan University(国立台湾大学)
  • Tsinghua University(清华大学)
  • Tongji University(同济大学)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑