arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

基于马尔可夫决策过程的罗宾斯问题算法

Algorithms for Robbins' Problem using Markov Decision Processes

Léonard Brice, F. Thomas Bruss, Anirban Majumdar, Jean-François Raskin

arXiv 2608.27419首次发表:更新:

AI 中文总结

本文将罗宾斯问题建模为无限马尔可夫决策过程,提出有限状态抽象以近似问题值,给出n≤100的更优近似结果,明确了近最优策略所需的简单记忆结构。

AI 中文摘要

本文研究罗宾斯问题,它是著名秘书选择问题的全信息变体。该问题的目标是在依次面试的n名候选人中,最小化所选候选人的期望排名,且需在面试完第m名候选人后立即做出选择或不选的决定(无法看到剩余n-m名候选人,也无法召回)。我们首先展示如何将罗宾斯问题的实例建模为无限马尔可夫决策过程(MDP),然后提出这些MDP的若干有限状态抽象,以近似固定n时该问题的值。已知要最小化期望最优排名需保留过去候选人值的完整记忆,这使问题分析颇具挑战性,我们则指出了足以获得近最优选择策略的简单记忆结构。此外,我们给出了候选人数量n≤100时罗宾斯问题的近似值,此前针对这类n尚无良好近似方法(仅n≤4时已知精确值,数值近似仅针对n不超过一位数的小值);对于5≤n≤100的所有n,我们给出了比此前更优的近似结果。

英文摘要

In this paper, we consider Robbins' problem, which is a full information variant of the well-known secretary selection problem. In this version of the problem, the goal is to minimize the expected rank of the selected candidate among $n$ that are interviewed sequentially, and a decision to select or not the $m^{th}$ candidate needs to be taken right after the interview (so without seeing the last $n-m$ candidates and without recall). We first show how to model instances of Robbins' problem as infinite Markov Decision Processes (MDPs). Then we propose several finite-state abstractions of these MDPs that allow us to approximate the value of the problem for fixed $n$. While it is known that the full memory of past candidates' values is necessary for optimal expected rank minimization, making the analysis of the problem challenging, we highlight simple memory structures that are sufficient for obtaining near-optimal selection strategies. Additionally, we provide approximate values for Robbins' problem for numbers of candidates $n$ up to 100 for which no good approximations were previously known (the exact value is only known for instances where $n \leq 4$ and numerical approximations were for small values of $n$ not exceeding one digit), for all $n : 5 \leq n \leq 100$, we give better approximation than what was previously known.

CommentsExtended version of article published in Principles of Verification

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑