arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.35160cs.AI

FONDANT:通过反链实现强规划与尽力而为规划

FONDANT: Strong and Best-Effort Planning via Antichains

Benjamin Aminof, Tuan Khai Nguyen, Sasha Rubin

首次发表
浏览论文内容

中文总结 AI 辅助

提出FONDANT规划器,基于反链表示状态集合,统一生成强策略和尽力而为策略,实现可靠完备的强规划与尽力而为规划,在覆盖率上优于现有规划器。

中文摘要 AI 辅助

在完全可观测非确定性(FOND)规划中,一个经典的解概念是强策略(在密切相关的反应式综合领域中亦称为获胜策略),即这样的策略确保在对抗性环境中达到目标。当强策略不可用或没有证据表明环境是对抗性时,可以求助于尽力而为策略,这种策略总是存在,并遵循经典的决策理论原则,即智能体不应使用被支配的策略。一种典型的基于位置的尽力而为策略的工作方式如下:从每个状态出发,如果该状态存在强策略(此类状态称为“强获胜”),则遵循强策略;否则,如果该状态存在弱策略(“弱获胜”),则遵循弱策略;否则,该状态不受约束(“失败”)。在这项工作中,我们引入了一个既适用于尽力而为规划又适用于强规划的可靠且完备的规划器。支撑该规划器的算法相当简单:它通过其⊆-最小元素来表示某些状态集合,例如获胜区域。该算法返回统一策略,即它返回一个策略π_t,该策略是从每个强获胜状态出发的强解,并返回一个策略π_w,该策略是从每个弱获胜状态出发的弱解,同时为失败状态集合提供证书。我们用一些简单的优化实现了该算法(称之为FONDANT),并在一个基准集上进行了评估,该基准集包含用于评估领先强规划器PR2和FOND-SAT以及尽力而为规划器BeSyftP的实例。在覆盖率方面,我们的实现在所有领域上至少同样好,并在某些领域上表现更优;在墙钟时间方面,它在中小型实例上较慢,但在较大实例上表现更优。

英文摘要

A classical solution concept in fully observable nondeterministic (FOND) planning, is the strong policy (aka winning strategy in the closely related area of reactive synthesis), i.e., such a policy ensures that the goal is reached in an adversarial environment. When strong policies are not available or there is no evidence that the environment is adversarial, one can resort to best-effort policies, which always exist, and which follow the classic decision-theoretic principle that an agent should not use a dominated strategy. A typical positional best-effort policy works as follows: from every state, it follows a strong policy if one exists from that state (such states are called ``strong-winning''), else a weak policy if one exists from that state (``weak-winning''), and else is unconstrained (``losing''). In this work, we introduce a sound and complete planner for both best-effort planning and strong planning. The algorithm that underpins the planner is quite simple: it represents certain sets of states, such as the winning regions, by their $\subseteq$-minimal elements. The algorithm returns uniform policies, i.e., it returns a policy $π_t$ that is a strong solution starting in every strong-winning state, and it returns a policy $π_w$ that is a weak solution starting in every weak-winning state, and it provides a certificate for the set of losing states. We implemented the algorithm with some simple optimizations (calling it FONDANT), and evaluated it on a benchmark set consisting of the instances that were used in the evaluation of leading strong planners PR2 and FOND-SAT, and the best-effort planner BeSyftP. On coverage, our implementation is at least as good on all domains, and outperforms on some domains; and on wall time, it is slower on small and medium-sized instances, and outperforms on larger instances.

发表机构

  • The University of Sydney(悉尼大学)
  • Technical University of Vienna(维也纳技术大学)

机构由 AI 辅助整理,请以论文原文为准。

↑