arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2610.11696cs.AI

面向智能序贯决策的三维表征框架

A 3D Characterization Framework for Intelligent Sequential Decision Making

Sadig Gojayev, Carolina Fortuna

首次发表
浏览论文内容

中文总结 AI 辅助

该研究提出三维表征框架,以汉诺塔谜题为基准,对比Neurosolver、FBRL、AutoToS等不同范式的序贯决策方法,发现LLM方法因动作空间约束弱导致成本更高。

中文摘要 AI 辅助

谜题被广泛用于评估人工智能(AI)系统的序贯决策推理能力,但来自不同范式的方法很少在统一条件下进行比较。为解决这一缺口,我们提出一种三维表征框架,该框架可通过以下三个维度助力AI方法分析:1)将其映射到马尔可夫决策过程(MDP)序贯决策形式体系;2)通过对其设计的人类先验排序体现自主程度;3)技能与计算成本。利用该框架,我们分析了代表性的基于图、强化学习及大语言模型(LLM)的方法在设计选择与性能特征上的差异,分别以Neurosolver、前向后向强化学习(FBRL)及自动思维搜索(AutoToS)(含思维搜索的双智能体扩展DA-ToS)为例进行实例化。分析依托汉诺塔谜题展开,该谜题提供了规则明确、复杂度可扩展的受控基准,支持在不断增大的问题规模下进行一致比较。三维表征显示,基于LLM的方法因动作空间设计约束较弱,将复杂度从架构转移至推理时验证,导致其内存与运行时成本显著高于Neurosolver与FBRL。

英文摘要

Puzzles are widely used to evaluate the reasoning capabilities of artificial intelligence (AI) systems for sequential decision making, yet approaches originating from different paradigms are rarely compared under unified conditions. To address this gap, we introduce a three-dimensional characterization framework that enables the analysts of AI methods by 1) projecting them to the Markov decision process (MDP) sequential decision making formalism, 2) degree of autonomy through human prior ranking of their designs and, 3) skill and computational cost. Using this framework, we analyze how representative graph-based, reinforcement learning, and large language model (LLM)-based approaches differ in their design choices and performance characteristics, instantiated respectively by Neurosolver, forward-backward reinforcement learning (FBRL), and automated thought-of-search (AutoToS), including a double-agent extension of thought-of-search (DA-ToS). The analysis relies on the Tower of Hanoi puzzle that provides a controlled benchmark with well-defined rules and scalable complexity, enabling consistent comparison across increasing problem sizes. The 3D characterization reveals that LLM-based methods, due to their weakly constrained action-space design, shift complexity from architecture to inference-time verification, leading to substantially higher memory and runtime costs than Neurosolver and FBRL.

发表机构

  • Jožef Stefan Institute(若热夫· Stefan 研究所)

机构由 AI 辅助整理,请以论文原文为准。

↑