arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

空间策略,而非动作:向量量化测地线作为LLM驱动智能体的工具

Spatial Strategies, Not Actions: Vector-Quantized Geodesics as Tools for LLM-Driven Agents

Gabriel Turinici

arXiv 2610.00613首次发表:更新:

发表机构

CEREMADE CNRS; Université Paris Dauphine - PSL(CEREMADE 法国国家科学研究中心; 巴黎第九大学 - PSL)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本研究提出一种结合向量量化测地线工具库与LLM高层协调的架构,在网格环境中提升智能体空间理解,以低成本达到高目标到达率。

AI 中文摘要

基于大型语言模型(LLM)的智能体常因缺乏空间理解且主要利用统计文本模式而受到批评。我们通过一种将几何工具与作为高层协调器的LLM相结合的架构,在网格世界环境中研究其空间理解能力。智能体首先收集测地线轨迹,然后通过向量量化提取代表性子集。离线阶段,LLM将每条选定轨迹关联到底层行为模式的自然语言描述,使其成为工具。在线阶段,LLM根据当前状态和目标选择合适的工具。低级控制由执行与工具关联轨迹的原始动作处理。从智能体AI的角度看,这种方法将学习分为两个层次:工具发现通过轨迹的无监督量化处理,而推理和决策由LLM处理。我们在一个部分可观察的动态二维网格环境中使用开放视觉语言模型(Qwen3.6-35B-A3B)测试该方法。将几何派生的工具库与智能体中心的缩放工具和碰撞检测工具配对,使快速、非推理配置达到目标到达率,与更昂贵的思维链版本相当,同时将决策成本从分钟级降至秒级。

英文摘要

Large language model (LLM) based agents are often criticized for lacking spatial understanding and mainly exploiting statistical text patterns. We investigate their spatial comprehension through an architecture combining geometrical tools with a LLM serving as a high-level orchestrator in grid-world environments. The agent first collects geodesic trajectories, which are then vector-quantized to extract a representative subset. Offline, the LLM associates a natural language description of the underlying behavioral patterns to each selected trajectory, making it a tool. Online, the LLM chooses the appropriate tool conditioned on the current state and goal. Low-level control is handled by primitive actions that execute the trajectory associated with the tool. From an agentic AI perspective, this approach separates learning into two levels: tool discovery is handled through unsupervised quantization of trajectories, while reasoning and decision-making are handled by the LLM. We test the approach in a partially observable dynamic 2D grid environment with an open vision-language model (Qwen3.6-35B-A3B). Pairing the geometry-derived tool library with an agent-centered zoom tool and a collision detection tool lets a fast, non-reasoning configuration match the goal-reaching rate of a much more costly chain-of-thought version, while cutting the cost of a decision from minutes to seconds.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑