arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2506.20911cs.CV

FaSTA$^*$:用于高效多轮图像编辑的带子程序挖掘的快慢工具路径智能体

FaSTA$^*$: Fast-Slow Toolpath Agent with Subroutine Mining for Efficient Multi-turn Image Editing

  • University of Maryland, College Park(马里兰大学 College Park分校)

机构由 AI 辅助整理,请以论文原文为准。

Advait Gupta, Rishie Raj, Dang Nguyen, Tianyi Zhou

更新

AI总结:

本文提出了一种名为FaSTA$^*$的神经符号智能体,结合LLM快速规划与局部A*搜索,通过归纳推理挖掘并复用子程序,在保持成功率的同时显著提升了多轮图像编辑的计算效率。

AI中文摘要:

我们开发了一种成本高效的神经符号智能体,以处理具有挑战性的多轮图像编辑任务,例如“检测图像中的长椅并将其重新着色为粉色。同时,移除猫以获得更清晰的视图,并将墙壁重新着色为黄色。”它结合了大型语言模型(LLMs)快速、高层的子任务规划,以及每个子任务缓慢、准确、工具使用和局部 A$^*$ 搜索,以寻找成本高效的工具路径——即对 AI 工具的调用序列。为了节省 A$^*$ 在类似子任务上的成本,我们通过 LLMs 对先前成功的工具路径进行归纳推理,以持续提取/优化常用的子程序,并将其作为新工具重用于未来任务的自适应快慢规划中,其中首先探索高层子程序,只有当它们失败时,才激活低层 A$^*$ 搜索。可重用的符号子程序大幅节省了在应用于类似图像的同类子任务上的探索成本,从而产生了一种类人的快慢工具路径智能体“FaSTA$^*$”:首先由 LLMs 尝试快速子任务规划,随后进行基于规则的子程序选择,这预计将覆盖大部分任务,而缓慢的 A$^*$ 搜索仅针对新颖且具有挑战性的子任务触发。通过与近期的图像编辑方法进行比较,我们证明 FaSTA$^*$ 在计算效率上显著更高,同时在成功率方面仍与最先进的基线保持竞争力。

英文摘要:

We develop a cost-efficient neurosymbolic agent to address challenging multi-turn image editing tasks such as "Detect the bench in the image while recoloring it to pink. Also, remove the cat for a clearer view and recolor the wall to yellow.'' It combines the fast, high-level subtask planning by large language models (LLMs) with the slow, accurate, tool-use, and local A$^*$ search per subtask to find a cost-efficient toolpath -- a sequence of calls to AI tools. To save the cost of A$^*$ on similar subtasks, we perform inductive reasoning on previously successful toolpaths via LLMs to continuously extract/refine frequently used subroutines and reuse them as new tools for future tasks in an adaptive fast-slow planning, where the higher-level subroutines are explored first, and only when they fail, the low-level A$^*$ search is activated. The reusable symbolic subroutines considerably save exploration cost on the same types of subtasks applied to similar images, yielding a human-like fast-slow toolpath agent "FaSTA$^*$'': fast subtask planning followed by rule-based subroutine selection per subtask is attempted by LLMs at first, which is expected to cover most tasks, while slow A$^*$ search is only triggered for novel and challenging subtasks. By comparing with recent image editing approaches, we demonstrate FaSTA$^*$ is significantly more computationally efficient while remaining competitive with the state-of-the-art baseline in terms of success rate.

↑