arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

持续改进与并行自主探索:用于搜索大型解空间的大语言模型智能体框架

Continuous Improvement and Parallel Autonomous Exploration: An LLM-Agent Framework for Searching Large Solution Spaces

Dulmini Hettiarachchi, Andre Rusli, Julio Christian Young, Sho Akiyama

arXiv 2608.04341首次发表:更新:

AI 中文总结

该研究提出一种LLM智能体框架,含持续改进奖励循环与并行自主探索机制,在产品到目录匹配任务中,5个并行智能体的合格覆盖度较单智能体及基线显著提升。

AI 中文摘要

我们提出了一个为大语言模型(LLM)智能体提供两种自主搜索大型解空间机制的框架。首先,基于保留数据评分的排行榜充当奖励信号,驱动每个智能体通过重复提交来优化其解决方案,该循环即使在单个智能体的情况下也能运行。其次,该框架支持完全自主地并行运行多个智能体,无需人工参与:智能体独立分析、调研方法、实现、自我评估、提交并修订,而仅由一个协调智能体处理后勤事务。在共享奖励下并行运行智能体,可拓宽解空间的探索范围,而非仅优化单一初始范式。我们将该框架实例化到产品到目录匹配任务中(这是一项核心电商检索任务,具有大型、类别结构化的解空间),将其表述为具有精确率-覆盖度操作点的选择性预测问题。单个智能体在其初始范式内进行优化,而并行自主智能体则会呈现出性质不同的解决方案。在该测试平台上,最佳合格覆盖度(每个类别P@1≥95%)在单个智能体下达到47.8%-57.4%,在5个智能体下达到62.8%-69.4%,而基线为33.3%。我们的贡献是该框架本身:一个持续改进的奖励循环和用于完全自主并行探索的基础,并有案例研究证据支持。

英文摘要

We present a framework that gives LLM agents two mechanisms for searching large solution spaces autonomously. First, a leaderboard scored on held-out data acts as a reward signal that drives each agent to refine its solutions over repeated submissions, a loop that operates even with a single agent. Second, the framework enables running many agents in parallel, fully autonomously, with no human in the loop: agents independently analyze, survey methods, implement, self-evaluate, submit, and revise, while a moderator agent handles only logistics. Running agents in parallel under the shared reward broadens the explored region of the solution space rather than refining the single seeded paradigm. We instantiate the framework on product-to-catalog matching (a core e-commerce retrieval task with a large, category-structured solution space), posed as selective prediction with a precision-coverage operating point. A single agent refines within its seeded paradigm, whereas parallel autonomous agents surface qualitatively different solutions. On this testbed, best qualified coverage (>=95% P@1 per category) reaches 47.8-57.4% with a single agent and 62.8-69.4% with five, against a 33.3% baseline. Our contribution is the framework itself: a continuous-improvement reward loop and a substrate for fully autonomous parallel exploration, backed by case-study evidence.

Comments5 pages, ACM KDD'26 Workshop on SciSoc Agents & LLMs

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑