arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

AutoFyn 技术报告:面向长时程智能体的非参数专家迭代

AutoFyn Technical Report: Non-Parametric Expert Iteration for Long-Horizon Agents

Adib Hasan, Daniel Schaffield, Akashnil Dutta, Tarik Adnan Moon

arXiv 2609.05446首次发表:更新:

发表机构

SignalPilot Labs; Prentis AI(SignalPilot实验室; Prentis AI)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

AutoFyn通过持久状态更新而非权重调整实现专家迭代,在数学、数据科学和网络安全领域显著提升性能,并产出多个漏洞公告。

AI 中文摘要

我们引入了 AutoFyn,一个受专家迭代算法启发的智能体框架,它通过从已验证的奖励信号中更新持久状态而非模型权重,使冻结模型在多个回合中适应。每个回合从全新的模型会话开始,持久信息仅通过显式接口(如持久记忆文件、报告和仓库状态)重新引入。在一个回合内,编排器探索、规划并构建多种替代方法,使用专门的智能体,而任务导向的验证器验证工作并提供客观奖励以衡量进展。该奖励被蒸馏回持久状态,从而更新下一回合的有效策略。在本技术报告中,我们形式化了这一循环,并描述了其持久状态和验证接口。然后,我们展示了其在三个领域的应用,即奥林匹克数学、数据科学和网络安全。在2026年国际数学奥林匹克竞赛的六个新问题上,每个有改进空间的模型在AutoFyn下的得分都高于其提供商自身的编码智能体。AutoFyn还构建了Spider 2.0 dbt基准上排名第一的智能体,并在http this URL、MetaMask、pnpm、Warp、LiteLLM、Langflow和Open WebUI中产生了16个维护者确认的漏洞公告。

英文摘要

We introduce AutoFyn, an agent harness inspired by the Expert Iteration algorithm, adapting a frozen model across many rounds by updating persistent state from verified reward signals rather than model weights. Each round begins from a fresh model session, and durable information is reintroduced only through explicit interfaces such as persistent memory files, reports, and repository state. Within a round, an orchestrator explores, plans and builds many alternative approaches with specialized agents, while a task-grounded verifier verifies the work and supplies an objective reward for measuring progress. This reward is distilled back into the persistent state, which updates the effective policy for the next round. In this technical report, we formalize this loop and describe its persistent state and verification interfaces. We then demonstrate its use in three domains, namely olympiad mathematics, data science, and cybersecurity. On the six fresh problems of the 2026 International Mathematical Olympiad, every model with room to improve scores higher under AutoFyn than in its provider's own coding agent. AutoFyn also built the top-ranked agent on the Spider 2.0 dbt benchmark, and has produced $16$ maintainer-confirmed vulnerability advisories in Next.js, MetaMask, pnpm, Warp, LiteLLM, Langflow, and Open WebUI.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑