AI 中文总结
该研究提出名为 Loreley 的 QD 方法,将完整仓库状态存入存档采样,在 Zstandard 实验中对比三种策略,发现 48 作业时 QD 未显优势但垫脚石机制生效,早期活动已产出多文件改进。
AI 中文摘要
顺序智能体搜索会从当前最优解累积变更,但会丢弃替代分支;独立提案则保留了广度,但会从根节点重新开始。Loreley 则相反,它在质量多样性(Quality-Diversity, QD)存档中保留完整的仓库状态,并将这些状态作为父节点采样,或作为后续编辑的上下文提供。候选解是在独立工作树(worktree)中生成的 Git 提交,并由项目提供的评估器进行评判。我们在匹配的 Zstandard 实验中比较了配置后的 Loreley QD、顺序最优解编辑和独立根节点提案:每个策略和每个区块有七个配对区块及 48 个物理候选作业(总计 1008 个),采用仅根节点初始化和各策略的原生并发机制。验证阶段在每个预算检查点选择一个优胜者;智能体隐藏的留存集用于测量固定候选解。在 48 个作业时,QD 比顺序最优解低 0.135%(配对效应的 95% BCa 区间:-0.556% 至 +0.161%),比独立根节点高 0.320%(区间:-0.082% 至 +0.686%)。两次对比均未确立 QD 的优势;顺序策略的观测 48 作业均值和中位数最高。存档保留及后续采样确实发生了。在仅对观测到的 QD 流应用回顾性单最优解规则下,七个最终 QD 优胜者中有四个在其主要父系祖先中具有非现任状态;加入灵感边后数量升至六个,但未显示提供的上下文导致了编辑。三次早期能力活动在两个 Python 库及一个独立 Zstandard 修订中产生了第 4 代多文件改进。Loreley 启用了预期的垫脚石机制,但受控实验在 48 个作业时未显示出终点效益。
英文摘要
Sequential agent search accumulates changes from its current champion but discards alternative branches; independent proposals preserve breadth but restart from the root. Loreley instead retains complete repository states in a Quality-Diversity (QD) archive and samples them as parents or supplies them as context for later edits. Candidates are Git commits produced in isolated worktrees and judged by a project-supplied evaluator. We compare configured Loreley QD, sequential champion editing, and independent root proposals in a matched Zstandard experiment: seven paired blocks and 48 physical candidate jobs per policy and block (1,008 total), with root-only initialization and each policy's native concurrency. Validation selected a winner at each budget checkpoint; an agent-hidden holdout measured the fixed candidate. At 48 jobs, QD was 0.135% below Sequential Champion (95% BCa interval for the paired effect: -0.556% to +0.161%) and 0.320% above Independent Root (-0.082% to +0.686%). Neither contrast established a QD advantage; Sequential had the highest observed 48-job mean and median. Archive retention and later sampling did occur. Four of seven final QD winners had a non-incumbent state in their primary-parent ancestry under a retrospective one-incumbent rule applied only to the observed QD stream. Including inspiration edges raised the count to six, without showing that supplied context caused an edit. Three earlier capability campaigns produced generation-4, multi-file improvements in two Python libraries and a separate Zstandard revision. Loreley engaged the intended stepping-stone mechanism, but the controlled experiment did not show an endpoint benefit at 48 jobs.
Comments15 pages, 3 figures, 8 tables. Code and evidence: https://github.com/NeapolitanIcecream/loreley