arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2603.28052cs.AI

Meta-Harness:端到端优化模型Harness

Meta-Harness: End-to-End Optimization of Model Harnesses

  • Stanford(斯坦福大学)
  • KRAFTON
  • MIT(麻省理工学院)

机构由 AI 辅助整理,请以论文原文为准。

Yoonho Lee, Roshen Nair, Qizheng Zhang, Kangwook Lee, Omar Khattab, Chelsea Finn

更新

AI总结:

本文提出Meta-Harness,通过自动搜索优化模型Harness代码,提升文本分类、数学推理和编程任务性能,减少上下文token使用量。

AI中文摘要:

大型语言模型(LLM)系统的性能不仅取决于模型权重,还取决于其Harness:决定存储、检索和呈现信息的代码。然而,Harness仍主要由人工设计,现有文本优化器因压缩反馈过于激进而不适用于此场景。我们引入Meta-Harness,一个外循环系统,用于搜索LLM应用中的Harness代码。它使用一个代理提案者,通过文件系统访问源代码,评分和执行痕迹的所有先前候选方案。在在线文本分类中,Meta-Harness在7.7分上优于最先进的上下文管理系统,同时使用4倍更少的上下文token。在检索增强的数学推理中,单个发现的Harness在200个IMO级问题上,平均提升五个测试模型的准确性4.7分。在代理编程中,发现的Harness在TerminalBench-2上超越了最佳的手工工程基线。这些结果表明,更丰富的先前经验访问可以实现自动Harness工程。

英文摘要:

The performance of large language model (LLM) systems depends not only on model weights, but also on their harness: the code that determines what information to store, retrieve, and present to the model. Yet harnesses are still designed largely by hand, and existing text optimizers are poorly matched to this setting because they compress feedback too aggressively. We introduce Meta-Harness, an outer-loop system that searches over harness code for LLM applications. It uses an agentic proposer that accesses the source code, scores, and execution traces of all prior candidates through a filesystem. On online text classification, Meta-Harness improves over a state-of-the-art context management system by 7.7 points while using 4x fewer context tokens. On retrieval-augmented math reasoning, a single discovered harness improves accuracy on 200 IMO-level problems by 4.7 points on average across five held-out models. On agentic coding, discovered harnesses surpass the best hand-engineered baselines on TerminalBench-2. Together, these results show that richer access to prior experience can enable automated harness engineering.

↑