arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

MoMHa:面向准确性、安全性与Token的LLM Harness多目标优化

MoMHa: Multi-Objective Optimization of LLM Harnesses over Accuracy, Safety, and Tokens

Subhojyoti Mukherjee, Md Mehrab Tanjim

arXiv 2609.30967首次发表:更新:

发表机构

Adobe Research(Adobe研究院)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

MoMHa将LLM harness设计视为多目标优化,通过智能体搜索平衡准确性、安全性与Token成本,在17个领域上超越基线,并实现最高安全得分与更低Token消耗。

AI 中文摘要

大多数关于改进大型语言模型的研究都将准确性视为唯一目标。我们认为,harness——即围绕模型构建提示、路由调用并解析输出的Python代码——是一等设计面,其质量本质上是多目标的:一个准确但拒绝任何不安全请求的harness,或一个消耗多一个数量级Token的harness,都不是好的harness。我们提出了Meta-Harness,一个将harness设计视为对三个领域目标(准确性、行为安全性和Token成本)的搜索的系统,由具有对先前harness源代码、执行轨迹和评分工件完全文件系统访问权限的智能体提议器(Claude Code)解决。我们的核心发现是,单阶段联合奖励提议器(MoMHa)优于所有替代方案,包括两阶段“先准确性后Token”的消融、仅标量反馈以及仅准确性的基线。我们在十七个领域上进行了评估:七个合成能力套件、七个真实世界公共基准(HumanEval、MBPP、Spider、FEVER、MMLU-Pro、LawBench、NuminaMath)以及三个源自U-SafeBench的用户特定安全领域,使用了涵盖四个系列的12模型舰队。在合成轨道上,MoMHa实现了0.482的联合均值,而十个基线为0.198-0.422,赢得了7/10的领域列;在真实世界轨道上,它得分为0.461,而最强基线(DSPy)为0.377,赢得了5/7列,证明了harness策略无需在12个目标模型中的8个上重新训练即可迁移到未见基准。MoMHa达到了最高的实测行为安全综合得分(U-SafeBench,0.781),并且每个示例比两阶段替代方案少使用95个Token。我们将发布所有harness代码、评估基础设施和跨模型日志。

英文摘要

Most work on improving large language models treats accuracy as the sole objective. We argue that the harness, the Python code surrounding the model that constructs prompts, routes calls, and parses outputs, is a first-class design surface whose quality is inherently multi-objective: an accurate harness that refuses no unsafe request, or that consumes an order of magnitude more tokens, is not a good harness. We present Meta-Harness, a system that casts harness design as search over three per-domain objectives (accuracy, behavioural safety, and token cost) solved by an agentic proposer (Claude Code) with full filesystem access to prior harness source, execution traces, and scoring artifacts. Our central finding is that a singlephase joint-reward proposer (MoMHa) outperforms every alternative, including a two-phase "accuracy then tokens" ablation, scalar-only feedback, and an accuracy-only baseline. We evaluate on seventeen domains: seven synthetic capability suites, seven real-world public benchmarks (HumanEval, MBPP, Spider, FEVER, MMLU-Pro, LawBench, NuminaMath), and three U-SafeBench-derived user-specific safety domains, using a 12-model fleet spanning four families. On the synthetic track MoMHa achieves a joint mean of 0.482 versus 0.198-0.422 for ten baselines, winning $7 / 10$ per-domain columns; on the real-world track it scores 0.461 versus 0.377 for the strongest baseline (DSPy), winning 5/7 columns, demonstrating that harness strategies transfer to unseen benchmarks without retraining on 8 of 12 target models. MoMHa attains the highest measured behavioral safety composite (U-SafeBench, 0.781) and uses 95 fewer tokens per example than the two-phase alternative. We will release all harness code, evaluation infrastructure, and crossmodel logs.

CommentsAccepted to the 40th Conference on Neural Information Processing Systems (NeurIPS 2026)

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑