arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2610.07362cs.LG

硬资源约束下的大语言模型评估动态预算分配

Dynamic Budget Allocation for LLM Evaluation under Hard Resource Constraints

Shai Feldman, Yaniv Romano

首次发表
浏览论文内容

中文总结 AI 辅助

针对硬资源约束下LLM评估的预算分配问题,提出HARP方法,通过自适应回流分配满足硬预算并构建下预测界,保证不超预算、覆盖率和无偏估计,实验验证其有效性。

中文摘要 AI 辅助

我们通过事件发生时间(time-to-event)来评估大语言模型(LLMs)在多轮交互中的表现:即产生感兴趣事件(如成功越狱或智能体任务完成)所需的交互步数。在计算资源有限的情况下,交互可能在事件发生前被终止,因此事件时间只能被部分观测(即被删失)。现有的用于校准事件时间界限的分配方法仅在期望意义上满足预算,并且可能在特定评估运行中超出可用预算。强制执行硬约束尤其具有挑战性,因为轨迹的成本最初是未知的。我们引入了硬预算分配与回流预测校准方法(HARP),这是一种满足硬资源约束并自适应地重新分配未使用预算的预算分配方法。我们展示了如何使用HARP来构建事件发生时间的下预测界(LPBs),并估计固定基准上的评估指标,如越狱率。尽管HARP在不同轨迹的采集决策中引入了依赖性,但我们证明了HARP永远不会超出目标预算,其LPBs具有有限样本覆盖保证,并且其指标估计是无偏的。在智能体任务成功、LLM越狱、有毒内容生成和RAG幻觉等实验表明,HARP在低方差下实现了接近名义水平的覆盖率,同时从不超出给定预算。

英文摘要

We evaluate large language models (LLMs) in multi-turn interactions through their time-to-event: the number of interaction steps required to produce an event of interest, such as a successful jailbreak or agentic task completion. Under limited compute, interactions may be terminated before the event occurs, so that event times are only partially observed (censored). Existing allocation methods for calibrating time-to-event bounds satisfy the budget only in expectation and can exceed the available budget on a particular evaluation run. Enforcing a hard constraint is particularly challenging as the cost of a trajectory is initially unknown. We introduce Hard-budget Allocation with Reflow for Predictive calibration (HARP), a budget allocation that satisfies hard resource constraints and adaptively reallocates unused budget. We show how to use HARP to construct lower predictive bounds (LPBs) on the time-to-event and to estimate evaluation metrics such as the jailbreak rate on a fixed benchmark. Although HARP induces dependence in acquisition decisions across different trajectories, we prove that HARP never exceeds the target budget, that its LPBs have finite-sample coverage guarantees, and that its metric estimates are unbiased. Experiments on agentic task success, LLM jailbreaks, toxic content generation, and RAG hallucinations show that HARP achieves coverage close to the nominal level with low variance, while never exceeding the given budget.

发表机构

  • Technion IIT(以色列理工学院)

机构由 AI 辅助整理,请以论文原文为准。

↑