arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

BudgetBench:面向本地大语言模型智能体记忆策略评估的预算分层协议与试点测试框架

BudgetBench: A Budget-Tiered Protocol and Pilot Harness for Memory Strategy Evaluation in Local Large Language Model Agents

Aditya Karnam Gururaj Rao, Arjun Jaggi

arXiv 2609.13149首次发表:更新:

AI 中文总结

BudgetBench提出预算分层协议与测试框架,通过扫描多档令牌预算评估本地大语言模型智能体记忆策略,揭示预算违规与质量非单调性,为固定预算评估提供可复用测量表面。

AI 中文摘要

对于本地大语言模型智能体而言,活动上下文是一种稀缺资源:内存容量、预填充延迟、缓存增长和服务目标都限制了每次调用可负担的输入令牌数量。我们提出了BudgetBench,一种活动预算协议和参考测试框架,将每次调用的输入令牌预算作为比较记忆策略时的自变量。在保持模型、任务、采样器和解码固定的情况下,它扫描2K、4K、8K、16K和32K令牌的预算,并记录质量、预算利用率、延迟以及作为首要结果的预算违规率。核心贡献在于这一可复用的测量表面:可替换的MemoryStrategy契约、显式的预算执行、确定性或版本化的评分器、提示审计元数据和可复现性工件,发布于此https URL。我们通过试点研究而非最终排名来验证该协议。在本地qwen2.5:1.5b试点(SWE-bench Verified和LongBench v2各89项)、使用精确分词的主机托管50项Qwen3 30B-A3B LongBench复现实验,以及由官方GPT-4o评估器评分的500项LongMemEval预言研究上,该框架暴露了预算合规失败、非单调质量曲线以及单预算评估所隐藏的工作点。预算化与全上下文方向仍未解决:本地切片接近无效,主机托管复现的点估计偏向全上下文。我们透明地报告结果,包括早期试点中的分词器近似低估了某些服务模型提示,因此其违规行是分词器近似的诊断,而非承载结论的结果;所有计时均为操作诊断。可复用的贡献在于协议、测试框架和失败报告纪律,这些是扩展固定预算记忆策略评估所必需的。

英文摘要

For local large language model agents, active context is a scarce resource: memory capacity, prefill latency, cache growth, and service objectives all constrain how many input tokens each call can afford. We present BudgetBench, an active-budget protocol and reference harness that treats the per-call input-token budget as the independent variable when comparing memory strategies. Holding the model, task, sampler, and decoding fixed, it sweeps budgets over 2K, 4K, 8K, 16K, and 32K tokens and records quality, budget utilization, latency, and, as a first-class outcome, budget-violation rates. The core contribution is this reusable measurement surface: a swappable MemoryStrategy contract, explicit budget enforcement, deterministic or versioned graders, prompt-audit metadata, and reproducibility artifacts, released at https://github.com/aviskaar/budgetbench. We substantiate the protocol with pilot studies rather than final rankings. Across a local qwen2.5:1.5b pilot (89 items each on SWE-bench Verified and LongBench v2), a hosted 50-item Qwen3 30B-A3B LongBench replication with exact tokenization, and a 500-item LongMemEval oracle study scored by the official GPT-4o evaluator, the harness exposes budget-compliance failures, non-monotonic quality curves, and operating points that single-budget evaluation hides. The budgeted-versus-full-context direction remains unresolved: the local slice is near-null and the hosted replication favors full context in point estimate. We report results transparently, including that the early pilot's tokenizer approximation undercounts some served-model prompts, so its violation rows are tokenizer-approximation diagnostics, not claim-bearing results; all timings are operational diagnostics. The reusable contribution is the protocol, harness, and failure-reporting discipline needed to scale fixed-budget memory-strategy evaluation.

Comments44 pages, 4 figures. Code and artifacts: https://github.com/aviskaar/budgetbench

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑