arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.13725cs.AIcs.LO

IBBench-Light:对外部指令的任务条件响应的配对评估

IBBench-Light: A Paired Evaluation of Task-Conditioned Responses to External Directives

Kainan Zhou, Zhaoyi Li, Janet Sung, Gangzhen Qian, Hang Xiao

首次发表
浏览论文内容

中文总结 AI 辅助

IBBench-Light通过配对评估外部指令下的任务条件响应,提出配对精确契约准确率(PECA)指标,并揭示停止策略对模型性能的关键影响。

中文摘要 AI 辅助

外部记录可能包含根据用户请求需要应用的程序或需要阅读的文本。IBBench-Light针对同一记录测试这两种用途。十二个语义基础为每个模型产生144个匹配对;四个量化指令模型产生了1,152个存档的贪婪响应。配对精确契约准确率(PECA)要求两个成员都满足其输出契约。Qwen在132个执行提示和109个处理提示上成功,但只有97个完整配对,这显示了边际平均值所遗漏的信息。我们审计了字面目标暴露和大小写规范化,然后添加了1,722次记录的CPU生成,以测试无指令控制、额外的十二个语义基础、基础内措辞变化以及生成停止。在固定的Phi重运行中,改变结束序列(EOS)集合将精确配对成功从0/144变为62/144。一个有界的IHEval比较使用相同的SmolLM2检查点和输出预算,同时保留其已发布的指令角色和评分器。该基准衡量条件任务和输出契约的成功。其任务边际和配对数量需要与停止策略一起阅读。

英文摘要

An external record may contain a procedure to apply or text to read, depending on the user's request. IBBench-Light tests both uses against the same record. Twelve semantic bases yield 144 matched pairs per model; four quantized instruction models produced 1,152 archived greedy responses. Paired exact-contract accuracy (PECA) requires both members to satisfy their output contracts. Qwen succeeds on 132 execute and 109 process prompts, but only 97 complete pairs, showing what marginal averages omit. We audit literal-target exposure and case normalization, then add 1,722 logged CPU generations to test directive-absent controls, twelve additional semantic bases, within-base wording changes, and generation stopping. In the pinned Phi rerun, changing the end-of-sequence (EOS) set changes exact paired success from 0/144 to 62/144. A bounded IHEval comparison uses the same SmolLM2 checkpoint and output budget while preserving its published instruction roles and scorer. The benchmark measures conditional task and output-contract success. Its task margins and paired count need to be read together with the stopping policy.

发表机构

  • Google LLC(谷歌有限责任公司)
  • Intuit Inc.(Intuit公司)
  • Fortinet, Inc.(飞塔公司)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑