发表机构
RIKEN Center for Computational Science(理化学研究所计算科学中心)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
该研究针对多候选LLM生成的评估场景,构建了五领域基准VTC-Bench并提出核心指标VTC,发现传统评估无法识别的模型行为差异,为多候选生成评估提供了新方案。
AI 中文摘要
许多大语言模型(LLM)应用在提供多个候选输出供比较、验证或组合时最具实用价值。然而,主流评估设置仍聚焦于单个输出,或将多个样本简化为单一成功或选定答案,这可能会遗漏输出是否包含若干真正不同的有用结果。我们引入了VTC-Bench——针对该评估场景的五领域基准,以及其核心评估指标验证的任务覆盖度(Validated Task Coverage, VTC)。该基准由精心选取的真实数据任务构建,可自动且可复现地检查输出质量与任务相关的独特性,无需基于模型的评判器。VTC衡量在k次尝试中获得的不同有用结果的数量。在多个模型和推理设置下,该基准得出了与传统评估不同的结论:从单次抽取质量来看表现最佳的配置,未必具有最佳覆盖度;而简单的输出变化度量也无法可靠地恢复与任务相关的覆盖度。这些结果表明,有限候选集可直接作为评估对象,揭示出传统逐输出评估无法发现的模型行为差异。
英文摘要
Many LLM applications are most useful when they provide several candidate outputs for comparison, validation, or combination. Predominant evaluation settings, however, still focus on individual outputs or reduce multiple samples to a single success or selected answer. This can miss whether the outputs include several genuinely different useful results. We introduce VTC-Bench, a five-domain benchmark for this setting, together with Validated Task Coverage (VTC) as its core evaluation quantity. The benchmark is built from carefully selected real-data tasks where both output quality and task-relevant distinctness can be checked automatically and reproducibly, without model-based judges. VTC measures how many distinct useful results are obtained within $k$ attempts. Across multiple models and inference settings, the benchmark leads to different conclusions from conventional evaluation: configurations that look strongest from single-draw quality are not necessarily those with the best coverage, and simple measures of output variation do not reliably recover task-relevant coverage. These results show that finite candidate sets can be evaluated directly as objects of interest, revealing differences in model behavior that are not apparent from conventional per-output evaluation.