基准并非单一整体:面向大语言模型评估的样本级审计与编排
Benchmarks Are Not Monolithic: Sample-Level Auditing and Orchestration for LLM Evaluation
AI总结:
该研究提出样本级审计的元评估框架,标注五大基准的内部异质性,可编排复合基准子集实现模型能力针对性评估,为基准重组提供原则性方法。
AI中文摘要:
基准数据集是评估大语言模型(LLMs)的核心,但通常被视为单一整体任务,掩盖了单个样本需求的显著差异。我们提出一种以数据集为中心的元评估框架,沿五个潜在维度对基准数据集进行样本级审计:1. 认知与知识需求;2. 语言与内容质量;3. 任务属性;4. 上下文;5. 伦理、安全与公平性。应用该框架,我们对MMLU、ARC、WinoGrande、HellaSwag和TruthfulQA这五个有影响力的基准进行标注,揭示了聚合准确率分数无法捕捉的显著内部异质性。我们展示这些标注如何支持跨数据集的复合基准子集的标准驱动编排,以实现对推理深度或伦理敏感性等模型能力的针对性评估。该方法将基准评估重新定义为数据集内省,为分析和重组现有基准以更好反映多样化评估需求提供了原则性方法。
英文摘要:
Benchmark datasets are central to evaluating Large Language Models (LLMs), yet they are typically conceived as monolithic tasks, obscuring substantial variation in the demands of individual samples. We introduce a dataset-centric meta-evaluation framework that audits benchmark datasets at the sample level along five latent dimensions: 1. Cognitive and Knowledge Demands, 2. Language and Content Quality, 3. Task Properties, 4. Context, and 5. Ethics, Safety, and Fairness. Applying this framework, we annotate five influential benchmarks -- MMLU, ARC, WinoGrande, HellaSwag, and TruthfulQA -- revealing pronounced internal heterogeneity that is not captured by aggregate accuracy scores. We show how these annotations enable criterion-driven orchestration of composite benchmark subsets across datasets, supporting targeted evaluation of model capabilities such as Reasoning Depth or Ethical Sensitivity. This approach reframes benchmark evaluation as dataset introspection, providing a principled methodology for analyzing and re-composing existing benchmarks to better reflect diverse evaluation needs.