通过专家指导生成有效情境特定基准的框架
A Framework for Generating Valid Context-Specific Benchmarks through Expert Guidance
浏览论文内容
中文总结 AI 辅助
本文提出一个通过专家输入引导合成数据生成的框架,以生成情境特定LLM基准,通过四个质量标准提升数据有效性,并经定量评估和案例研究验证其优于现有方法。
中文摘要 AI 辅助
本文提出了一种端到端的方法,通过结合专家输入与合成数据生成,来生成情境特定的大语言模型(LLM)基准数据集。现有的基准构建方法往往在有效性和可扩展性之间进行权衡:由领域专家设计的数据集可以产生高质量的评价,但创建过程缓慢且成本高昂,而合成生成数据虽可高效扩展,但常常导致不真实、冗余或超出范围的示例。为解决这一差距,我们引入了一个模式(schema),用于提取评价任务的目标、范围和情境等关键信息,并利用这些信息指导合成数据的生成。我们进一步定义了基于测量有效性(measurement validity)的四个标准来评估数据集质量:覆盖率、多样性、内容真实性和风格真实性。利用这些标准,我们展示了专家知情的支架(scaffolds)如何引导合成数据生成,使其更接近有效的基准。通过定量评价和与领域专家的真实世界案例研究,我们证明我们的方法在保持有效性的同时,相较于现有方法提高了基准数据的质量。我们还分析了不同类型的模式信息如何影响不同的数据集质量标准,并提供了在资源受限条件下应优先收集哪些信息的实用指导。
英文摘要
This paper presents an end-to-end approach for generating context-specific large language model (LLM) benchmark datasets by combining expert input with synthetic data generation. Existing benchmark construction methods often trade off validity and scalability: datasets designed with domain experts can produce high-quality evaluations but are slow and costly to create, while synthetically generating data may scale efficiently but often results in unrealistic, redundant, or out-of-scope examples. To address this gap, we introduce a schema eliciting key information about the goals, scope, and context of an evaluation task, and use this information to guide synthetic data generation. We further define four criteria grounded in measurement validity for assessing dataset quality: coverage, diversity, content realism, and stylistic realism. Using these criteria, we show how expert-informed scaffolds can guide synthetic data generation toward more valid benchmarks. Through quantitative evaluations and a real-world case study with domain experts, we demonstrate that our approach improves benchmark data quality over existing methods while preserving validity. We additionally analyze how different types of schema information affect different dataset quality criteria, and provide practical guidance on which information to prioritize collecting under resource constraints.
发表机构
- Carnegie Mellon University(卡内基梅隆大学)
机构由 AI 辅助整理,请以论文原文为准。