发表机构
Scale AI(Scale人工智能公司)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
研究针对评估应指导模型改进及提供训练数据的问题,提出CRAFT方法,将评分标准评估数据集转化为模型能力短板诊断,通过聚类能力描述、评估模型节点等生成微调数据,实验表明该方法能更精准诊断模型弱点并提升性能。
AI 中文摘要
评估不应仅衡量模型当前性能,还应指出改进方向并提供针对性训练数据。多数评估流程仅指出模型失败之处,未说明原因。本文介绍CRAFT方法,将基于评分标准的评估数据集转化为对模型能力短板的特定诊断。该方法把每个评分标准视为能力探针,提取能力描述并聚类成层次化能力树,在各节点评估目标模型,动态选择低性能节点以指导生成针对性的监督微调数据。在四个开源模型、两个专业领域及13个与诊断数据不相交的基准测试上的比较结果显示,CRAFT在重复温度解码下,在金融领域对所有四个模型平均表现最佳;在法律领域,四个模型中有三个表现最强,第四个模型也在最佳基线的解码方差范围内。在评分标准层面诊断弱点,能更清晰地了解模型的不足,并在基于此诊断进行微调后得到性能更佳的模型。
英文摘要
Evaluations should do more than measure a models current performance. They should tell us what to fix for the next model iteration and provide a way to generate targeted post training data. Most evaluation pipelines identify weak examples, topics, or categories, but they leave the underlying capability failure implicit: they say where a model fails, not why. We introduce CRAFT, a method that converts any rubric based evaluation dataset into a model specific diagnosis of weak capabilities. CRAFT treats each grading criterion as a capability probe: it extracts a capability description from every prompt rubric pair, clusters these descriptions into a hierarchical capability tree, scores the target model at every node, and selects low performing nodes dynamically across tree levels, at the granularity where each failure is clearest. The selected weak capabilities then direct the generation of targeted supervised finetuning data. Holding the data generation, finetuning, and evaluation setup fixed, we compare CRAFT against prompt level EvalTree clustering and untargeted random generation on four open source models, two professional domains (finance and legal), and 13 held out benchmarks disjoint from the diagnostic data. CRAFT achieves the strongest finance domain average for all four models under repeated temperature decoding; on legal domain, it is strongest for three of four models and remains within the decoding variance bands of the best baseline on the fourth. Diagnosing weaknesses at the level of rubric criteria, rather than prompts or categories, thus yields both a sharper picture of what a model cannot do and measurably better models after finetuning on that diagnosis.
Comments18 pages, 3 Tables, 2 Figures