发表机构
Brunel University of London(伦敦布鲁内尔大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
CompOrca利用五重LLM评判器对OpenOrca全语料进行合规标注,区分一致合规与不合规,实现高精度过滤并揭示现有拒绝检测方法的召回局限。
AI 中文摘要
研究微调如何塑造拒绝和不服从行为,需要识别那些拒绝、回避或以其他方式未能完成所请求任务的训练样本。但现有的标注最多覆盖几千个提示词的评估集。我们提出了CompOrca,一个对包含4,233,923个样本的完整OpenOrca语料库进行的合规性标注。每个样本都由一个开放权重的大语言模型评判器(LongCat-2.0,1.6万亿参数)的五次独立评估分类为合规或不合规,该语料库以一致合规(94.75%)、一致不合规(1.28%)和非一致行(3.97%)以及原始投票计数发布。单次评估将语料库的2.7-3.2%标记为不合规,而只有1.28%被全部五次标记,从而可以过滤最模糊的样本。针对450个人工标注的样本,其中150个被标注两次(人-人间κ=0.93),一致合规和一致不合规标签的精确率分别为97.3%和86.7%,后者是高度精确的子集,而非不合规的完整枚举。已发表的拒绝检测方法仅能召回不合规类别的0.4%至94.1%。我们在以下https URL发布包含每行标签和投票计数的完整语料库。
英文摘要
Studying how fine-tuning shapes refusal and noncompliance behaviour requires knowing which training examples refuse or otherwise fail to fulfil the request. Existing annotations cover evaluation sets, which are far smaller than training corpora. We present CompOrca, compliance labels for all 4,233,923 examples of the OpenOrca corpus. Every example was classified as compliant or noncompliant by five passes of an open-weight LLM judge (LongCat-2.0, 1.6T parameters). The corpus is released as unanimous compliance (94.75%), unanimous noncompliance (1.28%), and nonunanimous rows (3.97%), with the raw vote counts. A single pass flags 2.7-3.2% of the corpus as noncompliant, while only 1.28% is flagged by all five, so the most ambiguous rows can be filtered out. Against 450 human-annotated examples (150 annotated twice; human-human $κ=0.93$), the unanimous compliance and noncompliance labels are 97.3% and 86.7% precise. The noncompliance label is a high-precision subset of the corpus's noncompliance. Published refusal-detection methods recall between 0.4% and 94.1% of human-labelled noncompliance. We release the full corpus with its per-row labels and vote counts at https://huggingface.co/datasets/cemiu/CompOrca
CommentsAccepted to PlurVA-LLM Workshop @ AACL-IJCNLP 2026. Dataset available on HuggingFace