发表机构
Microsoft(微软)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
该研究发布了GLAN-QnA-KR韩语指令语料库,通过无种子分类法驱动的GLAN合成管道生成。此语料库规模大且具有独特属性,如重复问题少、污染率低,是Hugging Face Hub上可验证的最大单管道合成韩语指令语料库及唯一按该协议构建的韩语 >= 100k行语料库。
AI 中文摘要
我们发布了GLAN-QnA-KR,这是一个拥有303,581行的可公开重新分发的韩语指令-问答语料库,通过无种子分类法驱动的GLAN合成管道生成,以微软的Phi-3.5-MoE-instruct作为生成模型(生成时间:2024年12月;发布时间:2024年12月;许可协议:OpenRAIL)。该语料库涵盖1084个英文标签学科的扁平分类法,并配有韩语问答文本,难度范围为100-900,每条记录的问题字符中位数为313个,答案字符中位数为1098个。对于这种规模的合成指令数据,有两个非典型属性:一是在303,581行中只有1个重复问题,在5000个样本的探测中,Jaccard >= 0.9的字符三元组近似重复簇数量为零;二是针对KMMLU、KoBEST(五个子任务)和HAE-RAE-Bench的两层污染审计显示,在20000个采样的GLAN问题和七个评估集中,测试与语料库问题级字符三元组Jaccard的最大值为0.163,Jaccard >= 0.7时测试项目为零,多语言E5余弦最大值为0.901,余弦 >= 0.90时有一个测试项目,余弦 >= 0.95时为零。据我们所知,在发布时,这是在Hugging Face Hub上可验证的最大的单管道合成韩语指令语料库,也是唯一按照无种子分类法驱动协议构建的韩语 >= 100k行语料库。本说明以适合下游引用的形式记录了生成协议、语料库统计、污染审计和许可边界。
英文摘要
We release GLAN-QnA-KR, a 303,581-row openly redistributable Korean instruction-QA corpus produced via the seedless taxonomy-driven GLAN synthesis pipeline with Microsoft's Phi-3.5-MoE-instruct as the producer model (generation: 2024-12; release: 2024-12; licence: OpenRAIL). The corpus spans a flat taxonomy of 1,084 English-labelled disciplines paired with Korean question/answer text, a 100-900 difficulty scale, and a median of 313 question characters and 1,098 answer characters per record. Two properties are atypical for synthetic instruction data at this scale: (i) exact duplicate questions number only 1 in 303,581 rows and character-trigram near-duplicate clusters at Jaccard >= 0.9 number zero in a 5,000-sample probe, and (ii) a two-layer contamination audit against KMMLU, KoBEST (five sub-tasks), and HAE-RAE-Bench shows a maximum test-vs-corpus question-level character-trigram Jaccard of 0.163 with zero test items at Jaccard >= 0.7, and a maximum multilingual-E5 cosine of 0.901 with a single test item at cosine >= 0.90 and zero at >= 0.95, across 20,000 sampled GLAN questions and seven evaluation sets. At the time of release, this is, to our knowledge, the largest single-pipeline synthetic Korean instruction corpus verifiable on the Hugging Face Hub and the only Korean >=100k-row corpus built under a seedless taxonomy-driven protocol. This note documents the generation protocol, corpus statistics, the contamination audit, and the licensing boundary in a form suitable for downstream citation.
CommentsTechnical report; 7 pages, 4 tables. Dataset: https://huggingface.co/datasets/daekeun-ml/GLAN-qna-kr-300k