arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

度量-构建耦合夸大测得的合成方言恢复

Metric-Construction Coupling Inflates Measured Synthetic Dialect Recovery

Hyojung Han

arXiv 2610.00164首次发表:更新:

发表机构

ThakiCloud(ThakiCloud)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本研究揭示韩语方言合成数据恢复测量中,度量与构建共享清单会显著夸大结果,提出KoDialectBench基准及构建不相交测试,证明独立测量可避免高估。

AI 中文摘要

韩语方言语料库可用但不可再分发:可以发布权重,而无法基于底层数据重现训练和评估。我们探究合成数据能恢复多少这种监督信息,以及这种恢复能否独立于合成流程进行测量。我们贡献了KoDialectBench,包含五个地区、三个轴上的1000个项目,以标识符哈希和评分代码形式发布,使用户可以从自己获得许可的副本中重建项目。恢复强烈依赖于轴:我们最好的合成分支在地区识别上达到真实数据增益的91.2%,但在理解上仅为63.7%。在生成方面,答案取决于度量:部署的标记词典在方言性上报告92.3%,在地区匹配上报告119.3%,后者超过真实数据参考值,而基于参考的生成达到72.4%。我们发现标记度量的评分清单完全包含在我们的转换规则可以生成的清单中。我们通过一个精确形式的构建不相交分支直接测试构建访问的影响,该分支从规则中扣留20%的标记类型。在完全匹配的训练规模(8600个示例)下,它将保留标记清单上的方言性恢复从91.8%降至8.1%,地区匹配恢复从101.9%降至25.6%,而三个与流程无关的测量完全没有下降。一项补充评估器扫描定义了度量-构建覆盖率(MCC),并发现随着重叠从MCC=0增加到MCC=1,测得的方言性恢复单调增加。因此,共享的构建和评估清单会大幅夸大对合成数据恢复的估计。

英文摘要

Korean dialect corpora are available but not redistributable: weights may be released, while reproducing training and evaluation from the underlying data cannot be. We ask how much of that supervision synthetic data recovers, and whether that recovery can be measured independently of the synthesis pipeline. We contribute KoDialectBench, 1,000 items across five regions on three axes, released as identifier hashes and scoring code so users reconstruct the items from their own licensed copy. Recovery is strongly axis-dependent: our best synthetic arm reaches 91.2% of the real-data gain on region identification but 63.7% on comprehension. On generation the answer depends on the metric: the deployed marker lexicon reports 92.3% on dialectness and 119.3% on region match, the latter exceeding the real-data reference, whereas reference-based generation reaches 72.4%. We find the marker metrics' scoring inventory is entirely contained in the inventory our transformation rules can emit. We test the effect of construction access directly with an exact-form construction-disjoint arm that withholds 20% of marker types from the rules. At exactly matched training size (8,600 examples) it reduces dialectness recovery from 91.8% to 8.1% and region-match recovery from 101.9% to 25.6% on the held-out marker inventory, while the three pipeline-independent measurements do not fall at all. A complementary evaluator sweep defines metric-construction coverage (MCC) and finds measured dialectness recovery increasing monotonically as overlap rises from MCC=0 to MCC=1. Shared construction and evaluation inventories can therefore substantially inflate estimates of synthetic-data recovery.

Comments22 pages, 6 figures, 6 tables

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑