arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.18960cs.CL

当审计质量无法预测下游效用:面向低资源非洲NLP的合成数据选择器的反事实研究

When Audit Quality Fails to Predict Downstream Utility: A Counterfactual Study of Synthetic-Data Selectors for Low-Resource African NLP

Son Ha Xuan, Phat T. Tran-Truong, Xuan-Bach Le

首次发表
浏览论文内容

中文总结 AI 辅助

本研究通过反事实框架证明,在低资源非洲语言分类中,合成数据选择的审计质量无法预测下游效用,审计与下游排名显著分歧,强调应同时报告两类指标。

中文摘要 AI 辅助

质量感知的合成数据选择依赖于一个代理指标:LLM评判为优质的样本也应有助于下游模型的学习。在低资源非洲语言分类的受控重放实验中,我们证明该代理指标失效。在四种语言(阿姆哈拉语、豪萨语、斯瓦希里语、约鲁巴语)、两个分类任务(MasakhaNEWS、AfriSenti)以及五个预算匹配的选择器中,审计排名与下游排名出现分歧。在每个单元内,评判标签正确性与Macro-F1之间的Spearman相关系数均值为ρ=0.04(中位数0.00),表明这种不匹配并非聚合伪影。我们的反事实审计框架\method{}-V2在三个审计通道上同时产生最干净的选定池:最高的评判标签正确性(0.904对比naive的0.767,相对提升17.9%)、最低的捷径分数,以及硬拒绝率0.162对比naive的0.486。然而,AlpaGasus在下游Macro-F1上仍然领先(0.202对比\method{}-V2的0.163),且这种反转在五个非退化单元中持续存在。方法学上的教训是:在此受控设置中,审计质量是选定池的属性,而非下游效用的保证。因此,合成数据评估应在相同的保留集上报告审计和下游指标。我们发布了审计表、各选择器的保留池,以及将每个报告数字链接到其来源行的声明账本。

英文摘要

Quality-aware synthetic-data selection rests on a proxy: examples that an LLM judge rates as good should also help a downstream model learn. In a controlled replay in low-resource African-language classification, we show that this proxy breaks. Across four languages (Amharic, Hausa, Swahili, Yoruba), two classification tasks (MasakhaNEWS, AfriSenti), and five matched-budget selectors, audit rankings and downstream rankings diverge. Within each cell, the Spearman between judged label correctness and Macro-F1 across selectors has mean $ρ{=}0.04$ (median $0.00$), showing that the mismatch is not an aggregation artifact. \method{}-V2, our counterfactual audit framework, produces the cleanest selected pool on three audit channels at once: highest judged label correctness ($0.904$ vs.\ $0.767$ for naive, a $17.9\%$ relative gain), lowest shortcut score, and a hard-reject rate of $0.162$ vs.\ $0.486$ for naive. AlpaGasus nevertheless leads downstream Macro-F1 ($0.202$ vs.\ $0.163$ for \method{}-V2), and the inversion persists on the five non-degenerate cells. The lesson is methodological: in this controlled setting, audit quality is a property of the selected pool, not a guarantee of downstream utility. Synthetic-data evaluation should therefore report audit and downstream metrics on the same retained sets. We release the audit tables, per-selector retained pools, and a claim ledger that links every reported number to its source row.

发表机构

  • RMIT University(皇家墨尔本理工大学)
  • Ho Chi Minh City University of Technology (HCMUT), VNU-HCM(胡志明市理工大学(HCMUT),越南国立大学胡志明市分校)

机构由 AI 辅助整理,请以论文原文为准。

↑