arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.29266cs.AI

Baszta:波兰语多标签安全分类器的数据中心微调

Baszta: Data-Centric Fine-Tuning of a Polish Multi-Label Safety Classifier

Adam Górski, Mateusz Jąkalak, Rafał Jakubowski

首次发表
浏览论文内容

中文总结 AI 辅助

本研究通过微调波兰语模型开发多标签安全分类器,在分布外基准上优于对比模型,并发现鲁棒性需直接优化而非依赖分布内精度。

中文摘要 AI 辅助

我们开发了一个多标签波兰语内容安全分类器,通过对allegro/herbert-base-cased(124M)模型进行微调,覆盖五个类别(仇恨言论、粗俗内容、性内容、犯罪、自残),并使用Focal + R-Drop目标函数,然后在共享的分布外Gadzi Język基准上,将所得模型与Bielik Guard (Sójka)进行评估。两个系统都在相同的校准分割上进行了逐类别阈值调整。在该匹配协议下,我们的模型在微观F1上保持了一个小但统计显著的领先优势,而表观上的宏观F1领先优势并未保持:这是将调优模型与未调优模型进行比较所产生的假象。我们还报告了该微观数值的实际价值。由于Gadzi Język中97%的样本为犯罪阳性,一个对每个输入都标记犯罪且不标记其他内容的分类器在同一测试分割上已经获得了0.910的微观F1分数,因此微观指标无法区分任一系统与退化策略,而宏观指标才是能区分的指标。逐类别和逐协议的数值在第4节中报告。残余的分布外差距是校准问题而非判别问题。排序质量保持较高,而正概率坍缩,逐类别温度缩放恢复了损失,而Platt缩放和等渗回归未能做到。这种恢复被证明取决于校准集是否包含安全文本。Gadzi Język几乎不包含安全文本,因此在其上拟合的阈值会对每个安全输入都标记犯罪,而平衡的重新拟合以对抗性召回为代价换来了一个可部署的工作点。我们报告了两个工作点,而不仅仅是表现较好的那个。两种标准做法,即逐类别成本敏感加权和均值池化,各自提高了分布内宏观F1,同时降低了分布外数值,这表明鲁棒性必须直接选择,而不能从分布内准确性中继承。

英文摘要

We develop a multi-label Polish content-safety classifier by fine-tuning allegro/herbert-base-cased (124M) across five categories (hate, vulgarity, sexual content, crime, self-harm) using a Focal + R-Drop objective, and evaluate the resulting model against Bielik Guard (Sójka) on the shared out-of-distribution Gadzi Język benchmark. Both systems are given per-category threshold tuning on the same calibration split. Under that matched protocol our model holds a small but statistically significant lead in micro F1, while an apparent macro-F1 lead does not survive: it was an artifact of comparing a tuned model against an untuned one. We also report what that micro figure is worth. Because Gadzi Język is 97% crime-positive, a classifier that flags crime on every input and nothing else already scores 0.910 micro F1 on the same test split, so micro separates neither system from a degenerate strategy and macro is the column that does. Per-category and per-protocol figures are reported in Section 4. The residual out-of-distribution gap is one of calibration rather than discrimination. Ranking quality stays high while positive probabilities collapse, and per-category temperature scaling recovers the loss where Platt scaling and isotonic regression do not. That recovery turns out to be conditional on the calibration set containing safe text. Gadzi Język contains almost none, so thresholds fitted on it flag crime on every safe input, and a balanced refit buys a deployable operating point at the cost of adversarial recall. We report both operating points rather than only the flattering one. Two changes that are standard practice, per-class cost-sensitive weighting and mean pooling, each raise in-distribution macro F1 while lowering the out-of-distribution figure, which indicates that robustness has to be selected for directly rather than inherited from in-distribution accuracy.

发表机构

  • Billennium S.A.(Billennium公司)

机构由 AI 辅助整理,请以论文原文为准。

↑