发表机构
Le French News Lab(法国新闻实验室)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本研究构建了跨出版商法语新闻编辑部分类基准FrenchNews-7,采用微调CamemBERT分类器,实验显示其在召回率上优于零样本LLM基线,为新闻分类研究提供了可靠基准。
AI 中文摘要
我们提出FrenchNews-7,这是一个基于法国的跨出版商法语新闻编辑部分类基准,它结合了大型多出版商语料库、由URL衍生的七类分类法以及微调后的CamemBERT分类器。标签通过混合流程分配,该流程结合了出版商URL段和针对结构模糊案例的LLM注释,并通过评分者间研究(2名人类+2个LLM)进行审核,其中两两κ≥0.766,人类-人类κ=0.806。我们在分布内和保留出版商设置下评估了词汇、多语言和特定于法语的训练分类器,还在保留池上与零样本LLM基线(GPT-OSS-120B、Mistral Small 3.2、Llama-3.3-70B)进行了额外比较。最强的模型是使用完整文章文本的CamemBERT-base,它优于仅标题输入,可推广到未见过的出版商,并且在整体召回率(0.799)上超过了所有三个零样本LLM基线,差距集中在模糊的编辑部边界类别Economie和Societe。跨出版商评估显示边界稳定性不均:Sport、Culture & Loisirs和International可清晰迁移,而Economie(召回率=0.517)接近盲人类一致性(0.55),Societe(精确率=0.577)吸收了边界模糊性,两者都表明是编辑部惯例而非可恢复的分类器提升空间。微调后的CamemBERT-base模型、标签清单、参考收集脚本和可靠性层级指导表可在该模型URL和该数据集URL获取。
英文摘要
We present FrenchNews-7, a cross-publisher France-based French-language news editorial desk classification benchmark combining a large multi-outlet corpus, a URL-derived seven-class taxonomy, and a fine-tuned CamemBERT classifier. Labels are assigned via a hybrid pipeline combining publisher URL slugs with LLM annotation for structurally ambiguous cases, audited through an inter-rater study (2 humans + 2 LLMs; pairwise $κ\geq 0.766$, human--human $κ= 0.806$). We evaluate lexical, multilingual, and French-specific trained classifiers under both in-distribution and held-out-publisher settings, with additional comparison against zero-shot LLM baselines (GPT-OSS-120B, Mistral Small 3.2, Llama-3.3-70B) on the held-out pool. The strongest model, CamemBERT-base on full article text, outperforms headline-only input, generalizes to unseen outlets, and exceeds all three zero-shot LLM baselines on overall recall (0.799), with the gap concentrated in the ambiguous editorial-boundary categories Economie and Societe. Cross-publisher evaluation reveals uneven boundary stability: Sport, Culture & Loisirs, and International transfer cleanly, while Economie (recall = 0.517) is close to blinded human agreement (0.55), and Societe (precision = 0.577) absorbs boundary ambiguity, both suggesting editorial conventions rather than recoverable classifier headroom. The fine-tuned CamemBERT-base model, labeled manifest, reference collection scripts, and a reliability-tier guidance table are available at https://huggingface.co/LeFrenchNewsLab/camembert-base-frenchnews7 (model) and https://huggingface.co/datasets/LeFrenchNewsLab/frenchnews-7 (dataset).
Comments15 pages, 5 figures, includes appendices. Model and dataset available on HuggingFace