arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

通过措辞与选择进行框架构建:2022-2025年法国新闻标题的大规模二维审计

Framing by Wording, Framing by Selection: A Large-Scale Two-Dimensional Audit of French News Headlines, 2022-2025

Amr Sobhy

arXiv 2609.28487首次发表:更新:

发表机构

Le French News Lab(法国新闻实验室)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

针对2022-2025年法国新闻标题,提出二维框架分离措辞与选择框架,基于10,000条监督集和25家媒体90万+标题,发现显著性估计偏差并揭示群体显著性不平等,为最大规模法语标题审计。

AI 中文摘要

新闻标题通过选择内容和措辞方式对公共议题进行框架构建,然而计算框架研究通常将这两种操作合并为一个分数。我们引入了一个二维框架,将显著性框架(通过四种措辞手段衡量:情绪化词汇、责任归因、威胁框架、反问句)与选择框架(通过媒体层面的故事形态和高冲击力分布衡量)分离开来。我们使用三个LLM标注器,通过多数投票决议和人工仲裁,构建了一个包含10,000条法语标题的监督集,并针对两个独立于标注器的盲人研究验证了标签,然后将最强的分类器应用于来自25家法国媒体(2022-2025年)的902,111条去重标题。主要发现有三点。第一,显著性和选择性的分歧呈正相关,但仍有近一半的媒体层面方差无法解释,在四格媒体类型学中填充了解释上不同的非对角线单元格。第二,默认分类阈值系统性地夸大了语料库层面的显著性估计;一种基于精确度下限的重新校准协议纠正了这种扭曲。第三,群体提及分析揭示了显著不平等的显著性语境:提及犹太人、极右翼和穆斯林的标题具有最高的检测显著性率,而广泛的事件语境构成并不能完全解释这一点(残差是描述性的,而非同事件因果估计;每个群体的词典精确度一并报告)。据我们所知,这是迄今为止规模最大的以框架为重点的法语标题审计;我们发布了监督集、词典和分析代码。

英文摘要

News headlines frame public issues both by what they select and by how they word it, yet computational framing work typically collapses these operations into a single score. We introduce a two-dimensional framework that separates salience framing, measured through four wording devices (loaded vocabulary, blame attribution, threat framing, rhetorical question), from selection framing, measured through outlet-level story-form and high-charge distributions. We build a 10,000-headline French supervision set using three LLM annotators with majority-vote resolution and human arbitration, validate the labels against two annotator-independent blind human studies, and apply the strongest classifier to 902,111 deduplicated headlines from 25 French outlets (2022-2025). Three main findings emerge. First, salience and selection divergence are positively correlated yet leave nearly half of outlet-level variance unexplained, populating interpretively distinct off-diagonal cells in a four-cell outlet typology. Second, default classification thresholds systematically inflate corpus-level salience estimates; a precision-floor recalibration protocol corrects this distortion. Third, group-mention analysis reveals sharply unequal salience contexts: headlines mentioning Jews, the Far-right, and Muslims carry the highest detected salience rates, which broad event-context composition does not fully explain (residuals are descriptive, not same-event causal estimates; per-group lexicon precision is reported alongside). To our knowledge, this is the largest framing-focused French headline audit to date; we release the supervision set, lexicons, and analysis code.

Comments20 pages, 1 figure, includes appendices. Accepted for oral presentation at ICNLSP 2026

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑