arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.05889cs.DLcs.AIcs.CLcs.CY

长破折号嵌入国会:大语言模型时代初期(2021-2025)美国国会新闻稿中长破折号频率的全人口水平上升

The em-dash em-beds in Congress: A population-level rise in em-dash frequency in U.S. congressional press releases at the dawn of the large-language-model era, 2021-2025

Przemysław Czuma

首次发表
浏览论文内容

中文总结 AI 辅助

该研究分析2021-2025年美国国会新闻稿,发现无空格长破折号频率在2025年翻倍,证实LLM辅助写作的扩散,长破折号可作为全人口水平的相关标记。

中文摘要 AI 辅助

大语言模型(LLMs)会在其辅助生成的文本中留下细微的风格痕迹,其中讨论最多的是长破折号(U+2014),尤其是无空格形式的“词---词”,这种形式在排版后的英文散文中很常见,但在美国新闻写作中并不常见,美国联合通讯社(AP)风格要求使用带空格的破折号。本研究旨在探究该痕迹是否能在国会新闻稿中被测量。本研究采用预先注册的设计(OSF:https://doi.org/10.17605/OSF.IO/U5NEY),分析了来自480个众议院和参议院办公室的146239份通过爬虫获取的2021-2025年新闻稿(开放国会新闻数据集),分析内容包括:每1000字符清洗后文本中无空格散文式长破折号的密度,采用带长度偏移的泊松/负二项模型,按办公室聚类。结果显示,2021-2024年间,该密度保持在每1000字符0.10-0.12的范围内,2025年升至0.217,超过四年基线的两倍;包含此类长破折号的新闻稿占比从约13%升至24.8%。主要频率比(2023-2025年 vs 2021-2022年)为1.55(95%置信区间1.28-1.93;预先注册的精确截止值为1.528),略高于预先设定的1.5倍阈值。该增长是净新增的(连字符密度保持稳定),在作者层面(262个连续办公室中的75.6%出现增长;p≈1e-16)和224个办公室的闭合面板中均存在,且通过了证伪测试:三个安慰剂截止值均为无效,分析流程在2024/2025边界处无突变,持续办公的办公室承载了该增长。分段回归发现,在ChatGPT出现的截止点无突变,但在之后时期出现明显加速;2025年的增长在各政党和议会间对称分布。由于预先注册的验证门正式被突破,未满足完整预先注册的决策规则,因此将“随着模型成熟,LLM辅助写作广泛传播”的解释作为探索性结论。长破折号仍是全人口水平的标记,而非逐篇新闻稿的作者身份检测器,且该设计不支持因果主张。

英文摘要

Large language models (LLMs) can leave small stylistic traces in text written with their help. The most discussed is the em-dash (U+2014), especially the unspaced form word---word, which is normal in typeset English prose but unusual in U.S. press writing, where AP style calls for spaced dashes. This study asks whether that trace is measurable in congressional press releases. In a preregistered design (OSF: 10.17605/OSF.IO/U5NEY), 146,239 scraper-sourced releases from 480 House and Senate offices (2021-2025, the open congress-press dataset) were analyzed: density of unspaced prose-form em-dashes per 1,000 characters of cleaned text, Poisson/negative-binomial models with a length offset, clustering by office. Density stayed within 0.10-0.12 per 1,000 characters through 2021-2024, then rose to 0.217 in 2025, more than twice the four-year baseline; the share of releases with such an em-dash rose from ~13% to 24.8%. The primary frequency ratio (2023-2025 vs 2021-2022) was 1.55 (95% CI 1.28-1.93; exact registered cut-off: 1.528), just above the prespecified 1.5x threshold. The rise was net-new (hyphen density stable), held within authors (75.6% of 262 continuous offices increased; p ~ 1e-16) and in a closed panel of 224 offices, and survived falsification tests: three placebo cut-offs were null, the pipeline showed no step at the 2024/2025 boundary, and continuing offices carried the rise. A segmented regression finds no step at the ChatGPT cut-off but a clear post-period acceleration; the 2025 rise is symmetric across parties and chambers. Because the registered validation gate was formally breached, the full preregistered decision rule was not met; the interpretation (broad diffusion of LLM-assisted writing as the models matured) is offered as exploratory. The em-dash remains a population-level marker, not a per-release authorship detector, and the design supports no causal claim.

发表机构

  • Polish Association for Artificial Intelligence in Medicine(波兰医学人工智能协会)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑