arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

人工矫正:为什么大语言模型过度使用一种古典修辞格,以及如何缓解这一问题

Artificial Epanorthosis: Why large language models overuse a classical rhetorical figure, and how to mitigate it

Federico Boggia

arXiv 2607.21498首次发表:更新:

AI 中文总结

研究大语言模型过度使用矫正格修辞格的问题,基于模型与人类修辞风格差异及相关分类,提出用矫正格指数评分,通过测量发现校准错误,给出以LoRA适配器为中心的缓解技术、指令及适配器效果,强调校准到人类比率而非消除的重要性。

AI 中文摘要

一种两千年前西塞罗和昆体良记载过的修辞格——矫正格,系统性地出现在大语言模型的文本中,如“这不是一门课程。这是一场转变之旅”。本文认为这种过度使用是一种训练倾向,主要由富含宣传性散文的训练分布以及奖励自信、强调性措辞的偏好调整(基于人类反馈的强化学习)驱动;生成的从左到右的性质是放大器而非根本原因。基于模型与人类修辞风格不同的证据以及丰塔尼埃将矫正格归类为思维修辞格,提出通过矫正格指数(相对于人类比率的密度)根据特定体裁的人类基线对该修辞格进行评分的方案。对一个指令微调模型家族的三种规模进行的首次测量发现,在两个方向上都存在语域校准错误:模型在演讲中过度使用(约两倍,意大利语中接近三倍,集中在较大层级),在非正式问答写作中使用不足,而在论证、新闻和百科散文中与人类匹配。接着有三个建设性贡献:以轻量级LoRA适配器为中心的缓解技术调查;用意大利语证明一行指令可将该修辞格减少一半到近四分之三,监督微调适配器几乎可将其完全消除,缩放系数可将减少量调整到人类比率;以及认为目标是针对每种体裁校准到人类比率,而非消除。最后指出风险:真正的风险是我们开始像机器一样写作。

英文摘要

A rhetorical figure that Cicero and Quintilian catalogued two thousand years ago reappears, systematically, in the text of large language models: epanorthosis, the self-correction of the specimen «This is not a course. It is a journey of transformation». This essay argues that the overuse is a trained disposition, driven mainly by a training distribution rich in promotional prose and by preference tuning (RLHF) that rewards confident, emphatic phrasing; the left-to-right nature of generation is an amplifier rather than the root cause. Building on evidence that models diverge from human rhetorical style, and on Fontanier's classification of epanorthosis as a figure of thought, it sets out a programme that scores the figure against genre-specific human baselines through an Epanorthosis Index (density relative to the human rate). A first measurement, on three sizes of one instruction-tuned model family, finds mis-calibration by register in both directions: the models overshoot in oratory (about twofold, near threefold in Italian, concentrated in the larger tiers) and undershoot in informal question-and-answer writing, while matching humans in argument, journalism, and encyclopedic prose. Three constructive contributions follow: a survey of mitigation techniques centred on lightweight LoRA adapters; a demonstration, in Italian, that a one-line instruction cuts the figure by half to nearly three-quarters and that a supervised-fine-tuning adapter removes it almost entirely, with a scaling coefficient that dials the reduction back onto the human rate; and the argument that the target is calibration to the human rate for each genre, not elimination. It closes on the stakes: the real risk is that we begin to write like the machines.

Comments18 pages, 7 tables. v2: corrections to the classical sources (Quintilian, Cicero) and to several cited figures, and Appendix B corpus statistics aligned to the delivered dataset; measurements, results and conclusions unchanged. Data, code, and the trained LoRA adapter: https://federicoboggia.binatomy.com/pubblicazioni/

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑