自生成文本识别:LLM评估中的质量启发式、跨任务迁移与下游偏差
Self-Generated Text Recognition: Quality Heuristics, Cross-Task Transfer, and Downstream Bias in LLM Evaluation
浏览论文内容
中文总结 AI 辅助
该研究聚焦LLM的自生成文本识别(SGTR)能力,明确了导致过往研究结论分歧的实验设计因素,发现SGTR性能可跨任务迁移,且训练SGTR会引发下游偏差,强调需监控SGTR以保障AI安全。
中文摘要 AI 辅助
自生成文本识别(Self-Generated Text Recognition, SGTR)是指大型语言模型(LLM)识别自身输出的能力,它对依赖LLM作为评估器或监控器的AI安全防护构成风险。具体而言,LLM可能识别出来自同模型其他副本的输出,进而做出有偏差的判断或直接合谋。过往研究对当前模型是否具备显著SGTR能力得出了相互矛盾的结论。我们通过识别导致结果分歧的关键实验设计选择(我们称之为操作化方式)来协调这些发现。在六种操作化方式下评估13至21个模型,我们发现准确率随评估形式(文本的成对评估与单独评估)、对话结构(在用户标签与助手标签中呈现候选文本)以及用于生成候选文本的任务领域(例如编码与摘要)存在显著差异。我们证实了先前的观察结果:质量启发式——模型将作者身份归因于它们认为质量更高的文本——是主要的混淆因素。我们还发现,通过在一种评估配置中用SFT(监督微调)提升模型的SGTR性能,可将其泛化到其他配置。训练SGTR还会导致模型在AlpacaEval框架中作为评估者时更偏好自身的输出。最后,我们讨论了我们的评估对未来AI系统安全的影响:我们的研究表明,尽管存在混淆因素,部分模型仍具备实用的SGTR能力,且在一种环境中训练模型进行SGTR会更广泛地影响其自我识别与自我偏好。我们得出结论,SGTR应在安全关键型AI应用的设计中受到监控并被纳入考量。
英文摘要
Self-Generated Text Recognition (SGTR)--the ability of an LLM to identify its own outputs--poses risks to AI safeguards that rely on LLMs as evaluators or monitors: an LLM may recognize outputs from other copies of the same model and make biased judgments or collude outright. Prior work has drawn conflicting conclusions about whether current models possess significant SGTR capabilities. We explain these disagreements by identifying key experimental design choices--which we term operationalizations--that drive divergent results. Evaluating 13-21 models across six presentation operationalizations and four task-domain operationalizations, we find that accuracy varies substantially with evaluation format (pairwise vs individual assessments of text), conversation format (presenting candidate text in user tags vs assistant tags), and the domain of the task used to generate candidate text (e.g., coding vs summarization). We corroborate previous observations that a quality heuristic--models attributing authorship to text they perceive as higher quality--is a dominant confound. We also find that improving a model's SGTR performance via supervised fine-tuning (SFT) on one operationalization can generalize to others, and can increase the model's preference for its own outputs when it acts as a judge in the AlpacaEval framework. Our results suggest that, despite confounds, some models possess practical SGTR capabilities, and that SGTR should be monitored and considered in the design of safety-critical AI applications.
发表机构
- George Washington University(乔治·华盛顿大学)
- Geodesic Research(吉奥戴斯克研究机构)
- University of Cambridge(剑桥大学)
机构由 AI 辅助整理,请以论文原文为准。