发表机构
Huazhong University of Science and Technology(华中科技大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本研究通过大规模平台数据诊断出AI偏好与真实用户参与度之间的系统性差距,提出OMRA干预方法,将差距平均减少54.4%,并在人工评估中胜率62.4%。
AI 中文摘要
大型语言模型越来越多地被用于生成和评估在线内容,然而,它们所认为的与更高参与度相关的特质是否与真实用户的反应相符,目前仍不清楚。我们利用来自知乎、Quora和Reddit的25,978个问题下的117万个回答来研究这一问题,在四个问题内参与度水平上比较了真实平台回答与AI生成回答。我们引入了本体论偏好测量(Ontological Preference Measurement),该方法从逻辑、情感和表达三个维度表征回答。我们发现AI偏好与真实用户参与度之间存在系统性差距:随着目标参与度的提高,LLM会添加更多显式的逻辑结构,而真实用户的参与度则与情感和表达的显著性关联更强。我们将这种倾向称为逻辑过度绑定(logic overbinding)。基于这一诊断,我们提出了本体掩蔽推理自编码(Ontology-Masked Reasoning Autoencoding, OMRA),这是一种受控干预方法,它在保留立场、事实内容和连贯性的同时,掩蔽并重建过度解释的片段。在四个LLM家族中,OMRA将测量到的差距平均减少了54.4%。在人工评估中,OMRA在与匹配的真实平台回答进行的两两偏好判断中赢得了62.4%的胜率,尽管真实回答更常被判定为人类撰写。
英文摘要
Large language models are increasingly used to generate and evaluate online content, yet it remains unclear whether the qualities they associate with higher engagement match what real users respond to. We study this question using 1.17 million answers to 25,978 questions from Zhihu, Quora, and Reddit, comparing real platform answers and AI-generated answers across four within-question engagement levels. We introduce Ontological Preference Measurement, which represents answers along three dimensions: logic, affect, and expression. We find a systematic gap between AI preference and real user engagement: as target engagement increases, LLMs add more explicit logical structure, while real user engagement is more strongly associated with affective and expressive salience. We call this tendency logic overbinding. Based on this diagnosis, we propose Ontology-Masked Reasoning Autoencoding (OMRA), a controlled intervention that masks and reconstructs over-explained spans while preserving stance, factual content, and coherence. Across four LLM families, OMRA reduces the measured gap by an average of 54.4%. In human evaluation, OMRA wins 62.4% of pairwise preference judgments against matched real platform answers, even though the real answers are more often judged to be human-written.