arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.38621cs.AIcs.CL

当科学矛盾在翻译中丢失

When Scientific Contradictions Are Lost in Translation

  • Beakr

机构由 AI 辅助整理,请以论文原文为准。

Tal Zeevi, Trey W. Jensen, Maxwell Strome

AI总结:

本研究通过不可满足的异或约束系统,发现语言模型在科学报告语境下常偏好生物学预期而非约束满足,去除该偏好可显著提升正确赋值恢复率,表明科学验证依赖比较决策。

AI中文摘要:

两项科学发现可能彼此不一致,但并不一定相互矛盾。判断它们是否冲突,需要了解它们是否描述了可比较的测量。我们研究语言模型在此决策点上的行为。在一个受控任务中,我们生成一个不可满足的异或约束系统,并将其约束翻译成来自不同实验室的科学报告。一种赋值满足更多约束,而另一种赋值满足较少约束,但更符合预期的生物学。这造成了一个简单的困境:模型是选择最符合约束的赋值,还是选择更符合生物学预期的赋值?当约束被直接陈述时,GPT-5.6 Sol 和 Claude Opus 5 分别在 90% 和 96% 的情况下恢复了最受支持的赋值。然而,在科学散文中,模型的行为有所不同。Claude Opus 5 常常偏好生物学预期的赋值。去除该生物学偏好后,恢复更受支持的赋值的比例从 27% 增加到 79%(p<.001);当同一生物学偏好的记录附带形式化请求和明确的配对设计提示时,恢复率达到 92%(p<.001)。GPT-5.6 Sol 不太敏感,相应的变化均未达到统计显著性。这些结果表明,可靠的科学验证不仅依赖于形式推理,还依赖于模型如何决定哪些发现应被比较以及它们暗示何种关系。

英文摘要:

Two scientific findings can disagree without contradicting each other. Determining whether they conflict requires knowing whether they describe comparable measurements. We study how language models behave at this decision point. In a controlled task, we generate an unsatisfiable XOR constraint system and translate its constraints into scientific reports from different laboratories. One assignment satisfies more constraints, while another satisfies fewer but better matches expected biology. This creates a simple dilemma: does the model choose the assignment that best fits the constraints, or the one that better matches biological expectations? When the constraints are stated directly, GPT-5.6 Sol and Claude Opus 5 recover the best-supported assignment in 90% and 96% of cases, respectively. In scientific prose, however, the models behave differently. Claude Opus 5 often prefers the biologically expected assignment. Removing that biological preference increases recovery of the better-supported assignment from 27% to 79% (p<.001); recovery reaches 92% when the same Biology-favored record is accompanied by a formalization request and an explicit paired-design cue (p<.001). GPT-5.6 Sol is less sensitive, with neither corresponding change reaching statistical significance. These results suggest that reliable scientific verification depends not only on formal reasoning, but also on how models decide which findings should be compared and what relations they imply.

补充信息

↑