多通道缓解事实核查强化学习智能体中的来源信任捷径
Multi-Channel Mitigation of Source-Trust Shortcuts in Fact-Checking RL Agents
- University of Connecticut(康涅狄格大学)
- University of California, Santa Cruz(加州大学圣克鲁兹分校)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
针对事实核查强化学习智能体依赖来源标签而忽略证据内容的捷径,提出TrustSwap测试与信任交换增强(TSA)训练方法,在4B规模下显著降低判定翻转率并保持准确率,8B规模效果不明显。
AI中文摘要:
检索增强的事实核查器通常会为每个证据来源接收一个可靠性标签,例如高信任或低信任。这些标签应调整模型的置信度及其搜索更多证据的决策,而最终判定应遵循证据内容。我们引入了TrustSwap,一种反事实测试,它在保持所有证据文本不变的情况下交换、降低或移除来源标签,并分别测量其三个输出通道(判定、置信度和搜索决策)。在未训练和强化学习训练的模型上,跨越两个规模、三个数据集和两种提示,置信度和搜索在50次比较中有49次按预期响应标签,然而仅标签变化就改变了Qwen3模型4-23%的自信判定,对现有强化学习训练的事实核查器则高达50%。标准的GRPO微调在8B规模下于全部六种设置中放大了这一捷径。为减少该捷径,我们提出了信任交换增强(TSA),该方法在每条主张上使用其原始证据和标签交换后的证据,在相同金标准判定下训练GRPO。在4B规模下,TSA在六种设置中的四种将判定翻转率相对降低了7-35%,保持了准确率以及预期的置信度和搜索响应,在主要设置中优于基于奖励的替代方法,并泛化到未见过的标签移除扰动。添加一致性奖励有助于训练中的交换,但对未见过的扰动无效。在8B规模下,TSA的效果不可检测,这使得规模成为主要待解问题。
英文摘要:
Retrieval-augmented fact-checkers often receive a reliability label, such as HIGH or LOW trust, for each evidence source. These labels should adjust the model's confidence and its decision to search for more evidence, while the verdict should follow the evidence content. We introduce TrustSwap, a counterfactual test that swaps, lowers, or removes source labels while keeping every evidence text fixed, and measures its three output channels (the verdict, the confidence, and the search decision) separately. Across untrained and RL-trained models at two scales, three datasets, and two prompts, confidence and search respond to the labels as intended in 49 of 50 comparisons, yet a label change alone alters 4-23% of confident verdicts for Qwen3 models and up to 50% for an existing RL-trained fact-checker. Standard GRPO fine-tuning amplifies this shortcut at 8B in all six settings. To reduce it, we propose trust-swap augmentation (TSA), which trains GRPO on each claim with both its original and its label-swapped evidence under the same gold verdict. At 4B, TSA lowers the verdict flip rate by 7-35% (relative) in four of six settings, keeps accuracy and the intended confidence and search responses, outperforms reward-based alternatives in the main setting, and carries over to an unseen label-removal perturbation. An added consistency reward helps on the trained-on swap but not on unseen perturbations. At 8B, TSA's effect is not detectable, which makes scale the main open question.