arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

BiasTrace:将大语言模型(LLM)中的推理行为与有偏输出关联起来

BiasTrace: Linking Reasoning Behaviours to Biased Outputs in LLMs

Varsha Ramineni, Hossein A. Rahmani, Jerome Ramos, Karin Sevegnani, Emine Yilmaz

arXiv 2608.14161首次发表:更新:

发表机构

University College London; NVIDIA(伦敦大学学院; 英伟达)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本研究提出BiasTrace标注方案,将LLM的推理行为与有偏输出关联,构建大型标注数据集,发现有偏输出多源于微妙推理行为,且该方案可提升偏见检测与推理时缓解效果。

AI 中文摘要

大语言模型(LLM)会表现出社会偏见,可能产生不准确且带有歧视性的推断,在高风险应用中带来风险。尽管现有工作在偏见测量与缓解方面已取得进展,但大多聚焦于模型的最终输出,对产生有偏结果的机制理解有限。LLM推理领域的近期进展为研究偏见提供了新视角,但推理与偏见之间的关联仍未被充分理解。现有方法主要关注最终答案的正确性或明确的有偏语言,忽略了可能导致有偏结果的不同推理行为。我们提出BiasTrace,这是一种用于标注模型生成的推理轨迹中推理行为并将其与有偏结果关联的标注方案。BiasTrace可捕获偏见特定行为(如无依据的人口统计假设)以及可能隐性加剧偏见的通用推理模式(如过度思考)。我们将BiasTrace应用于偏见敏感场景中的推理轨迹,通过经验证的LLM作为评判者(LLM-as-a-judge)方法进行扩展,生成了一个大型标注数据集。分析显示,有偏输出往往源于微妙的推理行为而非明确的有偏语言,且推理级标注可提升偏见检测效果。我们进一步表明,BiasTrace标注的行为可用于推理时(inference-time)的偏见缓解。这些发现强调,研究更广泛的推理模式对更好理解LLM中的偏见至关重要。

英文摘要

LLMs exhibit social biases that can produce inaccurate and discriminatory inferences, posing risks in high-stakes applications. While prior work has made progress in measuring and mitigating bias, it largely focuses on final outputs of models, with limited understanding of the mechanisms that produce biased outcomes. Recent advances in LLM reasoning offers a new lens for investigating bias, yet the link between reasoning and bias remains poorly understood. Existing approaches focus primarily on final answer correctness or explicitly biased language, overlooking different behaviours in reasoning that can drive biased outcomes. We introduce BiasTrace, an annotation scheme for labelling reasoning behaviours in model-generated traces and linking them to biased outcomes. BiasTrace captures bias-specific behaviours (e.g., unsupported demographic assumptions) as well as general reasoning patterns that may implicitly contribute to bias (e.g. overthinking). We apply BiasTrace to reasoning traces in bias-sensitive contexts, scaled using validated LLM-as-a-judge methods, producing a large annotated dataset. Our analysis shows that biased outputs often stem from subtle reasoning behaviours rather than explicitly biased language, and that reasoning-level annotations improve bias detection. We further show that BiasTrace behaviours can be exploited for inference-time mitigation. These findings underscore the importance of examining a broader range of reasoning patterns to better understand bias in LLMs.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑