思考有助于公平吗?推理标记解决了一些偏见,但制造了更多
Does Thinking Help Fairness? Reasoning Tokens Resolve Some Biases but Create More
浏览论文内容
中文总结 AI 辅助
本研究通过消融实验发现,推理语言模型中的思考对反事实公平性具有不对称双重效应:虽解决部分偏见但创造更多(约5倍),并提出了CDPG和BTM两种工具来追踪和解释这一现象。
中文摘要 AI 辅助
在推理语言模型(RLMs)中,关于思考是否能解决或放大偏见一直存在争议。先前的研究在两个方向上都得出了相互矛盾的结论。通过在QwQ-32B、DeepSeek-R1-Distill-Qwen-32B和Qwen3-32B上对三个高风险决策任务(Adult、COMPAS、Credit)进行模型内思考与不思考的消融实验,我们发现思考对反事实公平性具有不对称的双重效应:它既解决了不思考基线产生的反事实翻转,又在接近饱和的模型置信度下创造了新的翻转。在所有九种(模型,数据集)组合中,创造的翻转数量大约是解决的翻转数量的5倍。为了解释这一效应,我们将思考轨迹本身视为公平性变化的可测量位点,并通过两个动态工具进行研究:1) 我们提出了反事实深度概率差距(CDPG)来追踪偏见沿思考深度的演变,并观察到偏见随思考传播和放大。2) 我们还构建了偏见转移矩阵(BTM)来展示反事实对的预测如何从不思考转变为思考,并发现不对称双重效应源于对状态联合转移。
英文摘要
Thinking in reasoning language models (RLMs) has been subject to debate on whether it resolves or amplifies bias. Prior works have shown competing conclusions in both directions. Using a within-model thinking-vs.-non-thinking ablation across QwQ-32B, DeepSeek-R1-Distill-Qwen-32B, and Qwen3-32B on three high-stakes decision tasks (Adult, COMPAS, Credit), we show that thinking has an asymmetric dual effect on counterfactual fairness: it both resolves counterfactual flips produced by the non-thinking baseline and creates new flips at near-saturating model confidence. In all nine (model, dataset) combinations, the created flips outnumber the resolved flips by roughly 5 times. To explain the effect, we treat the thinking trace itself as a measurable site of fairness change and study it through two dynamic instruments: 1) We propose Counterfactual Depth Probability Gap (CDPG) to track bias evolution along thinking depth, and observe that bias propagates and amplifies with thinking. 2) We also formulate the Bias Transition Matrix (BTM) to show how predictions of counterfactual pairs change from non-thinking to thinking, and find that the asymmetric dual effect originates in the pair-state joint transition.
发表机构
- University of Notre Dame(圣母大学)
- IBM Research, Dublin(IBM研究院(都柏林))
机构由 AI 辅助整理,请以论文原文为准。