AI 中文总结
本研究针对LLM推理中思维链提示引发的偏差问题,提出基于在线变点检测的自适应触发校正框架,在gpt-4o-mini等模型上验证了其在减少干预次数的同时保留准确率的优势。
AI 中文摘要
思维链提示会暴露并放大大语言模型(LLM)中间推理过程中的人口统计刻板印象,形成仅靠最终答案偏差校正无法解决的失效模式。在生成过程中缓解此类偏差存在一个根本的时机问题:干预过晚会让有偏差的推理扩散,而不必要的干预则会破坏原本正确的推理。现有方法大多通过事后评估完整推理链或在预定步骤进行干预来回避该决策,未解决何时推理轨迹发展到提供足够证据以保证校正的问题。我们将该决策表述为在线变点检测问题,每一步的偏差信号会更新CUSUM统计量,仅当累积证据超过在保留数据上校准的检测器特定阈值时才注入针对性校正。我们用来自下一个token概率的白盒信号和来自LLM评判器的黑盒信号实例化该框架,使其可部署于开放权重模型和托管模型。在gpt-4o-mini上,自适应黑盒触发机制在固定间隔干预下损失的消歧上下文准确率中恢复了大部分,且所需干预次数显著更少,该结果在使用独立评判器时依然成立。在六个开放权重模型上,白盒信号提升了所有六个模型的歧义项准确率,但降低了其中五个模型的消歧项准确率,因其无法区分无依据的刻板印象依赖与正确的、符合刻板印象的证据。
英文摘要
Chain-of-thought prompting can expose and amplify demographic stereotypes within an LLM's intermediate reasoning and create a failure mode that final-answer debiasing alone cannot address. Mitigating such bias during generation presents a fundamental timing problem: intervening too late allows biased reasoning to propagate, while unnecessarily intervening can disrupt otherwise correct reasoning. Existing approaches largely avoid this decision by either evaluating completed reasoning chains post hoc or intervening at predetermined steps, leaving open when a developing reasoning trajectory provides sufficient evidence to warrant correction. We formulate this decision as an online change-point detection problem. A per-step bias signal updates a CUSUM statistic and a targeted correction is injected only when accumulated evidence crosses a detector-specific threshold calibrated on held-out data. We instantiate the framework with a white-box signal derived from next-token probabilities and a black-box signal obtained from an LLM judge, enabling deployment with both open-weight and hosted models. On gpt-4o-mini adaptive black-box triggering recovers most of the disambiguated-context accuracy lost under fixed-interval intervention while requiring substantially fewer interventions. That result holds even with an independent judge. Across six open-weight models, the white-box signal improves ambiguous-item accuracy on all six but reduces disambiguated-item accuracy on five because it cannot distinguish unsupported stereotype reliance from correct, stereotype-congruent evidence.
Comments10 pages, 6 figures, Under review