相同证据,不同目标:从语言模型状态解码诊断证据如何关联因果问题
Same Evidence, Different Target: Decoding How Diagnostic Evidence Bears on Causal Questions from Language-Model States
浏览论文内容
中文总结 AI 辅助
该研究通过配对提示改变因果目标,分析Qwen2.5等模型的隐藏状态,证明其含可线性解码的诊断证据与因果目标关联信息,平衡准确率达0.654-0.659。
中文摘要 AI 辅助
当因果主张涉及不同总体、结局、估计量、路径或识别假设时,相同的诊断结果可能支持或挑战某一因果主张,却无法回应另一主张。当证据与目标共同变化时,正确答案可能反映有利或不利的措辞、词汇重叠或熟悉的诊断模式,而非证据与因果问题的匹配。我们引入配对提示,逐字重复相同的诊断证据,同时改变因果目标;每个提示根据证据与因果问题的关联方式被标记为支持、挑战、未解决或目标错误,仅当两个提示均被正确分类时,配对才被成功恢复。使用在单独开发集上训练的线性读出器,我们分析Qwen2.5-7B-Instruct、Qwen3-8B和Llama-3.1-8B-Instruct的倒数第二个Transformer块的最终标记隐藏状态。在涵盖9个诊断家族的49对主要基准上,平衡准确率介于0.654至0.659之间,成功恢复18至21对;两名独立人类评审员对98个提示中的95个(96.9%)赋予相同标签。在所有检查点上,平衡准确率和完整配对恢复率均超过保留开发场景组的排列空值。在Qwen2.5中,完整提示的平衡准确率超过受限输入,且两种差异的配对自举区间均大于零;在未使用所评估诊断家族的开发示例训练的读出器中,成功恢复21对,包括9个家族中各至少1对。隐藏状态读出器在平衡准确率和恢复配对数量上均超过基于答案选项对数几率的线性分类器和文本基线。这些结果表明,隐藏状态包含可线性解码的信息,用于判断诊断证据是否支持、挑战或无法回应因果目标。
英文摘要
The same diagnostic result can support or challenge one causal claim yet fail to address another when the claims concern different populations, outcomes, estimands, pathways, or identifying assumptions. When the evidence and target vary together, a correct answer may reflect favorable or adverse wording, lexical overlap, or a familiar diagnostic pattern rather than matching the evidence to the causal question. We introduce paired prompts that repeat the same diagnostic evidence verbatim while changing the causal target. Each prompt is labeled Favors, Challenges, Unresolved, or Wrong Target according to how the evidence bears on the causal question. A pair is recovered only when both prompts are classified correctly. Using linear readouts trained on a separate development set, we analyze the final-token hidden state from the penultimate transformer block of Qwen2.5-7B-Instruct, Qwen3-8B, and Llama-3.1-8B-Instruct. On the 49-pair primary benchmark spanning nine diagnostic families, balanced accuracy ranges from 0.654 to 0.659 and 18-21 pairs are recovered. Two independent human reviewers assigned the same label to 95 of the 98 prompts (96.9%). Across checkpoints, balanced accuracy and complete-pair recovery exceed permutation nulls that preserve development scenario groups. In Qwen2.5, full-prompt balanced accuracy exceeds both restricted inputs, with paired-bootstrap intervals for both differences above zero. Readouts trained without development examples from the evaluated diagnostic family recover 21 pairs, including at least one in each of the nine families. The hidden-state readout exceeds a linear classifier on answer-option logits and text baselines in balanced accuracy and recovered pairs. These results show that the hidden state contains linearly decodable information about whether diagnostic evidence favors, challenges, or fails to address the causal target.