什么驱动了大语言模型(LLM)的自我反思?武装冲突预测中不确定性路由的受控消融研究
What Drives LLM Self-Reflection? A Controlled Ablation of Uncertainty Routing in Armed Conflict Forecasting
浏览论文内容
中文总结 AI 辅助
本研究通过受控消融实验发现,武装冲突预测中LLM自我反思的增益源于类型化行动路由,而非诊断支架或分类学术语,且该机制在GPT-4o上复现,增益集中于结构新颖的冲突。
中文摘要 AI 辅助
自我反思被广泛认为可提升大语言模型(LLM)的推理能力,但驱动该提升的具体组成部分仍未被充分理解。本研究开展了包含六个条件的受控消融实验,分离出LLM自我反思的四个组成部分:证据暴露、诊断支架、分类学术语和行动路由。两项精确的零结果指向同一机制:其一,结构化诊断问题相较于非结构化反思未产生可测量的价值(F1值分别为0.296与0.297,p=1.000,95%置信区间[-0.041, +0.040]);其二,呈现完整的不确定性分类体系同时将行动空间简化为单一通用行动也未带来价值(F1增量为+0.008,95%置信区间重叠),排除了分类学术语作为驱动机制的可能。类型化行动路由则提供了一致的方向性增益(F1值为0.379,对照组为0.296),控制分类学术语的保守估计F1增量为+0.075,与单次基线相比的整体增益经自助法置信区间检验具有显著性(F1增量为+0.101,95%置信区间[+0.020, +0.185])。该术语-路由分解在GPT-4o上得到复现:分类学术语相较于通用反思未产生显著价值(p=0.773),而行动路由带来显著增益(p=0.025),证实该机制适用于不同的模型主干。增益集中于结构新颖的冲突:在缅甸(F1值从0.000提升至0.353)和乌克兰(F1值从0.167提升至0.500),仅含分类学术语的条件未比通用反思恢复更多信息,而行动路由打破了退化的先验。这些发现确定类型化行动路由——而非诊断支架或分类学术语——是元认知LLM预测智能体的有前景设计原则,同时推动在更多冲突类型中开展更大规模的评估。
英文摘要
Self-reflection is widely assumed to improve LLM reasoning, yet which component drives the gain remains poorly understood. We present a controlled six-condition ablation isolating four components of LLM self-reflection: evidence exposure, diagnostic scaffolding, taxonomy vocabulary, and action routing. Two precise null results converge on a single mechanism. First, structured diagnostic questions add no measurable value over unstructured reflection ($\text{F1} = 0.296$ vs $0.297$, $p = 1.000$, 95\% CI $[-0.041, +0.040]$). Second, presenting the full uncertainty taxonomy while collapsing the action space to a single generic action also adds no value ($Δ\text{F1} = +0.008$, overlapping 95\% CIs), ruling out taxonomy vocabulary as the mechanism. Typed action routing provides consistent directional gains ($\text{F1} = 0.379$ vs $0.296$); the conservative estimate controlling for taxonomy vocabulary is $Δ\text{F1} = +0.075$, and the overall gain over the single-shot baseline is significant by bootstrap CI ($Δ\text{F1} = +0.101$, 95\% CI $[+0.020, +0.185]$). The vocabulary-routing decomposition replicates on GPT-4o: taxonomy vocabulary adds no significant value over generic reflection ($p = 0.773$), while action routing provides significant gains ($p = 0.025$), confirming the mechanism holds across backbones. Gains concentrate on structurally novel conflicts: in Myanmar ($\text{F1}: 0.000 \rightarrow 0.353$) and Ukraine ($0.167 \rightarrow 0.500$), the vocabulary-only condition recovers no more than generic reflection while action routing breaks the degenerate prior. These findings identify typed action routing -- not diagnostic scaffolding or taxonomy vocabulary -- as a promising design principle for metacognitive LLM forecasting agents, while motivating larger-scale evaluation across conflict typologies.
发表机构
- University of North Texas(北得克萨斯大学)
- College of Computer Science and Engineering(计算机科学与工程学院)
机构由 AI 辅助整理,请以论文原文为准。