arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2607.24765cs.CLcs.AIcs.HC

通过事实-启发式-情感状态强制来测量和提高大语言模型中的行为一致性

Measuring and Improving Behavioral Consistency in Large Language Models through Fact-Heuristic-Emotion State Enforcement

Gi-Hun Lee, Joong Yull Park

首次发表
浏览论文内容

中文总结 AI 辅助

研究大语言模型行为一致性,通过认知内核模型CKM强制模型在决策前区分事实、假设和评估信号,经多实验评估发现其可降低输出变异性、减少决策翻转率,证明行为一致性可测且能部分改善。

中文摘要 AI 辅助

大语言模型(LLMs)在处理相同决策问题时多次运行会给出不同答案,甚至会因之前的答案作为上下文而改变决策。本文探讨在不改变模型权重的情况下,这种不稳定性能否被测量并部分降低。通过测试认知内核模型(CKM),它是一个提示级别的状态强制层,模型在决策前需将输入分为事实、启发式和情感三个认知角色。CKM本身不增加能力,而是迫使模型在行动前跟踪所使用的信息类型。通过对来自四个供应商的26个LLMs在韩语决策场景(歧义、伦理冲突、资源分配、错误处理)下的37403次观察进行评估,包括四个核心实验、一个4臂消融实验、一个5臂假限制消融实验和一个温度探测实验。结果表明CKM能降低重复输出的变异性,减少决策翻转率,其效果不仅是JSON格式化,在固定锚定状态下内在随机性可忽略,在采样随机性下优势更明显,且假消融实验表明其增益部分归因于结构支架和内容。CKM虽未提高推理正确性,但证明行为一致性可测量、因模型而异且可通过强制模型在决策前区分事实、假设和评估信号来部分改善。

英文摘要

Large language models (LLMs) can give different answers to the same decision problem across runs, and reverse a decision when their own prior answer returns as context. We ask whether this instability can be measured and partially reduced without changing model weights. We test the Cognitive Kernel Model (CKM), a prompt-level state-enforcement layer. Before deciding, the model must separate its input into three epistemic roles: Fact (given or verifiable), Heuristic (inferred or assumed), and Emotion (evaluative or priority signal). CKM adds no capability; it forces the model to track what kind of information it uses before acting. Formally it maintains a structured state S_t = {F_t, H_t, E_t} updated by a transition function. We evaluate CKM on Korean-language decision scenarios (ambiguity, ethical conflict, resource allocation, error handling) across 26 LLMs from four vendors and 37,403 observations, via four core experiments, a 4-arm ablation, a 5-arm sham-restriction ablation, and a temperature probe. Findings: (1) CKM reduces repeated-output variability (random-effects Hedges' g=1.09, 95% CI [0.83, 1.35], 31 model pairs); (2) state persistence cuts the decision-flip rate by 82% in newer models (g=1.52); (3) the effect is not JSON formatting alone (value-only recomputation, g=2.24); (4) intrinsic randomness under fixed anchor states is negligible; (5) the advantage grows under sampling stochasticity (g=2.87 at temperature 0.7); (6) a sham ablation attributes about 45% of the gain to structural scaffolding and 55% to Fact/Heuristic/Emotion content, and CKM is the only arm that both raises consistency and reduces flipping. CKM does not improve reasoning correctness. The narrower result: behavioral consistency is measurable, varies across models, and is partially improvable by forcing models to separate facts, assumptions, and evaluative signals before deciding.

发表机构

  • School of Mechanical Engineering, Chung-Ang University(韩国中央大学机械工程学院)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑