对话式AI中的熵:结构化不可预测性作为可推断的内在性
Entropy in Conversational AI: Structured Unpredictability as Inferrable Interiority
浏览论文内容
中文总结 AI 辅助
本研究提出结构化不可预测性作为对话AI设计目标,通过选择层更新隐藏状态并选择响应,实验表明该方法提升词汇新颖性但未建立路径依赖,且未测试感知心智主张。
中文摘要 AI 辅助
采样可以增加响应多样性,但不会产生依赖历史的行为。我们将一个不同的设计目标——结构化不可预测性——形式化为输出与一个持久的隐藏状态之间的条件依赖,这种依赖超出了观察者从对话记录中能推断的范围。一个选择层从容量有限的流中更新一个低维的风格与注意力状态,使用固定的基础模型生成多个响应,并选择新颖性和状态亲和性。评估使用独立的提示轮次脚本序列:基础模型接收当前轮次和渲染后的状态,但不接收之前的对话;跨轮次的依赖存在于包装器状态和响应选择器中。一个合成实现验证了流程,并在点级别匹配了四个前瞻性哈希冻结的散度特征。在最终的真实模型网格中(mlx-community/Qwen2.5-1.5B-Instruct-4bit;每臂56个序列),该机制相对于低方差和仅一致性对照组分别将词汇新颖性提高了0.073和0.023。其与新颖性匹配采样的文体一致性对比在注册的最小效应规则下等效为零,因此联合新颖性-一致性标准未通过。最初的两部分累积标准也未通过;在功效网格后冻结的修订最终网格对比发现,其一致性高于记忆重置消融(0.028,95%置信区间[0.018,0.039]),但未建立路径依赖。孪生分离未建立(0.003,95%置信区间[-0.011,0.019]);平均曲线的饱和曲率与冻结预测匹配,但在没有分离的情况下不支持路径依赖。探针级能力等效性在接近上限的测试组上保持在±0.10以内,而输出质量未评估。所有结果均为机器评分;未测试关于感知心智或意识的任何主张。
英文摘要
Sampling can increase response diversity without producing history-dependent behavior. We formalize a different design target, structured unpredictability, as conditional dependence between an output and a persistent hidden state beyond what an observer can infer from the transcript. A selection layer updates a low-dimensional style-and-attention state from a capacity-limited stream, generates several responses with a fixed base model, and selects for novelty and state affinity. Evaluation uses scripted sequences of independent prompt turns: the base model receives the current turn and rendered state, but not the preceding dialogue; cross-turn dependence resides in the wrapper state and response selector. A synthetic implementation validates the pipeline and matches four prospectively hash-frozen divergence features at point level. In the final real-model grid (mlx-community/Qwen2.5-1.5B-Instruct-4bit; 56 sequences per arm), the mechanism increased lexical novelty over the low-variance and consistency-only controls by 0.073 and 0.023, respectively. Its stylometric-consistency contrast with novelty-matched sampling was equivalent to zero under the registered smallest-effect rule, so the joint novelty-consistency criterion failed. The original two-part accumulation criterion also failed; a revised final-grid contrast, frozen after the powered grid, found higher consistency than the memory-reset ablation (0.028, 95% CI [0.018,0.039]), but does not establish path dependence. Twin separation was not established (0.003, 95% CI [-0.011,0.019]); the mean curve's saturating curvature matched the frozen prediction, which without separation does not support path dependence. Probe-level capability equivalence held within +/-0.10 on a near-ceiling battery, while output quality was not evaluated. All outcomes are machine-scored; no claims about perceived mind or consciousness are tested.
发表机构
- University of Bucharest(布加勒斯特大学)
机构由 AI 辅助整理,请以论文原文为准。