超越轨迹:将可解释推理状态读出与原生MoE路由耦合
Beyond the Trace: Coupling an Interpretable Reasoning-State Readout to Native MoE Routing
查看机构详情
- Fudan University(复旦大学)
- Shanghai Innovation Institute(上海创新研究院)
机构由 AI 辅助整理,请以论文原文为准。
浏览论文内容
中文总结 AI 辅助
本研究提出将可解释推理状态读出J64与原生MoE路由耦合的两级机制,提升推理过程可解释性与性能,R64作为低开销代理可保留其核心增益,还能优化分支选择与投票策略、调整推理行为。
中文摘要 AI 辅助
推理模型的输出仅为产生该输出过程的部分记录。本文针对混合专家(mixture-of-experts,MoE)推理引入两级内部读出机制:首先将词汇级J空间提炼为J64,这是从模型自身推理状态中学习得到的64轴语义框架,它能揭示已发出轨迹未呈现的可读过程状态,将推理 effort 与问题引发的压力区分开来;相较于以 token 占用率读取相同rollout并按相同方式聚合的基线方法,该方法在保留的AUC指标上提升了0.096至0.135。随后从原生专家路由统计数据中重构J64,得到低开销代理R64:在三个模型及两个系列中,其与J64的轴间中位数相关性为0.69至0.86,在gpt-oss-20b模型上可保留J64 95%至100%的预测增益。该读出机制支持两种时间分辨率的测试时决策:在完整候选集上,J64和R64可改进单分支选择,R64加权投票在8个设置中的7个场景下优于普通多数投票;生成过程中,滚动读出窗口驱动累积停止-重采样策略,其工作点仅基于训练问题确定,J64相较于同 permutation 控制组提升准确率1.1至5.9个百分点,仅路由的R64代理保留其中0.9至3.2个百分点。最后,针对J64命名机制的路由器编辑会诱导预期的推理行为,将诊断出的停滞从数值猜测转向精确符号执行。综上,J64使潜在过程状态可读,而路由使其可部署且可操作。
英文摘要
What a reasoning model writes is only a partial record of the process that produces it. We introduce a two-level internal readout for mixture-of-experts reasoning. We first distill vocabulary-scale J-space into J64, a 64-axis semantic frame learned from the model's own reasoning states. J64 reveals readable process state that the emitted trace does not show: it separates inference effort from problem-induced strain. It also adds 0.096 to 0.135 held-out AUC over a baseline that reads the same rollout as token occupancy and aggregates it in exactly the same way. We then reconstruct J64 from native expert-routing statistics. The result is R64, a low-overhead proxy: its median per-axis correlation with J64 is 0.69 to 0.86 across three models and two families, and on gpt-oss-20b it preserves 95 to 100% of J64's predictive gain. The readout supports test-time decisions at two temporal resolutions. Over completed candidate sets, J64 and R64 improve single-branch selection, and R64-weighted voting improves plain majority voting in seven of eight settings. During generation, rolling readout windows drive a cumulative stop-and-resample policy whose operating point is fixed on training questions alone. J64 improves accuracy by 1.1 to 5.9 points over a sibling-permuted control, and the routing-only R64 proxy retains 0.9 to 3.2 of those points. Finally, router edits aimed at the mechanism J64 names induce the predicted reasoning behaviors and shift a diagnosed stall from numerical guessing toward exact symbolic execution. Together, J64 makes latent process state readable, while routing makes it deployable and actionable.