arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

关于循环语言模型的思维链可监控性

On the Chain-of-Thought Monitorability of Looped Language Models

Han Wang, Ishwar B Balappanawar, Huan Zhang

arXiv 2610.02741首次发表:更新:

发表机构

University of Illinois Urbana-Champaign(伊利诺伊大学厄巴纳-香槟分校)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本研究首次系统评估循环语言模型的思维链可监控性,发现深层循环在压力测试下可降低特定任务的可监控性,但架构本身不必然导致更低监控性。

AI 中文摘要

思维链(CoT)监控为检测不良模型行为提供了一种有前景的方法。循环语言模型(LoopLMs)重复应用共享的Transformer层,增加了有效计算深度,并在不增加模型规模的情况下实现了额外的潜在计算。然而,循环架构对CoT可监控性的影响在很大程度上仍未得到探索。在这项工作中,我们首次对LoopLMs中的CoT可监控性进行了系统评估。我们研究了两个互补的设置:(1)在同一LoopLM家族内改变循环深度,以隔离额外循环计算的影响;(2)将LoopLMs与按参数规模、Transformer层数或有效深度匹配的非循环语言模型进行比较,以研究LoopLMs是否更不易监控。在来自MonitorBench的八个任务以及标准和压力测试设置中,我们观察到在特定Logic/Science/Engineering \ exttt{Cue Answer}任务的压力测试下,CoT可监控性出现任务依赖性的下降,而其他任务则表现出较弱或性质不同的趋势。我们的诊断表明,这些下降不能完全由任务难度、验证通过率或生成的令牌长度来解释;定性示例进一步表明,更深循环模型如何明确使用或归因所提供的线索发生了变化。我们的跨模型比较发现,没有证据表明LoopLMs在按规模或深度匹配的非循环语言模型上系统性更不易监控。总体而言,我们的结果表明,在压力测试下,更深的循环深度可以降低某些任务中的CoT可监控性,但仅循环Transformer架构本身并不一定意味着较低的可监控性。

英文摘要

Chain-of-thought (CoT) monitoring provides a promising approach for detecting undesirable model behavior. Looped language models (LoopLMs) repeatedly apply shared transformer layers, increasing effective computational depth and enabling additional latent computation without increasing model size. However, the effect of looped architectures on CoT monitorability remains largely unexplored. In this work, we provide the first systematic evaluation of CoT monitorability in LoopLMs. We study two complementary settings: (1) varying the loop depth within the same LoopLM family to isolate the effect of additional recurrent computation, and (2) comparing LoopLMs with non-looped language models matched by parameter size, transformer-layer count, or effective depth to study whether LoopLMs are less monitorable. Across eight tasks from MonitorBench and both standard and stress-test settings, we observe task-dependent reductions in CoT monitorability under stress tests on specific Logic/Science/Engineering \texttt{Cue Answer} tasks, while other tasks exhibit weaker or qualitatively different trends. Our diagnosis suggests that these declines are not fully explained by task difficulty, verification pass rate, or generated token length; qualitative examples further suggest changes in how deeper-loop models explicitly use or attribute provided cues. Our cross-model comparison finds no evidence that LoopLMs are systematically less monitorable than non-looped language models matched on size or depth. Overall, our results suggest that deeper loop depth can reduce CoT monitorability in some tasks under stress tests, but looped transformer architecture alone does not necessarily imply lower monitorability.

CommentsPreprint

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑