arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

已编码但不可执行:审计冻结大语言模型中几何约束的解码-生成-引导差距

Encoded but Not Actionable: Auditing the Decode-Generate-Steer Gap in Frozen LLMs for Geometric Constraints

Man Liang, Xinzhao Cheng, Faizan Wajid

arXiv 2608.17843首次发表:更新:

AI 中文总结

该研究以参数化CAD约束为测试平台,审计6个冻结LLMs的解码-生成-引导差距,发现几何关系可解码性与生成、引导能力存在差异,区分了编码与表达控制失败。

AI 中文摘要

大语言模型(LLMs)在结构化推理任务上表现出强大性能,但它们编码的内容是否能指导模型行为仍不明确。我们通过几何推理研究该问题,采用参数化CAD约束作为受控测试平台,以区分局部成对关系与草图级约束状态。通过探测6个冻结的仅解码器LLMs的隐藏状态,我们检验了四个属性:线性可解码性、强制选择生成、激活级影响及行为可引导性。预训练大幅提升了局部几何关系的解码能力,在通过打乱顺序控制考虑位置线索后,该优势依然存在。相比之下,草图级自由度(DOF)状态已可从随机初始化表示中高度解码,且仅随预训练适度提升,表明其大部分探测性能无需学习权重即可获得。进一步分析显示,可解码信息并非总能执行:生成常无法表达该信息,在两个经干预测试的主干模型中,修补实体位置的激活恢复效应消失,而可解码性在各层仍持续存在;均值差引导也无法可靠控制输出。这些结果表明,在测试场景中,可解码性、生成、激活级影响及可引导性可能存在差异。该审计提供了一种受控方式,以区分几何结构编码失败与编码信息的表达或控制失败。

英文摘要

Large language models (LLMs) have demonstrated strong performance on structured reasoning tasks, but what they encode and whether it informs model behavior remain unclear. We investigate this question through geometric reasoning, using parametric CAD constraints as a controlled testbed for separating local pairwise relations from sketch-level constraint status. By probing the hidden states of six frozen decoder-only LLMs, we examine four properties: linear decodability, forced-choice generation, activation-level influence, and behavioral steerability. Pretraining substantially improves the decoding of local geometric relations, and this advantage persists after accounting for positional cues with shuffled-order controls. In contrast, sketch-level DOF status is already highly decodable from randomly initialized representations and improves only modestly with pretraining, indicating that much of its probe performance is available without learned weights. Further analyses show that decodable information is not always actionable. Generation often fails to express this information, and on the two intervention-tested backbones, activation-restoration effects at the patched entity position vanish while decodability persists across depth. Mean-difference steering also does not reliably control outputs. These results show that decodability, generation, activation-level influence, and steerability can diverge in the tested setting. The audit provides a controlled way to distinguish failures to encode geometric structure from failures to express or control encoded information.

Comments13 pages, 7 figures, 8 tables, including appendices

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑