arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

创造力、诚实和设计遗忘在小型双曲语言模型中出现

Exposure is not manifestation: measurement target and output resolution jointly determine which behavioural-faithfulness evaluator wins

Kwan Soo Shin

arXiv 2607.09306首次发表:更新:

发表机构

PolymathMinds Lab; POSTECH; aSSIST University; Korean Educational Development Institute(多智思维实验室; 浦项科技大学; aSSIST大学; 韩国教育开发院)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

研究探讨小型双曲语言模型如何成为可信陪伴式AI,通过行为审计器检测合规差距,用线性读出检测相关问题,创意框架播种器更优,记忆操作系统实现设计遗忘,为可信陪伴式AI提供了小模型路径。

AI 中文摘要

语言模型虽在规模上得到优化,但作为助手向陪伴者转变时存在问题,人类评分者对此看法不一。本文展示了三个共享双曲基础的小型语言模型能回答相关问题。一个146M的行为审计器能检测评分者无法察觉的合规差距,其冻结表示的线性读出可检测陪伴者引发的谄媚等问题。创意框架播种器在比较中更受青睐,记忆操作系统实现了设计遗忘。创造力、诚实和设计遗忘构成了通往可信陪伴式人工智能的小模型路径。

英文摘要

Behavioural auditing asks whether a language model behaves as it claims, but detection scores are reported without separating two targets: whether a reply was produced under a behaviour-inducing condition (exposure) and whether the behaviour surfaced in it (manifestation). Scoring a compact 146-million-parameter auditor's frozen-representation read-out and a frontier judge against each label on the identical 720 replies, the gap between the instruments moves by roughly 0.2 AUROC when the target changes. Under the judge's deployed interface, a single verdict, the ranking reverses: the auditor leads on exposure, 0.804 against 0.718, and trails on manifestation, 0.690 against 0.811. Matching the output resolution from either direction, by asking the judge a target-specific question answered with a continuous confidence score or by thresholding the auditor's read-out, removes the reversal but not the interaction, which excludes zero at all three resolutions (0.207, 0.237 and 0.169). The target governs how far apart the instruments are; the interface governs whether that distance changes their order. The auditor's hyperbolic geometry confers no advantage here. A single behavioural-detection AUROC is under-specified: such claims are comparable only when they state the estimand, the evaluator, and its output interface.

CommentsSubstantially revised and narrowed version with a new title and estimand-centred analysis. Comparisons are now reported at three output resolutions, and the reproducibility package has been rebuilt. The author list was changed with the approval of all authors listed on v1-v2; previous versions remain publicly available. 17 pages, 3 figures, 3 tables

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑