arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

相似度门控认可反转:智能体系统中嵌入余弦阈值的有效性审计

Similarity Gates Approve Reversals: A Validity Audit of Embedding-Cosine Thresholds in Agent Systems

Scott E. Frias

arXiv 2608.10216首次发表:更新:

发表机构

Eigenforma; Freemind Labs(艾根福马; 自由思维实验室)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

该研究审计智能体系统中嵌入余弦阈值门控的有效性,发现其无法区分指令反转与释义、性能差且易受混杂因素影响,多数修复措施失效,仅部分配置可区分反转与释义。

AI 中文摘要

智能体框架配备了质量门控,这些门控通过嵌入余弦相似度比较文本块,并在固定阈值处做出决策。去重过滤器、语义缓存、漂移防护器和答案评分器门控用于回答“该文本是否仍然表达相同含义?”的问题,但该评分回答的是另一个问题:“措辞改变了多少?”。我们将这类门控作为测量工具进行审计。在这些门控旨在检测的案例中,两个问题的答案可能相反。很多时候,反转指令只需修改一个单词,而同意(即保持原意)往往需要改写整个句子。其结果是安全检查会反向触发。我们审计的生产漂移防护器在56个破坏含义的突变中,0个被检测到,而一个被认可的项目“ withhold the study drug”( withhold研究药物)变为“administer the study drug”(给予研究药物),其余弦相似度为0.9608。我们观察到五个已部署的操作点,在90个配置-阈值-任务单元中的平衡准确率从未超过0.700(中位数为0.525)。同样的混杂因素也破坏了评估:一个朴素构建的语料库继承了该混杂因素,可能返回相反的判断,在18个配置-任务单元中,有13个的决策AUROC恰好为0.000(所有18个单元中最高为0.040),而在平衡的2×2设计下,相同的9个配置的AUROC为0.440-0.815。这项工作中有两次,该门控捕获了我们自己的标题声明。明显的修复措施失效:编码器交换和重叠条件门控(样本内准确率0.750,保留集准确率0.533)在独立作者编写的保留集数据上表现接近随机,而NLI替代方案也没有更好的效果。嵌入在这方面仍有希望,因为9个配置中最强的两个在匹配重叠度下区分了反转和释义(AUROC为0.79-0.90),但只有配对审计才能揭示部署机制。我们发布了语料库方法、工具链和冻结结果,并认为以此方式门控的评分测量的是错误的东西,我们相信可以构建出有效的测量工具。

英文摘要

Agent frameworks ship quality gates that compare text blocks by embedding-cosine similarity and decide at a fixed cutoff. Deduplication filters, semantic caches, drift guards, and answer grader gates deploy to answer the question: "Does this text still mean the same thing?" But the score answers a different question: "How much did the wording change?" We audit this gate class as a measurement instrument. In the cases these gates exist to catch, the two can run in opposite ways. Many times, reversing an instruction is a single word edit, while agreement often rephrases a sentence. The consequence is a safety check that fires backwards. The production drift guard we audited caught 0 of 56 meaning-breaking mutations, and one approved item, "withhold the study drug" -> "administer the study drug", came in at cosine 0.9608. We observed five shipped operating points, and balanced accuracy across 90 configuration-threshold-task cells never exceeded 0.700 (median 0.525). The same confounder also corrupted evaluations. A naively built corpus inherits this confounder and can return an inverted verdict, with a decision AUROC exactly 0.000 in 13 of 18 configuration-task cells (at most 0.040 in all 18) against 0.440-0.815 for the same nine configurations under a balanced 2x2 design. Twice in the effort it captured our own headline claims. Obvious repairs fail: an encoder swap and an overlap-conditioned gate (0.750 in-sample, 0.533 held-out) land at chance on separately authored held-out data, and an NLI drop-in did no better. Embeddings do still bear hope here, as the strongest two of nine configurations separated reversal from paraphrase at matched overlap (AUROC 0.79-0.90), but only a matched-pair audit reveals the deployment regime. We release the corpus method, harness, and frozen results, and contend that scores gated this way measure the wrong thing. We believe a valid instrument is buildable.

Comments11 pages, 2 figures. Artifact: https://github.com/eigenforma/polaritycheck (DOI: 10.5281/zenodo.21796531)

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑