发表机构
University of Waterloo(滑铁卢大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
该研究提出Agentic形式主义陷阱与评估失调指数,通过多领域轨迹分析揭示作为评判者的大语言模型易受句法触发影响,且该漏洞与领域无关,需架构特定的警戒过滤器。
AI 中文摘要
我们提出了Agentic形式主义陷阱和评估失调指数($D_E$),量化作为评判者的大语言模型系统在对抗性负荷下如何将结构程序主义与语义真相混为一谈。对GAIA、SWE-bench、多挑战3个领域的22500条轨迹进行分析,我们提取了幻觉操作的语义分类法,经确定性词汇基础验证($p < 10^{-120}$)。逻辑元评估器分离出该评估器捕获的确切句法触发因素(ROC-AUC为0.8779),而零样本留一域外迁移证明该漏洞普遍与领域无关(平均ROC-AUC为0.7482)。架构分析显示,不同的模拟群体拓扑结构会引发数学上不同的语义盲点,证明无锚定的闭环评估不稳定、系统性发散,需要特定架构的警戒过滤器。
英文摘要
We introduce the \textit{Agentic Formalism Trap} and the Evaluative Dissonance Index ($D_E$), quantifying how LLM-as-a-Judge systems conflate structural proceduralism with semantic truth under adversarial load. Analyzing 22,500 trajectories across 3 domains (GAIA, SWE-bench, Multi-Challenge), we extract a semantic taxonomy of hallucination maneuvers, validated via deterministic lexical grounding ($p < 10^{-120}$). A logistic meta-evaluator isolates the exact syntactic triggers of this evaluator capture (ROC-AUC 0.8779), while a zero-shot Leave-One-Domain-Out transfer proves the vulnerability is universally domain-agnostic (mean ROC-AUC 0.7482). Architectural profiling reveals that distinct simulated swarm topologies induce mathematically disparate semantic blind spots, proving that unanchored closed-loop evaluation is unstable, systemically divergent and necessitates architecture-specific vigilance filters.