arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.28559cs.CRcs.AIcs.SE

谁在操控执行框架?通过智能体行为对大型语言模型进行指纹识别

Who Is Behind the Harness? Fingerprinting LLMs through Agentic Behavior

  • Tsinghua University(清华大学)

机构由 AI 辅助整理,请以论文原文为准。

Chuyi Wang, Xiaohui Xie, Tongze Wang, Fangchen Luo, Yong Cui

AI总结:

针对编码智能体中LLM身份识别难题,提出LIDAR主动黑盒指纹方法,通过探针对和双级特征,在36个模型上实现高准确率,证明执行行为可揭示模型身份。

AI中文摘要:

大型语言模型(LLM)越来越多地通过编码智能体执行框架(coding-agent harnesses)来运行,这些框架能够检查代码仓库、调用工具并修改文件。因此,替换此类智能体背后的模型可能会改变与安全相关的决策,包括是否验证更改或从故障中安全恢复。现有的LLM指纹识别方法主要从直接文本或令牌分布中推断模型身份。在编码智能体中,这些信号受到系统指令、控制器逻辑、工具和执行反馈的调节,限制了其可迁移性。我们提出了LIDAR(基于运行时决策与行为的LLM识别),一种针对编码智能体执行过程的主动黑盒指纹识别方法。三个编码探针对(coding probe pairs)在受控变化下分别暴露了编辑后验证、瞬时故障恢复和规范-测试冲突解决的行为。LIDAR利用互补的实例级和分布级特征来表示生成的轨迹,并通过轻量级概率标识符与干净参考进行比较。该方法无需访问模型权重、logits或提供商内部信息。在来自七个模型家族和两个智能体执行框架的36个模型上,LIDAR实现了高Top-1准确率和MRR,并优于四种现有的指纹识别和API审计基线。消融实验证实,两个特征级别、所有探针对及其受控变体均有贡献。这些结果表明,智能体执行行为提供了超越最终输出的模型身份证据。

英文摘要:

LLMs increasingly operate through coding-agent harnesses that inspect repositories, invoke tools, and modify files. Substituting the model behind such an agent can therefore change security-relevant decisions, including whether it verifies changes or recovers safely from failures. Existing LLM fingerprints largely infer identity from direct text or token distributions. In coding agents, these signals are mediated by system instructions, controller logic, tools, and execution feedback, limiting their transfer. We present LIDAR (LLM Identification from Decisions and Actions at Runtime), an active black-box fingerprinting method for coding-agent execution. Three coding probe pairs expose post-edit verification, transient-failure recovery, and specification--test conflict resolution under controlled changes. LIDAR represents the resulting trajectories with complementary instance-level and distribution-level features and compares them with clean references using a lightweight probabilistic identifier. It requires no access to model weights, logits, or provider internals. Across 36 models from seven families and two agent harnesses, LIDAR achieves high Top-1 accuracy and MRR and outperforms four existing fingerprinting and API-auditing baselines. Ablations confirm that the two feature levels, all probe pairs, and their controlled variants contribute. These results show that agent execution behavior provides model-identity evidence beyond final outputs.

↑