arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

推理层安全:防御对抗性推理与基础设施滥用

Inference-Layer Security: Defending Against Adversarial Inference and Infrastructure Abuse

Keifer Lee

arXiv 2609.38239首次发表:更新:

发表机构

New York Machine Learning Research Guild(纽约机器学习研究协会)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本文提出在推理层防御LLM服务的对抗性攻击,通过结构因果模型生成标记数据集,训练梯度提升检测器,并发现检测器精度高于评估标签,同时改进攻击类型归因方法。

AI 中文摘要

技术报告:将大型语言模型(LLM)作为服务运营所需的不仅仅是推理基础设施:提供商还必须防御旨在利用该服务的对抗性交互,包括用于有害用途的越狱、复杂的拒绝服务攻击和蒸馏攻击。我们在推理层研究这个问题,使用一个假设的前沿实验室——五行公司(Five Elements Inc.)作为贯穿示例。由于不存在公开的对抗性LLM使用标记数据集,我们引入了一个结构因果模型(SCM),该模型生成一个基于现实、标记的用户会话数据集,其中包含协调的多账户活动、平台反馈和三个级别的标签可观测性。在此数据集上,我们训练了一个实用的梯度提升检测器,将每个用户会话分类为良性或恶意,如果恶意,则按攻击类型分类。针对神谕标签,该检测器几乎完美解决了二分类任务(AUPRC 0.993),但针对真实信任与安全团队会持有的操作标签,同一模型的AUPRC仅为0.313:检测器比用于评估它的标签更准确。对于攻击类型归因,朴素argmax被98%的良性先验所主导(宏F1 0.295),而一个简单的阈值决策引擎将宏F1提高到0.489,且不牺牲准确性。该数据集已公开发布。

英文摘要

A Technical Report: Operating a large language model (LLM) as a service requires more than inference infrastructure: the provider must also defend against adversarial interactions that seek to exploit the service, including jailbreaking for harmful use, sophisticated denial of service, and distillation attacks. We study this problem at the inference layer, using a hypothetical frontier lab, Five Elements Inc., as a running example. Because no public labelled dataset of adversarial LLM usage exists, we introduce a structural causal model (SCM) that generates a realistically grounded, labelled dataset of user-sessions, with coordinated multi-account campaigns, platform feedback, and three tiers of label observability. On this dataset we train a practical gradient-boosted detector that classifies each user-session as benign or malicious and, if malicious, by attack type. Against oracle labels the detector very nearly solves the binary task (AUPRC $0.993$), yet against the operational labels a real Trust & Safety team would hold, the same model scores an AUPRC of only $0.313$: the detector is more accurate than the labels used to evaluate it. For attack-type attribution, a naive argmax is dominated by the $98\%$ benign prior (macro-F1 $0.295$), whereas a simple thresholded decision engine raises macro-F1 to $0.489$ without sacrificing accuracy. The dataset is publicly released.

CommentsA Technical Report

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑