AI 中文总结
HoneyRoute通过流式路由器检测恶意请求并路由至蜜罐模型,结合分析循环生成攻击者指纹,以低延迟高精度保护生产模型,显著降低资源消耗并提升检测性能。
AI 中文摘要
我们介绍了HoneyRoute,一个推理服务层,用于检测传入请求是否为恶意请求,如果是,则将其路由到专用的蜜罐模型,在对抗者的交互被持续收集用于情报的同时保护生产环境。现有防御措施在模型内存中嵌入陷阱或在协议层重建欺骗,使得服务层不受保护,并且无法将任何信息反馈到检测中。HoneyRoute结合了(i)一个流式路由器(一个冻结的0.8B嵌入骨干网络,带有按域划分的MLP头),(ii)一个双实现蜜罐(一个基于规则/提示工程的代码蜜罐或一个专用的同族副本),以及(iii)一个分析循环,将捕获的交互转换为攻击者指纹,用于路由器重新训练。在生产轨迹加上七个域的攻击语料库上,路由器在38毫秒的中位额外延迟下达到F1=.911,匹配了两层防护LLM级联F1的96%,而其延迟仅为后者的1/385,在13种对抗性变换下实现了0%的逃避率;在并发洪泛攻击下,使用真实的GCG后缀负载,将恶意份额转移可将生产模型令牌消耗减少97.8%;训练后的副本在92.9%的良性保留请求上与生产模型一致,而朴素的无条件诱饵注入降至7.6%,选择性伪装注入恢复到88.9%,描绘了可恢复的保真度-可追溯性边界;一个循环训练校正头将合法安全研究的错误路由减少了9倍,同时将检测F1提高到.933。
英文摘要
We introduce HoneyRoute, an inference-serving layer that detects whether an incoming request is malicious and, if so, routes it to a dedicated honeypot model, shielding production while the adversary's interaction is continuously harvested for intelligence. Existing defenses embed traps inside model memory or rebuild deception at the protocol layer, leaving the serving tier unprotected and feeding nothing back into detection. HoneyRoute couples (i) a streaming router (a frozen 0.8B-embedding backbone with per-domain MLP heads), (ii) a dual-implementation honeypot (a rule/prompt-engineered code honeypot or a dedicated same-family replica), and (iii) an analysis loop that converts trapped interactions into attacker fingerprints for router retraining. On a production trace plus a seven-domain attack corpus, the router reaches F1=.911 at 38 ms median added latency, matching 96% of a two-tier guard-LLM cascade's F1 at 1/385 of its latency with 0% evasion under 13 adversarial transformations; diverting the malicious share cuts production-model token consumption under concurrent flooding with real GCG-suffix payloads by 97.8%; the trained replica agrees with the production model on 92.9% of benign holdout requests, while naive unconditional bait injection collapses to 7.6% and selective camouflaged injection recovers to 88.9%, mapping the recoverable fidelity-traceability frontier; and a loop-trained correction head cuts misrouting of legitimate security research 9x while raising detection F1 to .933.
CommentsPreprint. 14 pages, 4 figures, 1 table, 25 references