通过LLM路由元数据泄露敏感主题:测量与缓解
Sensitive-Topic Leakage Through LLM Routing Metadata: Measurement and Mitigation
浏览论文内容
中文总结 AI 辅助
研究LLM路由元数据泄露敏感主题的隐私风险,通过大规模真实请求测量,提出缓解措施并评估其效果。
中文摘要 AI 辅助
LLM路由器根据每个请求的内容选择廉价或昂贵的模型,许多网关和一些云平台可以在内容日志关闭的情况下记录该选择。我们测量了这一超出令牌计数的隐私通道,考虑了噪声标签和重复提示。我们在170万个真实请求(WildChat-1M,LMSYS-Chat-1M)上运行了预注册研究,使用两个成本/质量路由器和域路由器,调查了11个系统的日志记录,并测试了后处理防御。在匹配长度下,偏移方向取决于类别和路由器。对于RouteLLM在50%操作点,骚扰和自我伤害请求在未见过的提示上比同类请求到达强模型的频率低19个百分点,医疗请求(探索性:LLM标签未通过其门控)在不同提示上低31个百分点(均为事后分析),性相关请求高10个百分点(次要);另一个路由器的四个均为负值。20个RouteLLM决策区分频繁医疗提问者,AUC为0.71,探索性且低于预注册主要终点的0.75(域路由器:0.92,上限估计)。按类别长度匹配的准确标签奇偶校验消除了真实流量上的差距,在RouterBench上最多花费0.2个准确率点(事后分析),其中路由器在敏感主题上的差距(13-42个百分点,预注册)超过了通过实现准确率增益的Oracle路由(1-11,事后分析)。按对话粘性、按用户预算带和池化奇偶校验均失败,后者因为类别的偏移在大小或符号上不同。事后精确按用户率仅隐藏偶数前缀强计数,并放弃了大部分自我评估的路由价值;它保留了奇数位置决策,从这些决策中,事后日志攻击在20个RouteLLM请求后达到AUC 0.73(探索性)。
英文摘要
LLM routers pick a cheap or expensive model per request by its content, and many gateways and some cloud platforms can log that choice with content logging off. We measure this privacy channel beyond token counts, accounting for noisy labels and repeated prompts. We run pre-registered studies on 1.7 million real requests (WildChat-1M, LMSYS-Chat-1M) with two cost/quality routers and a domain router, survey eleven systems' logging, and test post-processing defenses. At matched length, the shift's direction depends on category and router. For RouteLLM at the 50% operating point, harassment and self-harm requests reach the strong model 19 points less often than comparable ones on prompts unseen in exploration, medical requests (exploratory: LLM labels failed their gate) 31 points less often on distinct prompts (both post hoc), and sexual requests 10 points more often (secondary); the other router's four are negative. Twenty RouteLLM decisions separate frequent medical askers with AUC 0.71, exploratory and below the pre-registered primary endpoint's 0.75 (domain router: 0.92, an upper estimate). Per-category length-matched parity with accurate labels removes the gap on real traffic, costing at most 0.2 accuracy points on RouterBench (post hoc), where routers' gaps on sensitive subjects (13-42 points, pre-registered) exceed those of an oracle routing by realized accuracy gain (1-11, post hoc). Per-conversation stickiness, per-user budget bands, and pooled parity fail, the last as categories' shifts differ in size or sign. A post hoc exact per-user rate hides only even-prefix strong counts and forfeits most self-assessed routing value; it preserves odd-position decisions, from which a post hoc log attack reaches AUC 0.73 after 20 RouteLLM requests (exploratory).
发表机构
- Krixvon
机构由 AI 辅助整理,请以论文原文为准。