AI 中文总结
研究LLMs推理成本问题,提出RolePlay框架利用角色条件诱导低效行为放大推理成本,实验证明该框架优于现有方法,为LLM推理效率攻击提供新视角。
AI 中文摘要
大语言模型(LLMs)在现实世界应用中越来越多,但其自回归生成机制使恶意提示能操纵生成行为,导致过度令牌生成,增加计算消耗并威胁服务效率。现有方法有局限性。本文揭示LLMs中由角色一致性导致的未探索漏洞,基于此提出RolePlay框架,通过构建自适应角色诱导低效但语义连贯行为来放大推理成本。实验表明RolePlay优于现有方法,平均令牌放大达\(7.64\times\),最大放大率为\(207.64\times\),为LLM推理效率攻击提供新视角。
英文摘要
LLMs are increasingly deployed in real-world applications, making inference efficiency and service reliability critical concerns due to their substantial computational costs. However, the autoregressive generation mechanism of LLMs enables malicious prompts to manipulate generation behaviors, inducing excessive token generation that amplifies computational consumption and threatens service efficiency. Existing methods mainly rely on adversarial suffixes or explicit extension instructions, which introduce detectable behaviors and limit their applicability. In this paper, we reveal a previously unexplored vulnerability caused by persona consistency in LLMs, where models maintain assigned roles and reproduce corresponding behaviors even when they result in inefficient reasoning and excessive generation. Based on this observation, we propose RolePlay, a task-aware dynamic persona alignment framework that constructs adaptive personas to naturally induce inefficient yet semantically coherent behaviors for inference cost amplification. Extensive experiments across multiple LLMs and diverse task datasets demonstrate that RolePlay consistently outperforms existing inference extension methods, achieving an average token amplification of up to \bm{$7.64\times$} and a maximum token amplification ratio of \bm{$207.64\times$}. Our findings identify persona conditioning as a new attack surface for LLM inference efficiency and offer a new perspective on computational cost amplification.
Comments17pages