发表机构
University of Technology Sydney; City University of Macau(悉尼科技大学; 澳门城市大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本研究首次系统评估Jev作为应用决策层的安全与隐私风险,通过提示注入、数据投毒及推理攻击揭示其脆弱性,并探讨其决策能力在防御与恶意场景中的双重用途。
AI 中文摘要
Jev将自然语言问题转化为类型化答案和概率,具有低延迟和低成本的特点,使应用程序能够路由请求并选择工具。虽然这一接口使Jev能够作为决策层自然融入应用程序工作流,但这种新兴用途的安全和隐私影响在很大程度上尚未被探索。为弥补这一空白,我们使用官方Jev API和NanoJev(一种具有可控训练数据和更新的本地模型)进行了首次系统性研究,聚焦三个研究问题:(1)当Jev作为应用程序决策层部署时,会产生哪些安全威胁?(2)尽管返回受限的类型化输出,Jev能泄露哪些隐私信息?(3)Jev的通用决策能力如何被用于有益目的或被滥用?Jev的决策依赖于应用程序状态,并可能受到用户提供输入的影响。因此,我们调整了提示注入和对抗性后缀来操纵其决策。开源Jev的分发和更新引入了供应链风险,我们通过在NanoJev中植入后门(通过训练数据投毒)来检验这些风险。由于Jev的输出既反映应用程序状态,也反映训练期间学到的信息,我们进一步调整了成员推理、私有属性推理和内部知识推理攻击,以在其受限输出格式下恢复敏感信息。最后,Jev可作为防御性和恶意工作流的通用决策预言机。我们通过四项检测任务(涵盖提示注入、越狱输入、有害内容和AI生成文本)以及涉及越狱和模型提取的滥用场景来检验这种双重用途。我们的实证评估表明,Jev对所检验的安全和隐私威胁仍然脆弱,而其决策能力可支持有益和恶意用途。
英文摘要
Jev turns natural-language questions into typed answers and probabilities with low latency and cost, enabling applications to route requests and select tools. While this interface allows Jev to integrate naturally into application workflows as a decision layer, the security and privacy implications of this emerging use remain largely unexplored. To address this gap, we conduct the first systematic study of these implications using the official Jev API and NanoJev, a local model with controllable training data and updates, focusing on three research questions: (1) What security threats arise when Jev is deployed as an application decision layer? (2) What private information can Jev reveal despite returning constrained typed outputs? (3) How can Jev's general-purpose decision capability be used for beneficial purposes or misused? Jev's decisions depend on application state and may be influenced by user-provided inputs. We therefore adapt prompt injection and adversarial suffixes to manipulate its decisions. Open-source Jev distribution and updates introduce supply-chain risks, which we examine by implanting backdoors in NanoJev through training data poisoning. Since Jev's outputs reflect both application state and information learned during training, we further adapt membership, private attribute, and internal knowledge inference attacks to recover sensitive information despite its constrained output format. Finally, Jev can serve as a general-purpose decision oracle for defensive and malicious workflows. We examine this dual use through four detection tasks covering prompt injection, jailbreak inputs, harmful content, and AI-generated text, alongside misuse scenarios involving jailbreak and model extraction. Our empirical evaluation shows that Jev remains vulnerable to the examined security and privacy threats, while its decision capability can support beneficial and malicious uses.
Comments16 pages, 8 figures, 6 tables; The source code is available at \url{https://github.com/shihe98/Security_Privacy_Jev}