arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

决策劫持:针对Jev类型化概率决策的提示注入攻击

Decision Hijacking: Prompt Injection Attacks on Jev's Typed Probabilistic Decisions

Tiantong Wu, Wei Yang Bryan Lim

arXiv 2609.28613首次发表:更新:

发表机构

Nanyang Technological University(南洋理工大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

研究提示注入对非生成式决策模型Jev的影响,发现恶意内容改变动作概率但少致目标选择,自适应攻击提升成功率,表明模式化输出未消除风险。

AI 中文摘要

大多数关于提示注入的研究聚焦于生成式智能体,而对其在具有模式定义输出的模型上的影响尚不明确。我们在Jev这一非生成式决策模型中,使用510个重构的InjecAgent案例来考察这些影响。恶意内容会改变动作概率,但很少导致Jev选择攻击者的目标。覆盖标记(override markers)能减少这种影响,而上下文相关性声称的影响较小。利用分数反馈的自适应攻击使优化过程中发现的平均最高攻击者目标概率翻倍,而在新验证调用上的成功率从1.8%上升到3.5%。探索性分析将这些成功与较小的初始决策裕度或攻击者对观测的更大控制联系起来。综合来看,这些发现表明,模式定义的输出改变但并未消除提示注入风险,凸显了评估不可信内容如何影响允许动作集内选择的必要性。

英文摘要

Most studies of prompt injection focus on generative agents, leaving their effects on models with schema-defined outputs unclear. We examine these effects in Jev, a non-generative decision model, using 510 reconstructed InjecAgent cases. Malicious content shifts action probabilities but rarely causes Jev to select the attacker's target. Override markers reduce this influence, while claims of contextual relatedness have small effects. Adaptive attacks using score feedback double the mean highest attacker-target probability found during optimization, while success on fresh validation calls rises from 1.8% to 3.5%. Exploratory analysis links these successes to small initial decision margins or greater attacker control over the observation. Together, these findings show that schema-defined outputs change but do not eliminate prompt-injection risk, highlighting the need to evaluate how untrusted content influences choices within the allowed action set.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑