arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

市场信号注入:对LLM定价代理的对抗性上下文操纵

Market Signal Injection: Adversarial Context Manipulation of LLM Pricing Agents

Dohun Lee, Hyunwoo Park

arXiv 2609.18357首次发表:更新:

发表机构

Seoul National University(首尔大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本研究提出市场信号注入攻击,通过操纵数据呈现方式影响LLM定价代理行为,实验显示情绪攻击效果显著,并建议输入规范化等防御措施。

AI 中文摘要

大语言模型(LLM)定价代理可能会对市场数据的呈现方式做出响应,即使其数值保持不变。我们引入了市场信号注入(MSI),这是一种在不发出明确指令的情况下,操纵数字格式、竞争对手排序或定性市场评论的攻击方法。我们在模拟的伯特兰双寡头和三家寡头市场中评估了九个开放权重模型,并在双寡头市场中评估了三个专有模型。基于情绪的攻攻击产生了最大的行为转变,这些转变会传播到其他公司并改变利润和消费者剩余。易感性因模型系列而异,较大的模型并不总是更稳健。在固定需求参数的模拟下,匹配的中性文本对照和基于规则的代理支持基于框架的转变解释。在全部十一个重新评估的模型-条件对中,逐集保留的探针区分了基线和受攻击的激活:线性AUC为1.00,MLP AUC范围为0.93至0.99。这种可分离性本身并不能识别有害的定价决策。输入规范化消除了测试的情绪攻击,而决策边界锚定(结合了提示约束和输出投影)在测试的自适应攻击下提供了部分缓解。这些结果将数据呈现确定为LLM定价代理的攻击面,并促使防御措施考虑代理之间的交互。

英文摘要

Large language model (LLM) pricing agents may respond to how market data is presented, even when its numerical values remain unchanged. We introduce market signal injection (MSI), an attack that manipulates numerical formatting, competitor ordering, or qualitative market commentary without issuing explicit instructions. We evaluate nine open-weight models in simulated Bertrand duopoly and triopoly markets and three proprietary models in duopoly markets. Sentiment-based attacks produce the largest behavioral shifts, which propagate to other firms and alter profits and consumer surplus. Susceptibility varies across model families, and larger models are not consistently more robust. Matched neutral-text controls and a rule-based agent support a framing-based account of these shifts under the fixed demand parameters of our simulation. Episode-held-out probes distinguish baseline from attacked activations in all eleven re-evaluated model--condition pairs: linear AUC is 1.00 and MLP AUC ranges from 0.93 to 0.99. This separability does not by itself identify harmful pricing decisions. Input canonicalization removes the tested sentiment attacks, while decision boundary anchoring, which combines prompt constraints with output projection, provides partial mitigation under the tested adaptive attacks. These results identify data presentation as an attack surface for LLM pricing agents and motivate defenses that account for interactions among agents.

Comments30 pages, Accepted to FinNLP 2026 Workshop @ EMNLP 2026

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑