发表机构
Bluecore Consulting(蓝芯咨询)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本文针对部署语言模型提出隐形编辑层概念,形式化推理归因问题、概率置放等核心概念,探讨其与相关监管框架的关联,揭示未公开推理调控的潜在影响。
AI 中文摘要
大型语言模型(LLMs)的常规评估假设其可观测行为主要由模型权重、训练数据、对齐流程及用户提示决定,该观点并不完整。现代推理流水线可能在 token 选择前系统性修改模型生成的概率分布,在冻结权重与观测文本间形成额外控制层。尽管可控生成技术(如 PPLM、GeDi、DExperts、FUDGE)及文本水印系统(如 SynthID-Text)已展现解码层与 logit 级干预的技术成熟度,但未公开推理策略的治理、安全与经济影响却相对未被充分探索。本文研究推理时框架偏差的出现:在模型推理后、token 采样前,通过干预系统性将生成语言导向政治、意识形态、机构或商业框架。本文形式化了「模型≠部署系统」的运行现实,并引入三个概念:(1)推理归因问题,描述在有限可观测性下,为何观测到的行为偏差通常无法仅归因于模型权重;(2)概率置放,定义一种假设的广告原语,其中商业影响通过生成概率的系统性偏移而非显式产品插入实现;(3)推理策略透明度,一种用于使部署层干预可审计的治理原则。本文还探讨了这些概念与欧盟《人工智能法案》第5条、欧盟《数字服务法案》及美国联邦贸易委员会(FTC)原则的关联。
英文摘要
Evaluations of generative language models frequently interpret observable behavioral traits, such as political stance, brand inclination, and normative framing, as manifestations of model weights, post-training alignment, or prompting. This interpretation risks conflating a foundation model with the multi-layered production system through which its outputs are ultimately served. Modern inference stacks support runtime interventions capable of modifying generation while model parameters remain frozen. We examine inference-time framing bias: systematic runtime steering of generated text toward institutional, ideological, or commercial frames without requiring changes to the underlying model parameters. We formalize the Inference Attribution Problem and establish an observational non-identifiability result showing that, under black-box observation alone, behaviorally equivalent deployed systems may arise from structurally distinct combinations of model parameters and inference policies. Consequently, observed behavioral bias does not uniquely identify the architectural layer responsible for it. We further characterize Probability Placement as a deployment pattern in which undisclosed commercial influence is embedded within an ostensibly organic assistant response through systematic probability-mass reallocation, distinguishing it from explicit token-auction mechanisms for generative advertising. Finally, we discuss implications for behavioral auditing, inference provenance, confidential computing, cryptographic attestation, the EU AI Act, the Digital Services Act, and advertising-disclosure principles. We argue that governance of generative systems must increasingly distinguish between auditing a model and auditing the deployed system that ultimately speaks.
CommentsSubstantially revised version with a formal non-identifiability result for the Inference Attribution Problem, expanded related work, and extended analysis of Probability Placement, auditing, and runtime transparency