前瞻性解释风险:大语言模型之间的原则性通信控制
Prospective Interpretation Risk: Principled Communication Control Between LLMs
- University of Liverpool(利物浦大学)
- Queen Mary University of London(伦敦玛丽女王大学)
- University College London(伦敦大学学院)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
本研究提出前瞻性解释风险(PIR)框架,通过黑盒探针估计接收者解释偏差,并引入解释信息价值(VoII)指导消息修订,显著降低异构LLM系统中的解释失败率。
AI中文摘要:
大语言模型(LLM)智能体系统越来越依赖模型之间的相互通信,然而现有的不确定性和多智能体方法很少在消息发送之前估计特定接收者将如何解释该消息。这在异构系统中尤为重要,因为在该类系统中,有能力的接收者可能从同一条消息中重构出不同的任务。我们将此建模为具有潜在接收者类型的发送者-接收者问题,并定义了前瞻性解释风险(PIR):即接收者重构出非预期任务的概率。我们不对LLM的完整输入-输出行为进行建模,而是使用黑盒探针来关联消息、预期任务和接收者特定的重构,从而在将解释与下游能力故障分离的同时实现可扩展的监督。在离线阶段,异构的冻结接收者为接收者条件下的风险和预定义可变消息特征的效果提供监督。在部署时,历史记录会诱导接收者类型的后验分布,从而指导消息的修订和选择。我们引入了解释信息价值(VoII),仅当预期的通信收益超过其成本时才查询接收者信息。我们的理论刻画了接收者信息何时具有决策价值,并限制了此类查询。实验上,解释失败率在不同接收者之间变化4-13倍。接收者信息将PIR校准误差相对于不区分接收者的预测器降低了68%,这主要通过纠正接收者特定的风险水平来实现。PIR引导的修订将解释失败率相对于原始消息降低了44%,相对于通用重写降低了40%,这主要得益于一种对所有接收者都有帮助的修复。VoII在其优化的解释目标上,在匹配成本下优于信息增益和随机查询,将解释失败率从3.84%降至3.79%,同时仅查询了18.2%的回合。
英文摘要:
Large language model (LLM) agentic systems increasingly rely on models communicating with one another, yet existing uncertainty and multi-agent methods rarely estimate how a particular receiver will interpret a message before it is sent. This matters in heterogeneous systems, where capable receivers can reconstruct different tasks from the same message. We model this as a sender-receiver problem with a latent receiver type and define prospective interpretation risk (PIR): the probability that a receiver reconstructs a task other than intended. Rather than model an LLM's full input-output behaviour, we use black-box probes relating messages, intended tasks, and receiver-specific reconstructions, yielding scalable supervision while separating interpretation from downstream capability failure. Offline, heterogeneous frozen receivers provide supervision for receiver-conditioned risk and the effects of predefined mutable message features. At deployment, history induces a posterior over receiver types, guiding message revision and selection. We introduce value of interpretation information (VoII), querying for receiver information only when its expected communication benefit exceeds its cost. Our theory characterises when receiver information has decision value and bounds such queries. Empirically, interpretation-failure rates vary by 4-13x across receivers. Receiver information reduces PIR calibration error by 68% relative to a receiver-agnostic predictor, largely by correcting receiver-specific risk levels. PIR-guided revision reduces interpretation failure by 44% relative to the original message and 40% relative to a generic rewrite, mostly through a repair that helps every receiver. VoII outperforms information-gain and random querying at matched cost on the interpretation objective it optimises, lowering interpretation failure from 3.84% to 3.79% while querying 18.2% of episodes.