检测到效应不等同于学习基于该效应行动:LLM获取智能体的奖励信噪比下限
RAISE: Diagnosing Acquisition Collapse in Costly LLM Signals
浏览论文内容
中文总结 AI 辅助
该研究指出检测信号平均效应与学习基于该效应行动存在差异,提出奖励信噪比下限,引入SHE模型,在三个推荐数据集上发现学习获取策略失效,因数据集SNR低于下限。
中文摘要 AI 辅助
许多流程可对每个样本付出成本以获取模型衍生的辅助观测结果——比如大语言模型(LLM)的结构化推理、慢速预言机、昂贵测量值,之后必须判断何时该获取的信号值得使用。本文提出一个易被忽略的区分:检测这类信号平均有帮助,与学习针对每个实例基于该信号行动并非同一回事,且存在奖励信噪比(SNR)下限决定后者是否可行。即便信号是可信的,且选取实现奖励最高的b个样本的样本内预言机显示出可观的表观增益,任何可部署的策略都无法学习何时获取该信号:在单印象、聚类、 regime 及 uplift 树粒度下,学习到的路由策略从未优于随机策略,且匹配矩噪声安慰剂能重现预言机≥100%的表观增益——表观“可学习结构”实为噪声的顺序统计量。我们用两个概念解释这一现象:检测均值效应与学习每个实例的获取策略,以及奖励信噪比可检测下限:仅当奖励信噪比ρ超过ρ*(N)≈2.8/√N时,路由策略才可离线估计;正对照实验证实存在真实的低SNR限制而非流程故障。作为具体实例,我们引入结构化假设嵌入(SHE):一个冻结的LLM将用户历史转化为带排名、置信度评分、证据支撑的意图假设,融合进推荐系统。在三个公开数据集(MIND、REES46、Amazon-Beauty)上,SHE可信且可校准,但其价值依赖于 backbone 和 regime(相较于有序GRU有显著提升,+0.0114,95%置信区间[+0.0030, +0.0209],但全局冗余差距与零无差异),且在所有粒度下学习到的获取策略均失效,因为这三个数据集的SNR均低于下限。可实现的单元是设计时的regime门,而非每个实例的策略。我们发布了代码及一键复现脚本。
英文摘要
Large language models (LLMs) are increasingly used as costly, on-demand components in real systems, but calling them indiscriminately can waste substantial compute, latency, and serving budget. The key deployment question is therefore not only whether an LLM helps on average, but when it is worth calling. We identify a common failure mode, which we call acquisition collapse: an LLM signal can appear useful in aggregate or post hoc, yet still provide too little before-call information to support reliable selective use. We introduce RAISE (Reward-SNR Actionability in Signal Evaluation), a pre-routing diagnostic framework for testing whether available evidence supports selective use before committing to a routing strategy. We instantiate RAISE with Structured Hypothesis Embeddings (SHE), a frozen-LLM intent signal for recommendation using one LLM call per user, and evaluate it through controlled, retrospective, and fresh-cohort studies and a prospective offline pilot whose audit decisions are frozen before independent outcomes are revealed. Across these settings, predictable incremental benefit, not average lift alone, distinguishes settings with recoverable selective value; deployment additionally depends on cost and operational constraints. Seemingly strong oracle or subgroup gains can disappear under independent evaluation. More broadly, RAISE reframes costly inference as an information-acquisition problem: before paying for an expensive model, tool, sensor, or measurement, first test whether its value is predictable at decision time. This principle motivates cost-aware acquisition in settings ranging from agent tool use and stronger-model consultation to robotic sensing and clinical decision pipelines.