TRACE:基于合成奖励训练因果探索推理智能体
TRACE: Training Reasoning Agents for Causal Exploration with Synthesized Rewards
浏览论文内容
中文总结 AI 辅助
本文提出TRACE,通过模拟器注入干预生成合成奖励,将模糊的诊断推理转化为可扩展强化学习任务,在数字广告诊断中超越强基线。
中文摘要 AI 辅助
具有可验证奖励的强化学习(RLVR)推动了语言模型在数学和代码等领域的推理能力发展,在这些领域中,客观答案的验证成本较低。然而,对复杂数据进行诊断推理并不具备这一优势:确定异常的真实原因通常需要昂贵的专家调查,并且即使在事后也可能仍然存在歧义。我们提出疑问:这种验证的不对称性是否可以被设计出来。我们采样一个干预措施,将其注入一个受控模拟器中,并生成该干预措施所产生的观测数据。隐藏的干预措施提供了神谕标签和客观奖励,而智能体仍然必须调查带有噪声、混杂因素和分布式的证据。我们在TRACE中实现了这种方法,这是一个数字广告诊断环境,具有12个根本原因和细粒度的细分归因。智能体使用Python和SQL调查每个回合,并且必须识别根本原因以及(如果适用)受影响的细分分配。在一个保留的235回合测试集上,最强的提示基线Claude Opus 5达到了0.686的FullAttr@1。监督微调将Qwen3.5-35B-A3B从0.159提升到0.637,随后使用合成奖励的强化学习达到了0.757,超过了所有评估的提示基线,包括前沿闭源模型和提示的Qwen3.5-122B-A10B模型。所得策略使用的工具调用次数也远少于提示的35B基础模型。这些结果提供了证据,表明获得可扩展的客观训练信号可能比模型规模本身更为重要。更广泛地说,基于模拟的验证可以使原本模糊的诊断推理任务适用于可扩展的强化学习。
英文摘要
Reinforcement learning with verifiable rewards (RLVR) has advanced language-model reasoning in domains such as mathematics and code, where objective answers are inexpensive to check. Diagnostic reasoning over complex data lacks this advantage: establishing the true cause of an anomaly often requires costly expert investigation and may remain ambiguous after the fact. We ask whether this asymmetry of verification can instead be engineered. We sample an intervention, inject it into a controlled simulator, and generate the observations it would produce. The hidden intervention provides an oracle label and objective reward, while the agent must still investigate noisy, confounded, and distributed evidence. We instantiate this approach in TRACE, a digital-advertising diagnostic environment with 12 root causes and fine-grained segment attribution. Agents investigate each episode using Python and SQL and must identify both the root cause and, when applicable, the affected segment assignment. On a held-out 235-episode test set, the strongest prompted baseline, Claude Opus 5, reaches 0.686 FullAttr@1. Supervised fine-tuning raises Qwen3.5-35B-A3B from 0.159 to 0.637, and subsequent RL with synthesized rewards reaches 0.757, outperforming all evaluated prompted baselines, including frontier closed-source models and a prompted Qwen3.5-122B-A10B model. The resulting policy also uses substantially fewer tool calls than the prompted 35B base. These results provide evidence that access to a scalable, objective training signal can be a more important constraint than model scale alone. More broadly, simulation-based verification can make otherwise ambiguous diagnostic reasoning tasks amenable to scalable reinforcement learning.