arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

因果世界模型何时帮助模块化LLM智能体

When Do Causal World Models Help Modular LLM Agents

Xinyuan Song, Zekun Cai

arXiv 2610.00012首次发表:更新:

发表机构

Emory University; The University of Tokyo; LocationMind(埃默里大学; 东京大学; LocationMind)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本研究通过FedCausalCompose框架探讨因果世界模型在模块化LLM智能体中的价值,发现当跨模块接口统计可识别且以行动时可用的形式呈现时,因果结构最有助于提升规划性能。

AI 中文摘要

LLM智能体日益通过模块化系统行动,例如订单、支付、库存和发货服务,其中一个模块中的动作会改变另一个模块中哪些转换是有效的。标准世界模型通常拟合观测轨迹,但这并非干预时规划所需的数量:一条轨迹可能显示支付先于发货,但无法识别支付是否授权发货、库存是否中介了该效应,或者是否存在一个隐藏触发因素解释两者。我们通过FedCausalCompose研究这一差距,这是一个用于模块化LLM智能体的因果世界模型框架,其中局部动作为跨模块接口提供干预-响应证据。我们首先证明,在未阻断的后门路径下,观测世界模型会产生不可约的干预误差,接口恢复随干预-响应覆盖率的提高而改善,并且当覆盖率和局部机制误差得到控制时,一个神谕因果组合可以超越非因果下界。然后我们在诊断性智能体环境中测试所得预测。因果接口在结构化工具环境中帮助最大,其中API签名暴露前置条件和下游效应。相比之下,对话和叙事环境通常忽略原始边列表,除非一个短注意力锚点使因果信息与决策相关。这些结果确定了LLM智能体中因果世界模型的一个具体条件:当跨模块接口在统计上可识别且以智能体在行动时能使用的形式呈现时,因果结构才有帮助。

英文摘要

LLM agents increasingly act through modular systems, such as order, payment, inventory, and shipment services, where actions in one module change which transitions are valid in another. Standard world models usually fit observational traces, but this is not the quantity needed for intervention-time planning: a trace may show that payment precedes shipment without identifying whether payment authorizes shipment, inventory mediates the effect, or a hidden trigger explains both. We study this gap through FedCausalCompose, a causal world-model framework for modular LLM agents in which local actions provide intervention-response evidence for cross-module interfaces. We first show that observational world models incur an irreducible interventional error under unblocked back-door paths, that interface recovery improves with intervention-response coverage, and that an oracle causal composition can beat the non-causal lower bound when coverage and local mechanism errors are controlled. We then test the resulting prediction in diagnostic agent settings. Causal interfaces help most in structured tool environments, where API signatures expose preconditions and downstream effects. In contrast, dialogue and narrative environments often ignore raw edge lists unless a short attention anchor makes the causal information decision-relevant. These results identify a concrete condition for causal world models in LLM agents: causal structure helps when cross-module interfaces are both statistically identifiable and presented in a form the agent can use at action time.

CommentsUnder Review

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑