发表机构
SCSK Corporation(SCSK株式会社)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
MaSCoD通过多智能体框架在直接边判断前组织候选变量和结构模式,减少因果图生成中的关系遗漏,实验表明该方法能提升召回率和F1分数。
AI 中文摘要
大型语言模型(LLMs)已被应用于因果发现,但候选图生成很少将潜在相关因果关系的过早遗漏作为明确的设计目标。我们提出MaSCoD,一种多智能体框架,在直接边判断之前组织候选第三变量和局部结构模式。我们在Auto-MPG、DWD和Sachs数据集上使用GPT-5.4作为主要骨干模型、GPT-4o用于复现来评估MaSCoD。MaSCoD表现出依赖于数据集和骨干模型的保留-选择性特征,而非统一的优越性。在所有六种数据集-骨干模型组合中,Full(在直接边判断之前提供结构假设)相比No Phase 1(在判断过程中构建结构假设)实现了更高的平均召回率和F1分数,但同时也增加了假阳性率。在DWD上使用GPT-5.4以及在Sachs上使用GPT-4o时,观察到相对于所有评估基线的额外参考边保留,而非在所有设置中统一出现。部分消融实验表明,同时提供两种信息组件并不总是优于仅提供一种。对于GPT-5.4,分阶段分析显示,Full与No Phase 1之间的保留差距在直接边判断后已经存在,而协调阶段在Sachs上为Full引入了额外的参考边损失。这些发现支持将结构预组织作为遗漏控制的明确设计和评估目标,并激励将上下文构建与其在判断中的利用联合评估。
英文摘要
Large language models (LLMs) have been applied to causal discovery, but candidate-graph generation rarely treats premature omission of potentially relevant causal relations as an explicit design objective. We propose MaSCoD, a multi-agent framework that organizes candidate third variables and local structural patterns before direct-edge judgment. We evaluate MaSCoD on Auto-MPG, DWD, and Sachs using GPT-5.4 as the primary backbone and GPT-4o for replication. MaSCoD exhibits a dataset- and backbone-dependent retention-selectivity profile rather than uniform superiority. Across all six dataset-backbone settings, Full, which supplies structural hypotheses before direct-edge judgment, achieved higher mean Recall and F1 than No Phase 1, which instead constructs them within the judgment procedure, while also increasing false-positive rates. Additional reference-edge retention over all evaluated baselines was observed on DWD with GPT-5.4 and on Sachs with GPT-4o, rather than uniformly across settings. Partial ablations showed that supplying both information components did not always outperform supplying only one. For GPT-5.4, stage-wise analysis showed that the Full-No Phase 1 retention gap was already present after direct-edge judgment, while reconciliation introduced additional reference-edge loss for Full on Sachs. These findings support structural pre-organization as an explicit design and evaluation target for omission control and motivate evaluating context construction jointly with its utilization in judgment.
Comments31 pages, 4 figures, 18 tables. The first two authors contributed equally