发表机构
Mohamed bin Zayed University of Artificial Intelligence; Carnegie Mellon University(穆罕默德·本·扎耶德人工智能大学; 卡内基梅隆大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
该研究提出类比深度研究(ADR)任务及ADR-bench,发现LLM智能体找类比能力差。基于理论分析提出ADR原则,构建因果类比研究者(CANA)框架,提升历史类比生成效果,超越现有智能体,案例研究证实其利用历史类比的有效性。
AI 中文摘要
系统地将当前情况与历史上结构相似的过去事件进行比较,即历史类比,是前瞻性分析最有力的工具之一。在这项工作中,我们向大语言模型(LLM)智能体提出了一项名为类比深度研究(ADR)的新任务,并构建了第一个ADR基准ADR-bench,以研究LLM智能体在进行前瞻性分析时是否能够找到并利用历史类比。我们的调查揭示了一个关键障碍:LLM智能体在寻找类比方面很差,因为它们基于表面特征而非潜在机制进行匹配。我们认为ADR本质上是一个因果问题,因为它需要理解事件发生的原因。基于理论分析,我们提出了ADR所需的两个原则,包括机制对齐和交叉类比确认。基于理论结果,我们提出了一个名为因果类比研究者(CANA)的新智能体框架,指导LLMs找到并整合历史类比。CANA采用了一种简单而有效的结构分解表示,并整合了结构反馈以对历史类比识别和整合进行反思性改进。我们表明,CANA在历史类比生成方面带来了高达10%的提升,并在ADR-bench中超越了最先进的深度研究智能体。对当前事件的案例研究证实了CANA在利用历史类比方面的有效性。
英文摘要
Systematic comparisons between current situations and structurally similar past events in the historical, i.e., historical analogies, is among the most powerful tools for foresight analysis. In this work, we present a new task called Analogical Deep Research (ADR) to Large Language Model (LLM) agents and construct the first ADR benchmark ADR-bench to study whether LLM agents are able to find and leverage historical analogies when doing foresight analysis. Our investigation reveals a key obstacle: LLM agents are poor at finding analogies because they match on surface features rather than underlying mechanisms. We argue that ADR is inherently a causal question as it requires understanding why the event occurred. Based on our theoretical analysis, we propose two principles required for ADR, including the mechanism alignment and cross-analogy confirmation. Built upon our theoretical results, we propose a new agentic framework called Causal Analogical Researcher (CANA) that guides LLMs to find and integrate historical analogies. CANA incorporates a simple yet effective structural decomposition representation, and integrates structural feedback for reflective improvements of historical analogy identification and integration. We show that CANA brings up to 10% improvements in historical analogy generation, and surpasses the state-of-the-art deep research agents in the ADR-bench. Case studies with the ongoing events confirm the effectiveness of CANA in leveraging historical analogies.
CommentsOngoing project