AI 中文总结
针对现有实体解析被动范式无法适配异构流数据场景的问题,提出Agentic ER新范式,将其建模为自主智能体的序列决策过程,构建参考架构并明确研究挑战,开拓数据管理与智能体交叉的新研究方向。
AI 中文摘要
实体解析(Entity Resolution, ER)是数据管理中的基础问题,在数据清洗、知识图谱构建等任务中发挥关键作用。现有ER方法涵盖传统规则式、深度学习及基于大语言模型(LLM)的方法,但通常遵循“被动范式”,通过静态的一次性相似度计算检测重复项。这类方法无法捕捉现实世界ER任务固有的不确定性与上下文依赖性,尤其是在包含CSV、JSON、RDF转储、自由文本等异构格式流数据的数据湖中。在此类场景下,解决歧义往往需要迭代收集证据、跨多源推理,甚至选择性引入人类参与。为填补这一空白,我们倡导从被动范式向Agentic ER的转变,将ER建模为自主智能体执行的序列决策过程。这些智能体主动规划ER策略、获取外部证据、决定何时查询额外数据源或人类,并在准确率、成本与延迟之间优化权衡。我们将Agentic ER形式化为决策论问题,提出参考架构,明确核心研究挑战,并勾勒出适配智能体行为的新评估维度。通过引入Agentic ER,我们旨在建立数据管理与智能体交叉领域的全新研究方向。
英文摘要
Entity Resolution (ER) is a fundamental problem in data management, playing a critical role in tasks like data cleaning and knowledge graph construction. The existing ER approaches range from traditional rule-based to deep learning techniques and LLM-based methods, but typically operate under a ``passive paradigm'', as duplicates are detected through static, one-shot similarity computations. Such approaches fail to capture the inherently uncertain and context-dependent nature of real-world ER tasks, especially in data lakes with streaming content in heterogeneous formats such as CSV files, JSON files, RDF dumps, and free text. In such settings, resolving ambiguity often requires iterative evidence gathering, reasoning across multiple sources, even selective human involvement. To cover this gap, we advocate a paradigm shift from passive to Agentic ER, which frames ER as a sequential decision-making process that is performed by autonomous agents. These agents actively plan ER strategies, acquire external evidence, decide when to query additional sources or humans, and optimize trade-offs between accuracy, cost, and latency. We formalize Agentic ER as a decision-theoretic problem, we propose a reference architecture, we identify core research challenges, and outline new evaluation dimensions tailored to agentic behavior. By introducing Agentic ER, we aim to establish a new research direction at the intersection of data management and intelligent agents.
CommentsVision paper submitted to VLDB, 6 pages