发表机构
Faculty of Arts and Sciences; University of Sciences and Arts in Lebanon(文理学院; 黎巴嫩科学与艺术大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本研究提出整合多组件的基于证据的无参考智能体数据清洗框架,经多数据集实验发现,额外能力会在多指标间引入权衡,无配置在所有标准上最优。
AI 中文摘要
在无可信干净参考的数据清洗场景中,任务极具挑战性,因为异常值既可能是真实错误,也可能是有效观测值。本文研究不同智能体能力如何影响无参考数据清洗,并提出一种基于证据的框架,该框架整合结构化上下文、数据探查(profiling)、大语言模型(LLM)推理、可执行检查、受控证据检索、源排名、引文对齐、保守修复、可逆脚本及 provenance 日志记录。通过受控合成污染与原始数据描述性分析,在金融、临床及环境监测数据集上评估7种配置,共完成126次运行。评估包含两个对比基线及一个渐进式基于LLM的序列,该序列逐步添加可执行工具、证据检索、证据控制与保守修复。在合成评估中,确定性探查基线达到最高检测F1分数0.561;在基于LLM的配置中,完整保守配置达到最高F1分数0.421,但无任何配置在所有评估标准上表现最优。源排名配置实现最低无支持规则率,而决策级引文对齐仍较弱。完整保守配置未产生不安全或不必要修改(即便在添加保守策略前此类率已为0),且未执行直接修复。总体而言,结果表明额外能力会在检测、修复、证据基础、保守行为、可复现性及运营成本间引入权衡,而非带来一致改进。本研究为评估无参考智能体数据清洗中的这些权衡提供了结构化框架与实证方法。
英文摘要
Data cleaning without a trusted clean reference is challenging because unusual values may represent either genuine errors or valid observations. This paper studies how different agent capabilities affect reference-free data cleaning and proposes an evidence-grounded framework that combines structured context, profiling, LLM reasoning, executable checks, controlled evidence retrieval, source ranking, citation alignment, conservative repair, reversible scripts, and provenance logging. Seven configurations are evaluated across financial, clinical, and environmental-monitoring datasets using controlled synthetic corruption and original-data descriptive analysis, resulting in 126 completed runs. The evaluation includes two comparison baselines and a progressive LLM-based sequence that adds executable tools, evidence retrieval, evidence controls, and conservative repair. In the synthetic evaluation, the deterministic profiling baseline achieved the highest detection F1-score of 0.561. Among the LLM-based configurations, the full conservative configuration achieved the highest F1-score of 0.421, but no configuration performed best across all evaluation criteria. The source-ranked configurations achieved the lowest unsupported-rule rates, while decision-level citation alignment remained weak. The full conservative configuration produced no unsafe or unnecessary modifications, although these rates were already zero before the conservative policy was added, and it performed no direct repairs. Overall, the results show that additional capabilities introduce trade-offs among detection, repair, evidence grounding, conservative behaviour, reproducibility, and operational cost rather than producing consistent improvements. The study provides a structured framework and empirical methodology for evaluating these trade-offs in reference-free agentic data cleaning.
Comments21 pages, 3 figures, Submitted to New Generation Computing