AI 中文总结
KnowFeat是将五类结构化领域知识注入LLM智能体的特征工程框架,经三阶段验证后在12个公开基准中排名第一,在反洗钱数据集上也具备良好性能与溯源性。
AI 中文摘要
利用大语言模型(LLM)开展自动化特征工程可为表格数据生成语义有意义的特征,但现有方法缺乏结构化领域知识、严格验证及可解释溯源。我们提出KnowFeat,一种知识引导的特征工程框架,将领域知识组织为模式元数据、监管指标、检测规则、专家意见及法院文档证据共五类,并将其作为结构化上下文注入LLM智能体。该框架采用三阶段验证流程,通过代码执行、统计质量检查及模型有效性评估筛选候选特征;每个被接受的特征均带有溯源记录,可将其设计追溯至特定知识资产。在消除特征选择泄露的严格留出协议下,KnowFeat在7种方法的12个公开基准中排名第一(平均排名2.3,单侧Wilcoxon检验p值为0.017),在电信 churn 数据集上的AUC峰值提升达11.6个百分点;在真实世界比特币反洗钱(AML)数据集(Elliptic)及合成数字货币反洗钱基准(SimECNY)上,KnowFeat在保持竞争检测性能的同时具备完整溯源能力。
英文摘要
Automated feature engineering with large language models (LLMs) can produce semantically meaningful features for tabular data, yet existing methods lack structured domain knowledge, rigorous verification, and explainable provenance. We propose KnowFeat, a knowledge-guided feature engineering framework that organizes domain knowledge into five types -- schema metadata, regulatory indicators, detection rules, expert opinions, and court document evidence -- and injects them as structured context into an LLM agent. A three-stage verification pipeline filters candidates through code execution, statistical quality checks, and model effectiveness evaluation. Every accepted feature carries a provenance record tracing its design to specific knowledge assets. Under a strict held-out protocol that eliminates feature-selection leakage, KnowFeat ranks first (avg. rank 2.3) across twelve public benchmarks among seven methods (one-sided Wilcoxon p=0.017), with a peak gain of +11.6 pp AUC on a telecom churn dataset. On a real-world Bitcoin anti-money laundering (AML) dataset (Elliptic) and a synthetic digital currency AML benchmark (SimECNY), KnowFeat maintains competitive detection performance with full provenance traceability.
Comments12 pages, 13 tables, 2 figures. Under review