arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

基于大语言模型符号化决策过程实现的可解释列注释

Interpretable Column Annotation with LLM-Symbolized Decision Process Materialization

Mengqi Wang, Jianwei Wang, Qing Liu, Xiwei Xu, Zhenchang Xing, Michael Bain, Liming Zhu, Wenjie Zhang

arXiv 2607.25228首次发表:更新:

发表机构

Data61, CSIRO(数据61,联邦科学与工业研究组织)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

研究针对列注释方法存在的问题,提出SymCA框架,通过大语言模型赋能,将其实现为全局到局部的符号决策过程,含全局骨架归纳和局部基质演化两组件,实验证明该框架准确、稳健且可解释,性能优于基线。

AI 中文摘要

列注释(CA),包括列类型注释(CTA)和列属性注释(CPA),旨在识别表列的含义及其之间的语义关系。近期的CA方法通常使用各种神经模型来学习列表示并直接映射到标签类别,这牺牲了模型的可解释性和适应性,且忽视了丰富的标签语义,最终限制了准确性。为解决这些限制,我们提出了SymCA,这是一个由大语言模型赋能的可解释CA框架,将列注释实现为一个从全局到局部的符号决策过程。SymCA由两个组件组成:全局骨架归纳,它在标签空间上构建语义骨架;局部基质演化,它在骨架内演化预测基质。具体来说,全局骨架归纳模块利用大语言模型生成候选的受上位词启发的树形语义骨架,并采用基于最小贝叶斯风险(MBR)的共识策略来选择一个抗生成方差的稳健骨架。由于不同的内部节点需要不同的证据来区分其子节点,局部基质演化模块将每个内部节点实现为一个可执行且可演化的预测基质。在多个演化轮次中,每个基质使用当前的操作符集训练一个可解释的随机森林分类器,利用大语言模型提出特定节点的操作符修改,并使用探索-利用策略对有前景的基质进行优先级排序。大量实验表明,SymCA准确、稳健且可解释,在微F1上比最强基线平均高出6.42%,在宏F1上高出11.03%。

英文摘要

Column annotation (CA), including column type annotation (CTA) and column property annotation (CPA), aims to identify the meanings of table columns and the semantic relationships among them. Recent CA methods usually use various neural models to learn column representations and directly map them to label categories, thereby (1) sacrificing model interpretability and adaptivity, and (2) overlooking rich label semantics and ultimately limiting accuracy. To address these limitations, we propose SymCA, an LLM-empowered interpretable CA framework that materializes column annotation as a global-to-local symbolic decision process. SymCA consists of two components: (1) global skeleton induction, which constructs a semantic skeleton over the label space, and (2) local substrate evolution, which evolves predictive substrates within the skeleton. Specifically, to exploit label semantics while preserving an interpretable decision process, the global skeleton induction module leverages LLMs to generate candidate hypernym-inspired tree-structured semantic skeletons and employs a Minimum Bayes Risk (MBR)-based consensus strategy to select a robust skeleton against generation variance. Since different internal nodes require different evidence to distinguish among their child nodes, the local substrate evolution module materializes each internal node as an executable and evolvable predictive substrate. Over multiple evolution rounds, each substrate trains an interpretable random forest classifier with the current operator set, leverages the LLM to propose node-specific operator modifications, and uses an exploration-exploitation strategy to prioritize promising substrates. Extensive experiments demonstrate that SymCA is accurate, robust, and interpretable, outperforming the strongest baselines by an average of 6.42% in Micro-F1 and 11.03% in Macro-F1.

Comments13 pages, 7 figures

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑