arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2607.10588cs.AIcs.CL

用于规则驱动的细粒度分类的约束感知层次搜索

Constraint-Aware Hierarchical Search for Regulation-Driven Fine-Grained Classification

Siyu Wang, Wei Tan, Lulu Chen

首次发表
浏览论文内容

中文总结 AI 辅助

针对规则驱动的细粒度分类任务,现有方法存在不足。本文构建基准数据集,提出约束感知层次搜索框架,将监管文件转换为可搜索树,只检索有效局部候选节点,实验证明该方法准确率高且能提供可解释决策路径。

中文摘要 AI 辅助

海关关税分类、出口管制分类和基于标准的设备编码等任务需要在明确的监管层次结构下将输入实例分配到细粒度类别。与标准文本分类不同,这些任务中正确标签并非仅由语义相似性决定,还受规则定义的边界、阈值条件等影响。现有方法无法联合执行层次有效性、规则一致性和细粒度边界推理。本文将此设置为规则驱动的细粒度层次分类,构建四个基准数据集并验证标注。还提出约束感知层次搜索框架,实验表明该方法在所有四个数据集上实现最佳平均准确率并提供可解释决策路径。

英文摘要

Tasks such as customs tariff classification, export control categorization, and standards-based equipment coding require assigning an input instance to a fine-grained class under an explicit regulatory hierarchy. Unlike standard text classification, the correct label in these tasks is not determined by semantic similarity alone, but by rule-defined boundaries, threshold conditions, exclusion clauses, definitions, and local exceptions. As a result, two highly similar inputs may require different labels, while a retrieved passage that appears relevant may still be inapplicable under the governing rules. Existing flat classifiers, hierarchical text classification methods, and retrieval-augmented LLM systems are not designed to jointly enforce hierarchical validity, rule consistency, and fine-grained boundary reasoning. In this paper, we formulate this setting as regulation-driven fine-grained hierarchical classification, where an external instance must be assigned to a fine-grained class through a valid path in a regulatory hierarchy and supported by auditable evidence. We construct four benchmark datasets from representative regulation-intensive scenarios and validate the annotations through an expert-in-the-loop process. We further propose a constraint-aware hierarchical search framework that converts regulatory documents into a searchable tree, retrieves only valid local candidate nodes, and uses structured regulatory fields with evidence snippets to guide each next-hop decision. Experiments show that our method achieves the best mean accuracy on all four datasets and provides interpretable decision paths, with the largest gains on cases involving fine-grained neighboring categories and rule-based boundary conditions.

发表机构

  • Gusu Laboratory of Materials(姑苏材料实验室)
  • Chongqing Institute of Engineering(重庆工程学院)
  • Suzhou Digital China Wuxin Intelligent Technology Co., Ltd.(苏州神州数码五新智能科技有限公司)

机构由 AI 辅助整理,请以论文原文为准。

↑