面向证据-分类体系检索的分解式假设搜索
Factorized Hypothesis Search for Evidence-to-Taxonomy Retrieval
- The Fin AI
- Yale University(耶鲁大学)
- Cardiff University(卡迪夫大学)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
针对证据-分类体系检索的就绪差距问题,提出分解式假设搜索(FHS)方法,在两项任务中实现非最优方法中的最佳性能,且其并行首轮检索优于顺序优化。
AI中文摘要:
大型分类体系检索通常假设输入已表达目标概念,但在许多场景中,输入是间接证据,例如含义取决于行、列、数据类型和上下文的表格单元,我们将这种不匹配称为检索就绪差距。分析表明,当目标语义明确时,当前索引能可靠检索到目标,而原始证据常使其在排名中位置靠后。我们提出分解式假设搜索(Factorized Hypothesis Search, FHS),该方法在命名语义维度上维护多个部分解释,这些假设支持结构化查询生成、多假设检索和维度级候选验证。在金融分类体系标注和CodiEsp临床编码任务中,FHS在非最优方法中实现了最佳的Recall@1、MRR和最终准确率;将分解式假设路径替换为自由文本集成会导致头部排名性能大幅下降,而顺序优化未比FHS强大的并行首轮检索带来额外增益。
英文摘要:
Large-taxonomy retrieval often assumes that the input already expresses the target concept. In many settings, however, the input is indirect evidence, such as a table cell whose meaning depends on its row, column, datatype, and context. We call this mismatch the retrieval readiness gap. Our analysis shows that the current index retrieves the target reliably when its semantics are explicit, while raw evidence often leaves it deep in the ranking. We propose Factorized Hypothesis Search (FHS), which maintains multiple partial interpretations over named semantic dimensions. These hypotheses support structured query rendering, multi-hypothesis retrieval, and dimension-level candidate verification. On both financial taxonomy tagging and CodiEsp clinical coding tasks, FHS achieves the best Recall@1, MRR, and final accuracy among the non-oracle methods. Replacing the factorized hypothesis path with a free-text ensemble causes the largest drop in head-ranking performance, while sequential refinement provides no additional gain over FHS's strong parallel first round.