arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2607.24826cs.IRcs.MA

Aethel:用于多跳金融尽职调查的可重现图检索框架

Aethel: A Reproducible Graph-Retrieval Framework for Multi-Hop Financial Diligence

Krish Sapru

首次发表
浏览论文内容

中文总结 AI 辅助

针对二级私募股权交易中财务披露信息综合难题,Aethel框架结合二分个性化PageRank图检索等技术,将语料库建模为实体-段落图。在多跳金融尽职调查任务中,该框架在特定数据集上有较好表现,优势依赖语料库规模和实体索引质量,且发布相关工件保证可重现性。

中文摘要 AI 辅助

二级私募股权交易需要快速综合零散、无结构的财务披露信息,关键指标及其实体锚点分布在词汇重叠有限的不相关文档中。我们提出了Aethel,这是一个可重现的框架,它将二分个性化PageRank图检索与共指感知二分共指瞬移层和精心编排的专家代理架构相结合。Aethel将语料库建模为实体-段落图,并通过显式关系路径传播相关性以支持多跳金融尽职调查。我们在来自MuSiQue和2WikiMultiHopQA验证集的200个问题样本上评估检索层,比较稀疏词汇、密集双编码器、普通图和共指感知图检索。Aethel在2WikiMultiHopQA上的HR@5为100.0%,在MuSiQue上为88.5%,在牺牲顶级精度的同时提高了对普通图检索的覆盖率。我们还在一个由4123个金融披露块组成的语料库上评估检索,发现图检索在多跳召回率上优于密集检索,但在开放语料库规模上不超过强大的BM25基线。结果表明,基于图的检索提供了可解释的多跳证据路径,并且随着语料库大小的增加,其退化比密集检索更平缓,同时也表明其优势强烈依赖于语料库规模和实体索引质量。为了可重现性,我们发布了代码和评估工件。

英文摘要

Secondary private equity transactions require rapid synthesis of fragmented, unstructured financial disclosures, where critical metrics and their entity anchors are distributed across disjoint documents with limited lexical overlap. We present Aethel, a reproducible framework that combines bipartite Personalized PageRank graph retrieval with a coreference-aware Bipartite Coreference Teleportation layer and an orchestrated specialist-agent architecture. Aethel models corpora as entity-passage graphs and propagates relevance through explicit relational paths to support multi-hop financial diligence. We evaluate the retrieval layer on 200-question samples from the MuSiQue and 2WikiMultiHopQA validation sets, comparing sparse lexical, dense bi-encoder, vanilla graph, and coreference-aware graph retrieval. Aethel achieves HR@5 of 100.0% on 2WikiMultiHopQA and 88.5% on MuSiQue, improving coverage over vanilla graph retrieval while trading off top-rank precision. We also evaluate retrieval over a 4,123-chunk corpus of financial disclosures and find that graph retrieval outperforms dense retrieval on multi-hop recall but does not surpass a strong BM25 baseline at open-corpus scale. The results show that graph-based retrieval offers interpretable multi-hop evidence paths and degrades more gracefully than dense retrieval as corpus size grows, while also demonstrating that its advantage depends strongly on corpus scale and entity-index quality. Code and evaluation artifacts are released for reproducibility.

补充信息

↑