arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

文档衍生属性图中的隐藏关系:动态演化本体上的top-k块嵌入与逆距离加权

Hidden relationships in a document-derived property graph: top-k chunk embeddings and inverse-distance weighting over a dynamically evolving ontology

Bilge Kaan Karamete, Hunter Casten

arXiv 2609.00387首次发表:更新:

发表机构

Babel Street(巴别街公司)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

该研究针对文本提取知识图谱中语义相关实体脱节问题,提出一种二次处理方法,借助块嵌入、逆距离加权等技术,在多图数据库上实现了高效的潜在关联发现,兼具高边保真度与25倍加速效果。

AI 中文摘要

从文本中提取知识图谱的大型语言模型仅能捕获明确陈述的事实,常导致跨文档的语义相关实体相互脱节。我们提出一种可加、与引擎无关的二次处理步骤,在不改变已提取事实的前提下发现这些潜在关联。将每个文档分块并嵌入一次,通过现有块间的top-k最近邻查询,借助实体成员映射生成候选节点对。采用Shepard逆距离加权结合重缩放弦距离度量对候选对评分,避免了k-NN门控背后仿射余弦评分的阈值崩溃缺陷。无门控的每对累加器构成交换幺半群,确保流程严格与顺序无关,可增量扩展且无需重新计算之前的文档。在FalkorDB、Kinetica、ArangoDB和Neo4j上实现后,我们的方法显示,768维和240维嵌入分别保留了针对3072维基线的92%和72%边保真度,同时实现了25倍更快的top-k构建速度。

英文摘要

Large language models extracting knowledge graphs from text capture only explicitly stated facts, often leaving semantically related entities disconnected across documents. We present an additive, engine-neutral second pass that discovers these latent ties without altering extracted facts. Each document is chunked and embedded once; top-k nearest- neighbor queries across existing chunks yield candidate node pairs via entity membership maps. Candidate pairs are scored using Shepard inverse-distance weighting with a rescaled chord distance metric, avoiding the threshold-collapsing flaw of affine cosine scoring behind a k-NN gate. Un-gated per-pair accumulators form a commutative monoid, ensuring the pipeline is strictly order-independent and scales incrementally without recomputing prior documents. Implemented across FalkorDB, Kinetica, ArangoDB, and Neo4j, our method shows that 768- and 240-dimensional embeddings retain 92% and 72% edge fidelity against a 3072-D baseline while achieving a 25x faster top-k formulation.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑