发表机构
LIBIS; KU Leuven(LIBIS; 荷语鲁汶大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对同一计算问题在不同领域名称各异、现有嵌入无法识别的问题,提出剥离领域的计算指纹表示,显著提升跨领域解决方案导入的检索精度,并验证了其有效性。
AI 中文摘要
同一个底层计算问题在不同不相关领域中会以不同名称被解决:递归贝叶斯状态估计在控制领域被称为“卡尔曼滤波器”,在药代动力学中被称为“贝叶斯预测”,在地球科学中被称为“数据同化”。基于主题和引用的科学嵌入无法识别这种共享问题。我们将每篇论文提炼一次,生成一个剥离领域和方法名称的分面计算指纹,该指纹由自由文本机制骨架和受控计算分面组成。我们在此之上定义了一个可调、可选择性分面的距离。目标是解决方案导入:发现解决同一问题的跨领域配对,以便将定制实现替换为另一领域的标准专用求解器。在一个涵盖18个方法族、109篇论文的基准测试上,骨架将跨领域检索的平均精度从摘要的0.222提升到0.513,完整指纹达到0.557。值得注意的是,四种经过训练的科学嵌入器均低于纯摘要加TF-IDF:它们编码了主题和引用相似性,而这对于此任务而言是错误的信号。增益在于表示:从摘要到骨架的替换提升了每个嵌入器的性能,且流程是每篇论文一次缓存的LLM调用加上一个廉价的嵌入器。一项干预性的重新包装/数学编辑测试表明,指纹跟踪的是计算而非领域。在一个包含501篇论文的野生语料库中,已知的孪生对主导了排名顶部(前30名中有23个);在排除植入配对后,三位盲法LLM评审员将前5名中的3对和前30名中的8对评为真正的导入候选,而随机配对中为0对。人工验证是四次实际导入:其中一次,一个开放标准求解器复现了定制临床给药引擎的输出。我们发布了基准测试、代码和提炼提示。
英文摘要
The same underlying computational problem is solved across unrelated fields under different names: recursive Bayesian state estimation appears as a "Kalman filter" in control, "Bayesian forecasting" in pharmacokinetics, and "data assimilation" in geoscience. Topical and citation-based scientific embeddings cannot see this shared problem. We distill each paper once into a domain- and method-name-stripped faceted computational fingerprint, a free-text mechanism skeleton plus controlled computational facets. We define a tunable, facet-selectable similarity over it. The goal is solution import: surface cross-field pairs solving the same problem, so a bespoke implementation can be swapped for another field's standard, specialized solver. On a benchmark of 18 method families across 109 papers, the skeleton lifts cross-domain retrieval average precision over the abstract from 0.222 to 0.513, and the whole fingerprint reaches 0.557. Strikingly, four trained scientific embedders all fall below plain abstract+TF-IDF: they encode topical and citation similarity, the wrong signal for this task. The gain is the representation: the abstract-to-skeleton swap lifts every embedder, and the pipeline is one cached LLM call per paper plus a cheap embedder. An interventional re-skin / math-edit test shows the fingerprint tracks the computation, not the field. On a 501-paper wild corpus, known twins dominate the top of the ranking (23 of the top 30); with planted pairs excluded from the results, three blind LLM judges rate 3 of the top 5 and 8 of the top 30 pairs genuine import candidates, and 0 of 30 random ones. The human verification is the four executed imports: in one, an open standard solver reproduces a bespoke clinical dosing engine's output. We release the benchmark, the code, and the distillation prompt.
CommentsAccepted as a full paper at JCDL 2026 (The 2026 ACM/IEEE Joint Conference on Digital Libraries), Frisco, TX, October 13-16, 2026. 10 pages plus references, 2 figures, 8 tables. Code and benchmark: https://github.com/ErykKul/same-problem-different-field ; archived dataset (KU Leuven RDR): https://doi.org/10.48804/W3B9WC