arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

从位置短语到地理实体:面向人员搜索的任务适配检索

From Location Phrases to Geographic Entities: Task-Adapted Retrieval for People Search

Yanbo Li, Chujie Zheng, Jiahao Xu, Chetan Bhole, Lingyu Zhang, Puneet Singh Ahluwalia, Kevin Nguyen, Raghavan Muthuregunathan, Santhosh Sachindran, Sachin Ahuja, Fedor Borisyuk

arXiv 2608.28965首次发表:更新:

发表机构

LinkedIn(领英)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

该研究针对人员搜索中位置短语转地理实体的问题,提出任务适配的地理实体检索方法,在基准与迁移任务上提升了检索性能,可替代现有基于分类学的标准化器。

AI 中文摘要

人员搜索需要将自由形式的位置短语映射为用作结构化检索过滤器的地理实体。词汇标准化器能很好地处理规范名称,但对于别名、拼写错误、都会区表达以及同名歧义问题较为脆弱。我们将此任务表述为在固定本体上的分级集值实体检索。我们确定了三个耦合的设计要求:区分身份保留型变体与依赖知识的别名、控制有效同名实体中的假阴性、分离稳定转换与可变实体知识。我们通过带有校准别名支持的提示不对称双编码器、带边界的歧义感知负样本以及可编辑实体文档来实现这些要求,该文档支持本地化更新而无需重新训练。在固定的生产衍生开发基准和公开的GeoNames迁移任务上,与冻结编码器和标准token基线相比,任务适配的性能显著提升。受控开发 ablation 实验表明,专门的监督学习在标准任务微调与编码器缩放之外另有贡献。在GeoNames上,适配后的模型在零到中等字符重叠区间内提升了已知目标的Recall@1,而字符n元语法在总体Target Recall@5上仍保持小幅优势。在分层生产挑战集的盲法人工对比中,我们的模型将相关P@1从28.0%提升至46.0%(p=0.012);固定查询端点估计在非规范查询上有所改进,在频繁查询上与对照保持接近;随机实时实验未检测到参与度下降。这些结果支持将任务适配的地理实体检索作为现有基于分类学的标准化器的实用替代方案,在非规范查询上的相关性提升最大。

英文摘要

People search must map free-form location phrases to geographic entities used as structured retrieval filters. Lexical standardizers handle canonical names well but are brittle to aliases, misspellings, metropolitan expressions, and same-name ambiguity. We formulate this task as graded, set-valued entity retrieval over a fixed ontology. We identify three coupled design requirements: distinguishing identity-preserving variation from knowledge-dependent aliases, controlling false negatives among valid same-name entities, and separating stable transformations from mutable entity knowledge. We realize them in a prompt-asymmetric bi-encoder with calibrated alias support, bounded ambiguity-aware negatives, and editable entity documents that support localized updates without retraining. Across a fixed production-derived development benchmark and a public GeoNames transfer task, task adaptation improves substantially over frozen encoders and standard token baselines. Controlled development ablations show that specialized supervision contributes beyond standard task fine-tuning and encoder scaling. On GeoNames, the adapted model improves known-target Recall@1 throughout zero-to-moderate character overlap, while character n-grams retain a small aggregate Target Recall@5 advantage. In a blinded human comparison on a stratified production challenge set, our model raises relevant P@1 from 28.0% to 46.0% (p=0.012). Fixed-query endpoint estimates improve on non-canonical queries and remain close to control on frequent queries; a randomized live experiment detects no engagement regression. These results support task-adapted geographic entity retrieval as a practical replacement for the incumbent taxonomy-based standardizer, with the largest relevance gains on non-canonical queries.

Comments11 pages, 1 figure, 9 tables

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑