arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

H2CE:用于异构两阶段交叉编码器的POI重排序的地理语义交互建模

H2CE: Modeling Geo-Semantic Interactions for POI Reranking with Heterogeneous Two-Stage Cross-Encoders

Zhengwei Bai, Moreno D'Incà, Danielle Class, Alessandro Moschitti

arXiv 2610.11277首次发表:更新:

发表机构

Amazon; University of Trento(亚马逊; 特伦托大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

H2CE是一种异构两阶段交叉编码器,通过融合语义与数值特征实现POI重排序,在5743个查询的测试集上较XGBoost LTR和零样本LLM重排序器取得显著性能提升。

AI 中文摘要

本地搜索中的兴趣点(POI)重排序必须建模查询条件下的权衡,包括词汇语义、地理空间邻近度以及评分、评论数等数值质量信号,同时需满足实时服务的实用性约束。邻近的POI可能仅部分满足查询意图,而较远的POI可能提供更强的语义和质量证据。本文提出H2CE,一种用于延迟受限POI重排序的异构两阶段交叉编码器。H2CE以两种互补方式表示数值属性:将分桶的自然语言描述符插入交叉编码器输入以支持语义-数值注意力,同时将精确标量值通过专用多层感知器(MLP)处理以保留量级信息。生成的语义和数值嵌入通过隐空间聚合融合,实现超出标量加权和的非线性交互。H2CE采用两阶段架构:第一阶段对所有候选进行逐点评分以实现可扩展过滤,第二阶段对前K个候选进行两两比较并采用Copeland聚合,明确细粒度相对权衡,同时将两两比较的成本从O(N²)降至O(N+K(K-1))。在包含5743个查询的本地搜索测试集上,H2CE达到67.48%的NDCG@5,较XGBoost LTR提升22.82个百分点,较零样本大语言模型(LLM)重排序器提升35.89个百分点;两两阶段较单独的逐点模型使NDCG@5提升1.98个百分点。 ablation实验证实了数值特征、隐空间聚合、前K个两两重排序及对齐训练的价值。

英文摘要

Point-of-Interest (POI) reranking in local search must model query-conditioned tradeoffs among lexical semantics, geospatial proximity, and numerical quality signals such as rating and review count, while remaining practical under real-time serving constraints. A close POI may only partially satisfy the query intent, while a farther one may offer stronger semantic and quality evidence. We present H2CE, a Heterogeneous Two-stage Cross-Encoder for latency-bounded POI reranking. H2CE represents numerical attributes in two complementary ways: bucketized natural-language descriptors are inserted into the cross-encoder input to support semantic--numeric attention, while exact scalar values are processed by dedicated MLPs to preserve magnitude information. The resulting semantic and numerical embeddings are fused through latent-space aggregation, enabling nonlinear interactions beyond scalar weighted sums. H2CE then applies a two-stage architecture: Stage 1 scores all candidates pointwise for scalable filtering, and Stage 2 performs head-to-head pairwise comparison among the top-$K$ candidates with Copeland aggregation, making fine-grained relative tradeoffs explicit while reducing pairwise cost from O(N^2) to O(N+K(K-1)). On a 5,743-query local search test set, H2CE achieves 67.48% NDCG@5, improving over XGBoost LTR by +22.82% absolute and over a zero-shot LLM reranker by +35.89%. The pairwise stage adds +1.98% NDCG@5 over the pointwise model alone. Ablations confirm the value of numerical features, latent aggregation, top-K pairwise reranking, and aligned training.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑