arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.01645cs.IR

词汇差距即公平差距:公共福利获取检索系统中的语域不匹配

The Vocabulary Gap Is an Equity Gap: Register Mismatch in Retrieval Systems for Public-Benefits Access

Krish Sapru

首次发表
浏览论文内容

中文总结 AI 辅助

本研究发现公共福利检索系统存在语域不匹配的公平差距,构建基准后证实平实语域检索性能大幅下降,提出词汇桥接方法可有效缓解该问题,贡献了相关评估与缓解方案。

中文摘要 AI 辅助

检索增强型问答正越来越多地用于帮助人们了解公共福利资格,但这些系统检索的文档采用官方机构语域书写,而目标用户常以平实、非正式或非母语英语提问。我们表明这种语域不匹配会将高性能检索系统转变为不公平系统。我们构建了包含51项公开记录的联邦福利资格规则和25个信息需求的受控基准,每个需求均以官方机构语域和平实用户语域表述,同时保持黄金段落固定。在BM25、TF-IDF和术语图检索器上,正式语域评估近乎完美(Recall@5为96%-100%),但平实语域检索表现大幅下降(Recall@5为36%-44%)。对于BM25,Recall@1从84%降至16%,Recall@5从100%降至44%,相同信息需求下存在56个百分点的公平差距。我们将该机制追溯至可测量的词汇差距:正式查询与黄金段落共享0.63的内容词,而平实查询仅共享0.11,减少了5.9倍。一种刻意设计的简单、可审计的平实到正式词汇桥接方法恢复了大部分失效,将平实查询的BM25 Recall@5从44%提升至80%。本研究的贡献并非新检索器,而是针对标准检索评估所掩盖的高风险社会影响失效模式的评估协议、基准、机制诊断及透明缓解措施。

英文摘要

Retrieval-augmented question answering is increasingly used to help people navigate public-benefits eligibility, yet the documents these systems retrieve from are written in agency register while intended users often ask questions in plain, informal, or non-native English. We show that this register mismatch can turn a high-performing retrieval system into an inequitable one. We construct a controlled benchmark of 51 publicly documented federal benefit-eligibility rules and 25 information needs, each phrased in both agency register and plain user register while keeping the gold passage fixed. Across BM25, TF-IDF, and a term-graph retriever, formal-register evaluation is nearly perfect (Recall@5 96-100%), but plain-register retrieval collapses (Recall@5 36-44%). For BM25, Recall@1 falls from 84% to 16% and Recall@5 from 100% to 44%, a 56-point equity gap on identical information needs. We trace the mechanism to a measurable vocabulary gap: formal queries share 0.63 of their content terms with the gold passage, while plain queries share only 0.11, a 5.9x reduction. A deliberately simple, auditable plain-to-formal lexicon bridge recovers much of the failure, lifting plain-query BM25 Recall@5 from 44% to 80%. The contribution is not a new retriever; it is an evaluation protocol, benchmark, mechanistic diagnosis, and transparent mitigation for a high-stakes social-impact failure mode that standard retrieval evaluation hides.

补充信息

↑