arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

所选对象能被读者获取吗?审查接地语言模型流水线中的身份交接

Does the Selected Object Reach the Reader? Auditing Identity Handoffs in Grounded Language-Model Pipelines

Siddharth Vohra, Runmin Jiang, Xiaomo Li, Min Xu

arXiv 2609.04579首次发表:更新:

发表机构

Carnegie Mellon University; Amazon Web Services AI Native(卡内基梅隆大学; 亚马逊云科技AI原生部门)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本文审查接地语言模型流水线的对象身份交接问题,对比不同检索方法的对象返回效果,发现对齐对象可提升精确匹配得分,发布了Returned-Object Profile工具。

AI 中文摘要

接地语言模型流水线可分为三个阶段:选择对象、为该对象检索段落、利用证据回答问题。若所选对象必须被读者获取,丢失该对象即会破坏身份交接。基准召回检查与数据集关联的对象,二者可能存在差异。我们审查了三个选择器族的600个HybridQA问题。在所选对象与数据集追踪段落匹配的1463条可解析记录中,精确键查找和精确标题匹配每次都能返回该对象。在相同解码所选标题的情况下,仅正文BM25在截断值为5时,有389条记录(26.6%)未返回该对象,而带重排序的混合检索仅在14条记录(1.0%)中未返回。在1792条可解析记录中,两种身份存在差异的有329条。在原始问题排名下,二者的前5检查有106条记录(5.9%)不一致。冻结阅读器比较显示,对齐对象的存在与精确匹配得分提高28.6至31.0个百分点相关。在特意选取的64个样本队列中,移除该段落会显著降低精确匹配得分,而移除相似长度的比较段落不会产生相同降幅。我们发布了Returned-Object Profile(ROP),这是目标、返回ID字段、截断值、成员规则及完整预期总体的可执行记录,附带数据和离线重放功能。

英文摘要

Grounded language-model pipelines can be divided into three stages: selecting an object, retrieving passages for it, and using that evidence to answer. If the selected object must reach the reader, losing it breaks the handoff. Benchmark recall checks the dataset-linked object, which can differ. We audit 600 HybridQA questions across three selector families. On 1,463 resolvable records where the selected object matches the dataset-traced passage, exact key lookup and exact title matching return the object every time. With every ranked rule given the same decoded selected title, body-only BM25 omits it on 389 records (26.6%) at cutoff five, while hybrid retrieval with reranking omits it on 14 (1.0%). The two identities differ on 329 of 1,792 resolvable records. With original-question rankings, their top-five checks disagree on 106 records (5.9%). Frozen reader comparisons associate the aligned object's presence with 28.6 to 31.0 points higher exact match. In a deliberately selected 64-item cohort, removing that passage sharply lowers exact match, while removing a similar-length comparison passage does not reproduce the drop. We release the Returned-Object Profile (ROP), an executable record of the target, returned-ID field, cutoff, membership rule, and complete expected population, with data and an offline replay.

Comments15 pages, 1 figure, 23 tables. Accepted to the GroundLM Workshop (Grounding Language Models: Learning Faithfully and Efficiently) at EMNLP 2026

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑