arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

MUDDLE:测量干扰项和长度影响下的文档理解能力

MUDDLE: Measuring Understanding of Documents under Distractor and Length Effects

Jason Luo, Saibilila Abudukelimu, Judy Song, Andrew Feng, Shivank Garg, Vasu Sharma, Kevin Zhu

arXiv 2608.29477首次发表:更新:

AI 中文总结

针对文档问答系统混淆干扰项与长度影响的问题,提出受控基准MUDDLE分离二者效应,实验发现gpt-5-mini受主题相似难干扰项影响更大,发布数据与代码供可复现研究。

AI 中文摘要

文档问答系统越来越多地针对检索到的文档集合而非单一纯净源文档回答问题,因此对干扰上下文的鲁棒性与阅读能力同等重要。当这类系统出错时,往往难以确定是上下文过长还是干扰项与主题过于接近导致,因为现有研究倾向于将这两种效应混为一谈。我们提出了MUDDLE,一个将二者分离的受控基准。MUDDLE使用270个人工标注的问题,每个问题关联单一源文档,并在五种条件下实例化:仅源文档、源文档搭配2个或4个主题相似的难干扰项、源文档搭配2个或4个随机干扰项。随机干扰项在长度和来源上与难干扰项匹配,因此两组间的准确率差距反映的是主题相似性而非长度。所有五种条件均以Markdown、页面图像和原始PDF呈现,但此处报告的干扰项扫描在Markdown中运行,因为源文档加其干扰项超出了当前图像和PDF输入限制。我们使用LLM评判器在三个模型族上对答案打分。在完整的Markdown扫描中,对于gpt-5-mini,在两种上下文规模下,难干扰项比长度匹配的随机文档更能降低准确率,而随机文档的准确率接近无干扰项基线。该效应较小但方向一致,且对于gpt-5-mini,跨上下文规模汇总后,难干扰项的表现显著差于长度匹配的随机干扰项。我们发布了数据和评估代码,以便对上下文退化进行可复现研究。

英文摘要

Document question-answering systems increasingly answer questions over collections of retrieved documents rather than one clean source, so robustness to distracting context matters as much as reading ability. When such systems fail, it is often unclear whether the context was too long or the distractors were too close to the topic, because prior work tends to conflate these two effects. We present MUDDLE, a controlled benchmark that separates them. MUDDLE uses 270 human-annotated questions, each tied to a single source document, and instantiates every question in five conditions: the source alone, the source with two or four topically similar hard negatives, and the source with two or four random distractors. The random distractors are matched to the hard negatives in length and provenance, so an accuracy gap between the two arms reflects topical similarity rather than length. All five conditions are rendered in markdown, page images, and raw PDF, but the distractor sweep reported here is run in markdown, since a source plus its distractors exceeds current image and PDF input limits. We score answers with an LLM judge across three model families. In the complete markdown sweep, hard negatives lower accuracy more than length-matched random documents at both context sizes for gpt-5-mini, while random documents stay near the no-distractor baseline. The effect is small but directionally consistent, and for gpt-5-mini hard negatives significantly underperform length-matched random distractors when pooled across context sizes. We release the data and evaluation code for a reproducible study of context degradation.

Comments12 pages. Accepted to the Context Beyond the Window (CBW) workshop at COLM 2026 (non-archival). Code and data: https://github.com/luoojason/muddle

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑