arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2610.04118cs.CL

LongSocialBench:长上下文大语言模型能否理解在线讨论串?

LongSocialBench: Do Long-Context LLMs Understand Online Discussion Threads?

Xinyi Liu, Rinat Khaziev, Dilek Hakkani-Tür, Tarek F. Abdelzaher

首次发表
浏览论文内容

中文总结 AI 辅助

提出LongSocialBench基准,含1462道人工验证题,评估长上下文模型对在线讨论串中结构化社会话语的理解,发现现有模型远低于人类表现,缺失能力在于对回复树的社会理解。

中文摘要 AI 辅助

长上下文大语言模型现在能够处理整个在线讨论串,但理解其中的社会性话语需要的不仅仅是阅读长文档:模型必须追踪父回复关系、转折点、局部子树、跨分支对比以及参与者轨迹。为了测试这种结构感知的社会推理能力,我们引入了LongSocialBench,这是一个包含1,462个经过验证的人工编写的多项选择题的基准,这些题目来自94个完整的Hacker News、Stack Exchange和Reddit r/ChangeMyView讨论片段,中位长度约为73K个token。每个题目将一个完整的序列化讨论和回复结构与一个四选项问题配对,要求模型恢复结构化的社会证据。发布的题目经过可回答性、选项唯一性和证据依据的验证。在18个模型和29种评估设置中,当前的长上下文工作流仍远低于人类表现。最佳个体结果来自Claude-Opus-4.7,在回答前提示其消除错误选项时达到63.0%,而独立人类读者的准确率为72.4%。在所有18个模型的平均表现中,全上下文基线得分为43.9%。提供黄金证据范围后,这一分数提高到55.0%,表明即使在识别出相关讨论区域后,仍存在大量错误。LongSocialBench表明,缺失的能力不仅仅是上下文访问或提示,而是对结构化回复树的社会理解。

英文摘要

Long-context LLMs can now ingest entire online discussion threads, but understanding their social discourse requires more than reading a long document: models must track parent-reply relations, turning points, scoped subtrees, cross-branch contrasts, and participant trajectories. To test this structure-aware social reasoning, we introduce LongSocialBench, a benchmark of 1,462 verified human-authored multiple-choice items drawn from 94 complete Hacker News, Stack Exchange, and Reddit r/ChangeMyView episodes, with a median length of approximately 73K tokens. Each item pairs a complete serialized discussion and reply structure with a four-option question, requiring models to recover structured social evidence. Released items are verified for answerability, option uniqueness, and evidence grounding. Across 18 models and 29 evaluation settings, current long-context workflows remain far below human performance. The best individual result comes from Claude-Opus-4.7, which reaches 63.0% when prompted to eliminate incorrect options before answering, compared with 72.4% for independent human readers. Averaged across all 18 models, the full-context Baseline scores 43.9%. Supplying the gold evidence scope raises this to 55.0%, showing that substantial errors remain even after the relevant thread region is identified. LongSocialBench shows that the missing capability is not context access or prompting alone, but social understanding over structured reply trees.

发表机构

  • Semantic Code Lab(语义代码实验室)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑