发表机构
Brock University; Emory University(布鲁克大学; 埃默里大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
WordPolo通过语义反馈找词任务,评估LLM和LRM的迭代推理与搜索策略,引入进展指标揭示模型能力,强调过程与结果并重的基准测试。
AI 中文摘要
大型语言模型(LLMs)和大型推理模型(LRMs)通常在具有挑战性的基准测试中仅通过数据集准确率进行评估,这无法提供对其推理过程质量或忠实度的任何洞察。我们提出了WordPolo,一个找词任务,参与者必须利用语义相似度反馈来发现一个未知的目标词。玩家从零知识开始,进行猜测,并接收距离分数(1=正确,数值越大表示越远)。成功需要解读分数以在语义空间中导航并系统地缩小搜索范围。这种设计使得迭代推理和自适应搜索策略既可直接观察,又是成功所必需的。我们在1,500个谜题上评估了最近的LLMs(GPT-4.1、Llama 4、Claude 3.5 Haiku、Qwen 3)、LRMs(o4-mini、Deepseek-R1)、人类以及一种新颖的启发式方法。除了解决率(范围从4%到62%)之外,我们引入了基于进展的指标,这些指标揭示了模型经常取得有意义的进展,而这些见解仅靠准确率是无法获得的。我们的分析表明,推理模型可能因过度思考和思考不足而受到阻碍,而成功的模型则表现出类似人类的策略。WordPolo证明了需要同时测试推理过程和结果的基准测试,提供对模型能力的全面衡量。我们的代码和数据集可在https://这个URL找到。
英文摘要
Large Language Models (LLMs) and Large Reasoning Models (LRMs) are typically evaluated on challenging benchmarks through dataset accuracy alone, providing no insight into the quality or faithfulness of their reasoning processes. We present WordPolo, a word-finding task where participants must discover an unknown target word using semantic similarity feedback. Players start with zero knowledge, make guesses, and receive distance scores (1 = correct, higher = further away). Success requires interpreting scores to navigate semantic space and systematically narrow the search. This design makes iterative reasoning and adaptive search strategies both directly observable and necessary for success. We evaluate recent LLMs (GPT-4.1, Llama 4, Claude 3.5 Haiku, Qwen 3), LRMs (o4-mini, Deepseek-R1), humans, and a novel heuristic on 1,500 puzzles. Beyond solve rates (which range from 4% to 62%), we introduce progression-based metrics that reveal models often make meaningful progress, insights that accuracy alone would miss. Our analysis shows how reasoning models can be hindered by overthinking and underthinking, while successful models exhibit human-like strategies. WordPolo demonstrates the need for benchmarks that test both reasoning process and outcomes, providing holistic measurements of model capabilities. Our code and dataset can be found at https://wordpolo-demo.vercel.app/.