TurnBench-MS: A Benchmark for Evaluating Multi-Turn, Multi-Step Reasoning in Large Language Models
TurnBench-MS: 一个评估大语言模型多轮多步推理能力的基准
AI总结 TurnBench-MS通过互动破码任务评估大语言模型的多轮多步推理能力,揭示当前模型在复杂推理任务中的显著不足。
Comments Accepted to Findings of the Association for Computational Linguistics: EMNLP 2025
Journal ref Findings of the ACL: EMNLP 2025, pp. 19892-19924, 2025