Omanic: Towards Step-wise Evaluation of Multi-hop Reasoning in Large Language Models
Omanic:迈向大语言模型多跳推理的逐步评估
机构 * The University of Tokyo(东京大学) ; Yale University(耶鲁大学) ; Stanford University(斯坦福大学) ; Xiaomi EV(小米EV) ; Soongsil University(顺天大学)
AI总结 针对大语言模型在多跳问答中中间步骤推理失败难以诊断的问题,提出Omanic基准,通过分解为单跳子问题并分析步骤级错误,揭示后期跳数瓶颈、事实知识下限和错误传播,微调后提升多个推理基准性能。
Comments EMNLP 2026 Findings