发表机构
Manulife(宏利金融)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本文提出TAM基准,评估语言模型在长程程序性推理中的能力,发现其在ICD-10-CM编码和联邦量刑任务上精确匹配率极低,揭示现有基准高估了模型推理能力。
AI 中文摘要
大型语言模型(LLMs)在广泛的自然语言任务上取得了强劲的性能,最近的基准测试表明它们越来越擅长多跳推理。然而,这些基准测试通常是短程的,只需要少量的检索或推理步骤,并且对于涉及遵循跨越数百页、具有复杂且相互依赖的指南的真实世界任务,它们提供的可靠性证据有限。在本文中,我们引入了应用程序手册上的任务(TAM),这是一个用于评估长程程序性推理的基准。我们通过从两个领域策划真实世界任务来构建TAM:ICD-10-CM临床编码(将医疗状况映射到诊断代码)和美国联邦量刑(计算犯罪量刑指南结果,特别是犯罪等级),并带有经过人工验证的标签。每个任务都需要遵循包含数万条规则的权威手册,并跨不同部分执行一系列相互依赖的步骤以产生精确答案。我们评估了通用提示方法,包括检索增强生成、ReAct风格提示和基于GPT-5的智能体工具基线,发现最佳精确匹配性能仍然极低:在ICD-10-CM编码上为1%,在量刑任务上为15.5%。这些结果表明,当前的基准测试可能高估了LLM的推理能力,并遗漏了一个关键挑战:可靠地遵循冗长的、基于规则的程序。完整的TAM数据和代码可公开获取。
英文摘要
Large language models (LLMs) have achieved strong performance on a wide range of natural language tasks, and recent benchmarks suggest that they are increasingly adept at multi-hop reasoning. However, these benchmarks are typically short-horizon, requiring only a small number of retrieval or inference steps, and provide limited evidence of reliability on real-world tasks that involve following manuals spanning hundreds of pages with complex, interdependent guidelines. In this paper, we introduce Tasks over Application Manuals (TAM), a benchmark for evaluating long-horizon procedural reasoning. We construct TAM by curating real-world tasks from two domains: ICD-10-CM clinical coding (mapping medical conditions to diagnostic codes) and U.S. federal sentencing (computing crime sentencing guideline outcomes, specifically offense levels), with human-validated labels. Each task requires following an authoritative manual with tens of thousands of rules and executing a sequence of interdependent steps across different sections to produce an exact answer. We evaluate general-purpose prompting approaches, including retrieval-augmented generation, ReAct-style prompting, and an agent-harness baseline on GPT-5, and find that the best exact-match performance remains extremely low: 1% on ICD-10-CM coding and 15.5% on sentencing tasks. These results show that current benchmarks may overestimate LLM reasoning ability and miss a key challenge: reliably following long, rule-based procedures. The complete TAM data and code are publicly available.