ISA-Bench:跨指令集架构的计算推理基准
ISA-Bench: A Benchmark for Computational Reasoning Across Instruction Set Architectures
- IIT Hyderabad(印度理工学院海得拉巴分校)
- Microsoft Research, India(微软研究院(印度))
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
ISA-Bench通过受限指令集的编程游戏基准,评估大型语言模型在陌生计算模型中的推理能力,发现推理模型表现更优,但语法不熟悉是主要失败原因,并引入REG分析揭示策略与表达间的脱节。
AI中文摘要:
大型语言模型的代码生成基准主要评估资源丰富的语言,如Python和Java,这些语言中模型受益于大量的训练数据。它们对于在陌生计算模型中的推理提供的证据有限:从单条减法指令推导算术、协调跨通信节点的并行程序,或将逻辑门连接成电路。我们提出了ISA-Bench,一个具有受限指令集的编程游戏基准。对于每个游戏,我们提供了完整的执行栈(解析器、虚拟机、验证器),支持自动化评估并提供结构化反馈以进行迭代改进。推理模型在平均求解率上高于代码专用和通用模型,但不熟悉的语法仍然是失败的主要来源。模型在迭代反馈下能解决更多任务,尽管不同架构之间的收益差异显著。我们引入了推理-执行差距(REG)分析,揭示了识别可行计算策略与将其表达为目标ISA中的正确程序之间反复出现的脱节。代码已开源。
英文摘要:
Large language model code generation benchmarks primarily evaluate well-resourced languages like Python and Java, where models benefit from abundant training data. They provide limited evidence about reasoning in unfamiliar computational models: deriving arithmetic from a single subtract instruction, coordinating parallel programs across communicating nodes, or wiring logic gates into circuits. We present ISA-Bench, a benchmark of programming games with constrained instruction sets. For each game we provide a full execution stack (parser, VM, and verifier), enabling automated evaluation with structured feedback for iterative refinement. Reasoning models achieve higher average solve rates than code-specialized and general-purpose models, but unfamiliar syntax remains a major source of failure. Models solve more tasks with iterative feedback, though the gains vary substantially across architectures. We introduce a reasoning--execution gap (REG) analysis that reveals a recurring disconnect between identifying a plausible computational strategy and expressing it as a correct program in the target ISA. Code is open-sourced.