arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

Pseudo2CodeQA:面向基于大语言模型的代码生成中结构化算法推理的基准测试

Pseudo2CodeQA: A Benchmark for LLM-Based Structured Algorithmic Reasoning in Code Generation

Shadikur Rahman, Umme Ayman Koana, Syed Muhammad Danish

arXiv 2608.09068首次发表:更新:

AI 中文总结

本文提出Pseudo2Code基准测试及Pseudo2Code Agentic Framework,实验表明该框架性能优于现有基线,验证了结构化伪代码对代码生成的提升作用。

AI 中文摘要

大语言模型(LLMs)在自然语言转代码生成方面已取得令人瞩目的性能,但它们遵循结构化算法推理的能力仍未得到充分研究。本文引入Pseudo2Code,这是一个旨在系统评估结构化伪代码对代码生成质量和算法忠实性影响的基准测试。该基准测试包含300个经人工验证的真实编程任务,涵盖多个领域及简单、中等、困难三个难度等级,每个任务包含问题描述、结构化伪代码、参考实现和可执行测试套件。为确保基准测试的可靠性,本文采用了双阶段人工验证协议,并发布了完全可执行的基准测试实例。除基准测试外,本文还提出了Pseudo2Code Agentic Framework,这是一个多阶段流程,利用伪代码作为代码生成的显式中间推理表示。本文使用基于评分标准的评估框架对商业和开源语言模型进行评估,该框架衡量正确性、完整性、相关性、清晰度、推理质量和伪代码 adherence(贴合度),并辅以基于执行的测试。实验结果表明,所提出的Pseudo2Code Agentic Pipeline在性能上始终优于强大的商业和开源基线模型,总分为4.78,而最强基线模型的总分为4.31。此外,一项涉及100个基准测试任务的人工评估研究显示,人工判断与自动评估之间存在高度一致性。本文的研究结果提供了实证证据,证明结构化伪代码可提升代码生成中的功能正确性、推理质量和算法忠实性。本文发布Pseudo2Code,以支持未来关于结构化推理、可解释代码生成和可靠AI辅助软件开发的研究。

英文摘要

Large Language Models (LLMs) have achieved impressive performance in natural language-to-code generation; however, their ability to follow structured algorithmic reasoning remains insufficiently understood. We introduce Pseudo2Code, a benchmark designed to systematically evaluate the impact of structured pseudocode on code generation quality and algorithmic faithfulness. The benchmark consists of 300 manually validated real-world programming tasks spanning multiple domains and three difficulty levels (Easy, Medium, and Hard). Each task contains a problem description, structured pseudocode, reference implementation, and executable test suite. To ensure benchmark reliability, we adopt a dual-stage human validation protocol and release fully executable benchmark instances. Beyond the benchmark, we propose the Pseudo2Code Agentic Framework, a multi-stage pipeline that leverages pseudocode as an explicit intermediate reasoning representation for code generation. We evaluate both commercial and open-source language models using a rubric-based evaluation framework that measures correctness, completeness, relevance, clarity, reasoning quality, and pseudocode adherence, complemented by execution-based testing. Experimental results demonstrate that the proposed Pseudo2Code Agentic Pipeline consistently outperforms strong commercial and open-source baselines, achieving an overall score of 4.78 compared to 4.31 for the strongest baseline model. Furthermore, a human evaluation study involving 100 benchmark tasks shows strong agreement between human judgments and automated assessments. Our findings provide empirical evidence that structured pseudocode improves functional correctness, reasoning quality, and algorithmic faithfulness in code generation. We release Pseudo2Code to support future research on structured reasoning, interpretable code generation, and reliable AI-assisted software development.

Comments8 pages, 3 figures

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑