发表机构
Anvenssa AI; Pimpri Chinchwad College of Engineering(安文萨人工智能公司; 平普里钦奇瓦德工程学院)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本研究开发了一种结合Gemini LLM与定制MCTS的自主AI编码智能体,通过自评判系统优化代码生成,在复杂逻辑任务上成功率达92%,性能优于零样本生成模型。
AI 中文摘要
软件工程需求的持续变化催生了对能从自然语言输入生成安全源代码的自动化工具的巨大需求。传统大语言模型(LLM)的性能受限于其“一次生成”能力,这会导致逻辑幻觉,并在复杂操作中降低算法性能。本研究提出一种自主AI编码智能体,通过其结构化决策方法建立LLM生成内容与可投入生产的软件之间的关联。该框架使用Gemini 2.5 Flash API实现核心推理能力,同时采用定制化的蒙特卡洛树搜索(MCTS)方法,将代码生成挑战作为搜索操作解决。该智能体使用“自评判”评估系统测试不同实现方法,根据准确性和难度等级对其排序,再通过反向传播优化其运行框架。系统通过基于Flask的Web界面运行,提供即时反馈和语法高亮功能。实验结果显示,基于MCTS的方法在复杂逻辑提示上达到92%的成功率,优于标准零样本生成模型。
英文摘要
The ongoing changes in software engineering requirements have created a substantial need for automated tools which can create secure source code from natural language input. The performance of traditional Large Language Models (LLMs) becomes limited by their "one-shot" capability which results in logical hallucinations together with reduced algorithmic performance during complicated operations. The research presents an autonomous AI Coding Agent which establishes a connection between LLM-generated content and production-ready software through its organized methodology for decision making. Our framework uses the Gemini 2.5 Flash API for essential reasoning capabilities while employing a tailored Monte Carlo Tree Search (MCTS) method to solve code generation challenges as a search operation. The agent uses a "Self-Critic" evaluator system to test different implementation methods which it ranks according to their accuracy and difficulty level before it improves its operational framework through backpropagation. The system operates through a Flask-based web interface which delivers instant feedback together with syntax highlighting features. Our experimental results show that the MCTS-based method achieves a 92% success rate on complex logical prompts while surpassing standard zero-shot generation models.
Comments10 PAGES WITH PLAGIARISM REPORT ON 10TH PAGE