用于跨语言手写OCR的基于大语言模型驱动的自动化机器学习:使用GPT-5、GPT-4o和Claude十四行诗4的闭环神经架构搜索
LLM-Driven AutoML for Cross-Lingual Handwritten OCR: Closed-Loop Neural Architecture Search with GPT-5, GPT-4o, and Claude Sonnet 4
浏览论文内容
中文总结 AI 辅助
研究针对跨语言手写OCR,提出用GPT-5等大语言模型驱动的自动化机器学习框架,能自主生成、训练等优化神经网络架构,经实验评估,该框架无需人工设计等就能发现高效准确模型,实现跨语言手写识别。
中文摘要 AI 辅助
我们提出了一个完全自动化的闭环自动化机器学习框架,该框架使用GPT-5、GPT-4o和Claude十四行诗4作为跨语言手写光学字符识别的自主神经架构设计器。每个大语言模型利用先前试验的性能反馈,独立生成、训练、评估并迭代优化神经网络架构。该框架通过270个独立实验在阿拉伯语、波斯语和英语手写数据集上进行评估。它无需人工架构设计、特定领域预处理或超参数调整,就能持续发现准确且计算高效的模型。生成的模型平均测试准确率超过93%,最佳准确率为98.1%,推理延迟在41到44毫秒之间。结果表明大语言模型可作为神经架构搜索的有效自动化机器学习代理,实现跨语言可扩展、脚本自适应且可重现的手写识别。
英文摘要
We present a fully automated closed-loop AutoML framework that uses GPT-5, GPT-4o, and Claude Sonnet 4 as autonomous neural architecture designers for cross-lingual handwritten optical character recognition. Each large language model independently generates, trains, evaluates, and iteratively refines neural network architectures using performance feedback from previous trials. The framework is evaluated on Arabic, Persian, and English handwriting datasets through 270 independent experiments. It consistently discovers accurate and computationally efficient models without manual architecture design, domain-specific preprocessing, or hyperparameter tuning. The generated models achieve mean test accuracies above 93 percent, a best accuracy of 98.1 percent, and inference latency between 41 and 44 milliseconds. The results demonstrate that large language models can function as effective AutoML agents for neural architecture search, enabling scalable, script-adaptive, and reproducible handwriting recognition across languages.