arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

从BERT到前沿智能体:八年语言模型进展、能力-成本曲线的崩溃及任务定向模型的兴起

From BERT to Frontier Agents: Eight Years of Language-Model Progress, the Collapse of the Capability-Cost Curve, and the Rise of Task-Targeted Models

Pranav Kumar Kaliaperumal

arXiv 2608.13675首次发表:更新:

AI 中文总结

该文梳理2018-2026年语言模型进展,发现2024年末后编码能力年提升近6倍、能力-成本曲线崩溃,专用模型成趋势,Qwen 2.5数学测试及置信度工具验证了相关结论,研究材料全公开。

AI 中文摘要

2018年10月至2026年7月间,AI模型从BERT这类简单系统发展为能解决复杂数学问题、编写软件的大型智能体。自2024年末以来,解决实际编码问题的能力每年提升近6倍;同期成本大幅下降,OpenAI的预算模型GPT 5.6 Luna以仅1至6美元每百万token的成本达到旗舰级能力,以极低价格击败旧版本。当前顶尖性能分散在各类专用模型中:Claude Opus 5在前端编码领域领先,Claude Fable 5在仓库级编码中表现出色,GPT 5.6 Sol则在终端任务中占据主导。在使用Qwen 2.5模型的小学数学测试中,基础方法解决了100道题中的58道,高级采样方法解决了多达79道;置信度排名工具在其前50个选项中正确识别了47道正确答案,证明其对分类任务极具实用性,所有研究材料均已完全公开。

英文摘要

Between October 2018 and July 2026 AI models progressed from simple systems like BERT to massive agents that solve complex math and write software. The ability to resolve real coding issues improved by nearly six times per year since late 2024. During this time costs dropped sharply with OpenAIs budget model GPT 5 point 6 Luna matching flagship capabilities for just one to six dollars per million tokens beating older versions at a fraction of the price. Top performance is now split across specialized models as Claude Opus 5 leads in frontend coding Claude Fable 5 excels at repository level coding and GPT 5 point 6 Sol dominates terminal tasks. In a grade school math test using the Qwen 2 point 5 model basic methods solved 58 of 100 problems while advanced sampling solved up to 79. A confidence ranking tool correctly identified 47 right answers in its top 50 choices proving highly useful for sorting tasks with all research materials made fully public.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑