arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.00468cs.SE

下一代模型能保留什么?基于单提示的大语言模型技术基准测试

What Survives the Next Model? Benchmarking LLM-Based Techniques Against Single-Prompts

  • University of Virginia(弗吉尼亚大学)

机构由 AI 辅助整理,请以论文原文为准。

Nahian Salsabil, Joy Saha, Simantika Bhattacharjee Dristi, Nicholas Phair, Nusrat Jahan Mozumder, Matthew B. Dwyer, Sebastian Elbaum

AI总结:

该研究分析ICSE 2026的35篇LLM技术论文,发现37%-63%的论文中,带单提示的新一代LLM性能优于旧技术,提出需关注与未来模型协同扩展的持久挑战。

AI中文摘要:

软件工程研究界已积极将大语言模型(LLMs)整合到复杂技术中以解决各类任务。然而,这种投入的战略价值尚不明确,因为后续前沿模型的原生能力可能迅速淘汰现有技术。为评估该研究投入,我们分析了ICSE 2026的35篇基于LLM的技术论文,评估其复杂工具是否能被最简单的替代方案超越:即使用新一代模型执行的单个自动生成提示,无任何迭代优化。我们发现,37%至63%的论文中,带单个提示的新一代模型原生性能优于仅一年前提出的精心设计工具。我们还发现,代码生成或修复等建设性技术更易被单提示替代,另有部分论文依赖的策略能为模型提供额外见解,新一代LLM会放大这些技术的效果。我们的发现引发了对作为临时模型缺陷解决方案的技术的成本效益问题,以及需聚焦于与未来模型协同扩展的持久挑战的思考。我们的源代码和结果已公开于此https URL。

英文摘要:

The software engineering research community has enthusiastically embraced the integration of Large Language Models (LLMs) into complex techniques to solve a wide variety of tasks. However, the extent to which this investment is strategic remains unclear, as the native capabilities of successive frontier model generations can rapidly render existing techniques obsolete. To assess this research investment, we analyze 35 LLM-based technique papers from ICSE 2026. We evaluate whether their complex tools can be outperformed by the simplest possible alternative: a single, automatically generated prompt executed on a newer generation model, without any iterative refinement. We find that for between 37% and 63% papers, a newer model with a single prompt natively outperforms the heavily engineered tooling proposed just a year prior. We identify that constructive techniques like code generation or repair are more amenable to substitution by a single-prompt. We also identify a surviving set of papers relying on strategies that provide additional insights to the model where newer LLMs will amplify the proposed technique. Our findings raise questions about the cost-benefit proposition of techniques designed as workarounds to temporary model deficits and the need to focus on enduring challenges that scale synergistically with future model generations. Our source codes and results are made publicly available at https://github.com/less-lab-uva/What-Survives-the-Next-Model.

↑