arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

大型语言模型(LLMs)能否测试终端用户界面?

Can LLMs Test Terminal User Interfaces?

Chao Peng, Ruida Hu, Ajitha Rajan, Tegawendé F Bissyandé, Jacques Klein, Cuiyun Gao

arXiv 2608.03743首次发表:更新:

发表机构

University of Edinburgh; Harbin Institute of Technology; University of Luxembourg(爱丁堡大学; 哈尔滨工业大学; 卢森堡大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

该研究针对终端用户界面(TUIs)缺乏专门测试方法的问题,构建多框架基准测试对比前沿LLMs与随机探索,发现自动生成输入可提升测试效果,验证了自动化TUI测试的可行性。

AI 中文摘要

终端用户界面(TUIs)结合了图形用户界面(GUIs)的有状态、面向屏幕的特性与终端部署方式,目前在开发者工具中已广泛应用,但缺乏专门的测试方法。我们调研了197个真实世界的TUI应用:仅12%的测试代码会对界面进行操作,其中45%的测试从未发送输入,仅检查静态界面帧。我们将这些应用转化为覆盖ratatui/Rust、bubbletea/Go、textual/Python和ink/TypeScript的无头基准测试,每个都打包为带检测的Docker镜像。我们记录了可靠场景下的代码行和组件(widget)覆盖率、渲染的终端状态以及崩溃情况。在相同的 wall-clock 时间预算下,我们对比了四个前沿大型语言模型(LLMs)与随机探索的表现,没有任何模型占据绝对优势。随机探索是一个强大的时间预算基准,但它的崩溃发现优势来自更高的吞吐量;而每次交互中,LLM引导的效率更高,且能唯一到达由输入控制的故障点。自动生成启动输入带来了最大的实际收益,可启动原本无法启动的应用。代码行覆盖率无法有效预测崩溃发现,削弱了其作为测试有效性代理指标的作用。自动化TUI测试是可行的,但远未解决,且可靠的基准比模型选择更重要。我们发布了覆盖率工具tuicov和测试框架tuibot。

英文摘要

Terminal User Interfaces (TUIs) combine the stateful, screen-oriented behaviour of GUIs with terminal deployment and are now common in developer tools. Yet they lack a dedicated testing methodology. We survey 197 real-world TUI applications: only 12% of test code exercises the interface, and 45% of those tests never send input, checking a static frame instead. We turn these applications into a headless benchmark spanning ratatui/Rust, bubbletea/Go, textual/Python, and ink/TypeScript, packaging each as an instrumented Docker image. We record line and widget coverage where reliable, rendered terminal states, and crashes. Under equal wall-clock budgets, we compare four frontier LLMs with random exploration. No model dominates. Random is a strong time-budgeted baseline, but its crash advantage comes from higher throughput: per interaction, LLM guidance is more efficient and uniquely reaches input-gated faults. Automatically deriving launch inputs yields the largest practical gain, enabling applications that otherwise never start. Line coverage poorly predicts crash discovery, weakening it as a proxy for test effectiveness. Automated TUI testing is feasible but far from solved, and honest baselines matter more than model choice. We release the coverage tool tuicov at https://github.com/tui-testing/tuicov and the testing framework tuibot at https://github.com/tui-testing/tuibot.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑