arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

FaultLens:为生成的操作程序学习紧凑的行为测试套件

FaultLens: Learning Compact Behavioral Test Suites for Generated Operational Programs

Zeming Liu, Hang Lyu, Jingtao Zhang

arXiv 2608.26746首次发表:更新:

AI 中文总结

FaultLens是一种学习紧凑行为测试套件的方法,通过混合故障驱动贪心组件与多样性组件,在低计算成本下实现高故障覆盖率,可降低尾部回归风险。

AI 中文摘要

生成的操作程序通常通过少量手写示例或穷尽回归套件进行验证。前者可能会遗漏稀疏的边界和交互故障,而后者可能会产生不必要的高成本。我们提出FaultLens,一种在保留与执行证据可审计连接的同时学习紧凑行为测试套件的方法。它仅执行一次丰富的探针域,将故障-探针杀伤关系存储为稀疏结果缓存,并仅从早期程序世代中学习探针排序。一个故障驱动的贪心组件利用已知的杀伤结构,而一个与突变无关的多样性组件覆盖探针家族、案例、模板和时间区间。当新程序包含排序构建中不存在的故障机制时,它们的交替混合仍然有用。我们在四个环境、十个执行种子、1200个测量运行摘要、2160个受控程序转换和4120200个执行的程序-探针对上评估了20个生成的操作策略。在1960个预期的故障转换中,1779个在有限审计域内的某个地方改变了合约或输出;另外200个对照保留了行为。在第1-3代上学习的32个探针混合体使用穷尽域的1.2-2.0%覆盖了第4-5代中576/582(99.0%)的可动态杀伤故障。在整个故障家族被排除在训练之外时,多样性将场景-家族宏观覆盖率从84.6%提高到94.9%。在下游部署研究中,一个保守的准入规则将严重的尾部回归从15/20程序-环境组减少到0/20。FaultLens提供了一种优先证据机制,而非正确性证明,并明确了其预算、证据源、泛化分割和遗漏情况。

英文摘要

Generated operational programs are often validated with either a few hand-written examples or exhaustive regression suites. The former can miss sparse boundary and interaction faults, while the latter can be unnecessarily expensive. We introduce FaultLens, a method for learning compact behavioral test suites while preserving an auditable connection to executed evidence. It executes a rich probe domain once, stores the fault-probe kill relation as a sparse outcome cache, and learns probe orderings only from earlier program generations. A fault-driven greedy component exploits known kill structure, while a mutation-independent diversity component covers probe families, cases, templates, and temporal bins. Their alternating hybrid remains useful when a new program contains a fault mechanism absent from ordering construction. We evaluate twenty generated operational policies across four environments, ten execution seeds, 1,200 measured run summaries, 2,160 controlled program transformations, and 4,120,200 executed program-probe pairs. Of 1,960 intended faulty transformations, 1,779 alter a contract or output somewhere in the finite audit domain; 200 additional controls preserve behavior. A 32-probe hybrid learned on generations 1-3 covers 576/582 (99.0%) dynamically killable faults in generations 4-5 using 1.2-2.0% of the exhaustive domain. With an entire fault family withheld from training, diversity raises scenario-family macro coverage from 84.6% to 94.9%. In a downstream deployment study, a conservative admission rule reduces severe tail regressions from 15/20 program-environment groups to 0/20. FaultLens provides a prioritized evidence mechanism, not a proof of correctness, and makes its budget, evidence source, generalization split, and misses explicit.

Comments13 pages, 7 figures. Includes a complete reproducibility artifact in the source package

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑