AI 中文总结
本研究评估了不解析shell语法的策略门在Linux上的裁决性能,发现其高准确率但存在成本问题,并提出修复方案。
AI 中文摘要
在执行代理动作之前对其进行裁决的网关,其有效性取决于对动作的理解程度。我们研究了一个从不解析shell语法的策略门:它接收类型化的、已实现的动作(动词、操作数、已解析的区域、程序对象身份),并决定允许、询问或拒绝。在我们先前研究[1]的Windows基准的Linux孪生版本上,我们考察了其裁决的忠实度、翻译后保留的内容,以及随着系统使用,裁决是否仍保持成本可承受。一个冻结的61案例表在两轮中得分61/61,无错误允许;一项50操作符变异测试杀死了50个变异体中的46个(92.0%),所有四个存活变异体均被分类。一项82行审计产生43个重新表达、21个载体差异、16个研究特定不适用行和两个未解决案例;在标记为“无对应”的25行中,四个在此仍未映射。由执行器咨询时,25个逃逸变为零,且无良性负载被阻止。在一次探索性单网关快照中,三个未认证端点标签产生案例级非拒绝多数为81.8%-98.0%,但立即执行多数为4.0%-52.5%。裁决不读取累积状态;凭证验证则读取,扫描其整个账本。我们报告了该成本及我们拟采取的修复措施。
英文摘要
Gateways that adjudicate an agent's actions before they execute are only as good as their understanding of the action. We study a policy gate that never parses shell syntax: it consumes a typed, realised action (verb, operands, resolved zones, program-object identity) and decides ALLOW, ASK or DENY. Working on a Linux twin of the Windows benchmark of our previous study [1], we ask how faithfully it adjudicates, what survives translation, and whether deciding stays affordable as the system is used. A frozen 61-case table scores 61/61 in two rounds with no false allow; a 50-operator mutation campaign kills 46 of 50 mutants (92.0%), with all four surviving mutants classified. An 82-row audit yields 43 re-expressions, 21 carrier differences, 16 study-specific inapplicable rows and two unresolved cases; of 25 rows labelled "no counterpart," four remain unmapped here. Consulted by the executor, 25 escapes become zero with no benign payload blocked. In an exploratory one-gateway snapshot, three unauthenticated endpoint labels produce case-level non-refusal majorities of 81.8-98.0%, but immediate-execution majorities of 4.0-52.5%. Adjudication reads no accumulating state; credential verification does, scanning its whole ledger. We report that cost and the fix we would make.
Comments13 pages, 8 figures, 6 tables. Supplementary material and reproducibility artifact are included as ancillary files