arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

从主观判断到可审计标准:协议引导的网站冗余度AI审计

From Subjective Judgments to Auditable Standards:Protocol-Guided AI Auditing of Website Redundancy

Ge Kong, Yongtong Cao

arXiv 2608.21476首次发表:更新:

发表机构

Beihang University; Beijing Institute of Technology(北京航空航天大学; 北京理工大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本文提出CORA审计方法,通过多维度测量分离网站冗余的不同含义,在测试台上表现优于基线,对不达标工具弃权自动评分,是受控基准的可审计候选程序。

AI 中文摘要

网站冗余度并非单一固定含义:同一重复元素可能在一项任务中造成干扰,却在另一项任务中提供备份。本文提出CORA(反事实可观测冗余审计),该方法分别测量重复负载、正常使用税和故障域恢复储备,每次运行都会保留截图、稳定元素标识和任务轨迹。一个带版本控制的视觉语言模型会生成标注,随后通过类型化验证和发布检查确定是否可报告校准维度;失败或格式错误的输出会保留在固定分母中。在透明机制测试台上,因子化的CORA表示将恢复储备与正常使用税分离,且比标量负载基线更准确地预测扰动成功。模型研究进一步表明,仅可重复性是不够的:两个小型本地视觉语言模型生成了可重复输出,但均未满足所有发布要求。因此CORA对这两个工具弃权(不执行)自动评分,同时保留原始响应和故障记录。独立检查工具确认类型化验证器和加固发布闸口符合其规范;这些测试未建立语义基础或生产网站上的准确性。总体而言,研究结果表明CORA是所研究受控基准的可审计候选程序,而非通用标准。人类一致性、AI与人类的准确性以及独立生产网站上的验证仍是未解决的实证问题。

英文摘要

Website redundancy does not have a single fixed meaning. The same repeated element may distract during one task and provide backup during another. We introduce CORA (Counterfactual, Observable Redundancy Audit), which measures repetition load, normal-use tax, and failure-domain recovery reserve separately. Each run retains screenshots, stable element identities, and task traces. A versioned vision-language model proposes the annotations. Typed validation and release checks then determine whether a calibrated dimension can be reported; failed or malformed outputs stay in the fixed denominator. On a transparent mechanistic testbed, the factorized CORA representation separated reserve from normal-use tax and predicted perturbed success more accurately than scalar-load baselines. The model studies then showed why repeatability is not enough: two small local vision-language models produced recurring outputs, but neither instrument met all release requirements. CORA therefore withheld automated scores from both instruments while retaining the raw responses and failure records. Separate checker fixtures confirmed that the typed validator and hardened release gates implement their specifications; these tests do not establish semantic grounding or accuracy on production sites. Taken together, the results position CORA as an auditable candidate procedure for the controlled benchmark studied here rather than a general standard. Human agreement, AI-versus-human accuracy, and validation on independent production sites remain open empirical questions.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑