AI 中文总结
研究针对白盒代理式XSS发现中编码代理声明不可信的问题,提出RECEIPT验证框架,通过环境隔离等手段使发现可信,经实验评估,在预算内发现多个未知漏洞,相比其他方式能确认更多真实利用且无假阳性。
AI 中文摘要
跨站脚本攻击(XSS)仍是最普遍且具破坏性的网络漏洞之一。基于大语言模型的编码代理通过结合源代码推理与对运行中应用的交互式测试,为XSS发现提供了一种有前景的方法。然而,编码代理的声明本身不可信。我们刻画了白盒代理式XSS发现中的三种奖励攻击行为,并提出理想验证器应满足的三个要求。我们展示了RECEIPT,一个通过强制环境隔离、漏洞证明约束、角色分离和裁决绑定,使代理报告的XSS发现可信的验证框架。每个确认建立两个属性:脚本在真实浏览器中运行,有效载荷在攻击者角色下植入并在受害者角色的浏览器中执行。这种受限的重放过程使验证具有确定性和可重复性。我们在从流行开源项目中抽取的95个真实世界的Web应用目标上评估了RECEIPT。在每个应用20美元的预算内,RECEIPT发现了24个以前未知的XSS漏洞,其中12个在负责任披露后已得到维护者认可,并且在36%的已知漏洞恢复目标中恢复了标记的CVE。与使用自我判断的同一代理和黑盒扫描器相比,RECEIPT确认了更多真实利用,同时不承认有假阳性。
英文摘要
Cross-Site Scripting (XSS) remains one of the most prevalent and damaging classes of web vulnerabilities. LLM-based coding agents offer a promising approach to XSS discovery by combining source-code reasoning with interactive testing against a running application. However, a coding agent's claims cannot be trusted on their own. We characterize three reward-hacking behaviors in white-box agentic XSS discovery and propose three requirements that an ideal verifier should meet. We present RECEIPT, a verification framework that makes agent-reported XSS findings trustworthy by enforcing environment isolation, PoC constraints, role separation, and verdict binding. Each confirmation therefore establishes two properties: the script runs in a real browser, and the payload was planted under the attacker role and executed in the victim role's browser. This constrained replay procedure makes validation deterministic and reproducible. We evaluate RECEIPT on 95 real-world web-application targets drawn from popular open-source projects. Within a $20 per-application budget, RECEIPT found 24 previously unknown XSS vulnerabilities, 12 of which have already been acknowledged by maintainers after responsible disclosure, and recovered the labeled CVE in 36% of known-vulnerability recovery targets. Compared with the same agent using self-judgment and with black-box scanners, RECEIPT confirms more real exploits while admitting no false positives.