arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2610.04693cs.AIcs.CLcs.CR

Penumbra:面向监管义务的样本高效对抗性搜索

Penumbra: Sample-Efficient Adversarial Search for Regulatory Obligations

Anthony Rhodes

首次发表
浏览论文内容

中文总结 AI 辅助

针对监管义务的对抗性搜索,Penumbra通过自适应分配和编辑预算扩展,以样本高效方式找到合规与违规边界上的相邻回复对,显著提升失败面覆盖率。

中文摘要 AI 辅助

智能体正进入金融、医疗和法律等领域,在这些领域中,违规行为不会留下词汇层面的痕迹,却会带来实际的处罚。遗漏是否重大,披露是否充分,取决于回复遗漏了什么。探查此类义务意味着寻找那些仅需一次最小编辑即可翻转合规状态的回复,而每次探查都需消耗一次生成和两次判定,因此监管红队测试的约束条件是样本效率,而非数量。我们提出Penumbra,一种对抗性搜索方法,它在不断扩大的编辑预算下从已验证的锚点出发,直到双评估委员会改变其裁决,并输出跨越该变化的两条相邻回复。分配是自适应的,目标是覆盖失败面:每个候选者所解决的不同的(义务×失败模式)单元。在匹配的预算下,自适应分配在59%的候选者上达到均匀分配的全预算覆盖率;在相同记录数下,它覆盖的失败模式是朴素枚举的1.43倍,且增益局限于其针对的轴。在对金融咨询章程的60项筛选义务中,Penumbra返回144对回复,每对包含一条合规和一条违规回复,委员会将它们置于边界两侧,它们仅相差几个词,而直接要求模型同时生成两者时,产生的文本几乎毫无共同之处。另一份用于临床分诊的章程在18项义务上重现了这一点:以相同的紧密度得到49对,且最难的模式相同。此类配对展示了义务自身条款何时不再起决定作用,这正是部署在该义务下的智能体必须被测试的内容,而搜索以随边界而非文本规模增长的成本找到它们。

英文摘要

Agents are entering finance, healthcare and law, sectors where a violation leaves no lexical signature and carries real penalties. Whether an omission is material, or a disclosure sufficient, depends on what the response left out. Probing such an obligation means finding responses one minimal edit from flipping compliance, and every probe costs a generation and two adjudications, so the binding constraint on regulatory red-teaming is sample efficiency, not volume. We introduce Penumbra, an adversarial search that walks from a verified anchor under an expanding edit budget until a two-evaluator committee changes its verdict, and emits the two adjacent responses that straddle the change. Allocation is adaptive, and the objective is coverage of the defeat surface: distinct (obligation x defeat mode) cells resolved per candidate. At matched budget, adaptive allocation reaches uniform allocation's full-budget coverage on 59% of the candidates; at equal records it covers 1.43x the defeat modes of naive enumeration, and the gain is confined to the axis it targets. On 60 screened obligations of a financial advisory constitution, Penumbra returns 144 pairs, each a compliant and a violating response that the committee places on opposite sides of the boundary, differing by a handful of words where a model asked for both directly produces texts sharing almost nothing. A second constitution, for clinical triage, reproduces this on 18 obligations: 49 pairs at the same tightness, with the same modes hardest. Pairs like these show where an obligation's own terms stop deciding, which is what an agent deployed under it must be tested against, and the search finds them at a cost that scales with the boundary, not the text.

发表机构

  • Confidential Core AI

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑