arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2610.10629cs.AI

仅将控制流作为唯一可变表面:信贷流程中LLM智能体的合规性受限自演化及带测量的准入门

The Harness as the Only Mutable Surface: Compliance-Bounded Self-Evolution of LLM Agents in Credit Pipelines, with a Measured Admission Gate

Ravil Akhtyamov

首次发表
浏览论文内容

中文总结 AI 辅助

该研究提出将LLM智能体的自演化限制在运行时控制流内,构建带哈希链记录准入门的双循环引擎,实验显示该门可减少有害变更,且符合欧盟AI法案相关要求。

中文摘要 AI 辅助

自改进的大语言模型(LLM)智能体可使信贷流程适应变更后的规则,但自行重写的智能体会破坏主管审核的人工制品:即已命名的变更、记录的测试及审批。我们认为,仅当自演化被限制在运行时控制流(指令文本、工具调用逻辑及原语组合)内且模型权重保持固定时,该自演化才是可审核的,如此每项适配均为附带原因与测试的差异。我们基于该约束构建了双循环引擎,其中一个准入门会在部署前写入哈希链记录,并在模拟环境中对该门进行测量,采用模拟智能体及种子搜索提议器而非语言模型。在三种严重程度的三类监督重新解释中,每类各10个种子,带门限的循环在7449个候选变更中准入144个,且均未在保留历史数据上恶化误差,在所有低、中严重程度单元中均将误报率恢复至 oracle 水平且未提高漏报率。若将门替换为无约束系统应用的检查(近期轨迹中可见误差更少),相同循环准入了309个有害变更,且90次运行中有49次漏报率高于10%:误报下降是因为筛查被放宽。在预变更标签上评估时,该门拒绝了所有候选变更,因此重新解释必须编码为可重新标记历史的规则。参数与范围变更可局部修复,结构性变更仅可通过原语替换修复;在最高结构性严重程度下,该门的固定容忍度在一半种子中阻止了正确替换。我们将这些机制映射至欧盟《人工智能法案》中高风险信贷评分的条款,并指出2026年4月美国模型风险指南未将智能体AI纳入其范围。

英文摘要

Self-improving LLM agents can adapt a credit pipeline to a changed rule, but an agent that rewrites itself destroys the artefact a supervisor reviews: a named change, a recorded test, an approval. We argue that self-evolution is reviewable only if it is confined to the runtime harness (instruction text, tool-call logic and primitive composition) while model weights stay fixed, so that every adaptation is a diff with a cause and a test attached. We give a dual-loop engine built on that bound, with one admission gate that writes a hash-chained record before deployment, and we measure the gate in simulation, with a simulated agent and a seeded-search proposer rather than language models. Across three families of supervisory re-interpretation at three severities, 10 seeds each, the gated loop admitted 144 of 7,449 candidate changes, none of which worsened error on held-out history, and restored the false-positive rate to the oracle level without raising missed flags in every low- and mid-severity cell. With the gate replaced by the check an unbounded system applies (fewer errors visible in recent traces), the same loops admitted 309 harmful changes and left missed flags above 10% in 49 of 90 runs: false positives fell because the screen was loosened. Evaluated on pre-shift labels, the gate rejected every candidate, so a re-interpretation must be encoded as a rule that relabels history. Parametric and scope shifts were repaired locally, a structural one only by primitive replacement; at the highest structural severity the gate's fixed tolerance blocked the correct replacement in half the seeds. We map the mechanisms to the EU AI Act's provisions for high-risk credit scoring and note that the April 2026 US model-risk guidance excludes agentic AI from its scope.

发表机构

  • Digital Economy Lab(数字经济实验室)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑