发表机构
HKUST; HKBU; NTU(香港科技大学; 香港浸会大学; 新加坡南洋理工大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
研究针对LLM智能体在金融合规中的规则落地问题,构建ReguSim环境与ReguBench基准,发现可见规则无法完全消除违规动作,结构化监控基线表现优于纯提示LLM,为金融合规评估提供新框架。
AI 中文摘要
金融市场中的大语言模型(LLM)智能体可能引用规则,但仍会提交违反可执行约束的订单或误读监控证据。我们引入ReguSim(受控金融合规环境)和ReguBench(带目标标记的监控基准),以分离四类产物:陈述性推理、尝试性动作、执行强制、监控证据。在使用DeepSeek V4 Pro和Gemini 3.5 Flash的交易者运行中,可见规则减少但未消除被拒绝的动作,激励或角色设定会改变行为。一项桥接研究显示,除非展示执行证据,否则交易者的理由可能误导独立监控。在监控方面,简单结构化基线的表现要么匹配要么超过仅用提示的LLM。这些结果将金融合规评估定义为对基于规则的动作和证据使用的审计,而非单一合规分数。
英文摘要
LLM agents in financial markets may cite rules yet still submit orders that violate executable constraints or misread surveillance evidence. We introduce ReguSim, a controlled financial-compliance environment, and ReguBench, a target-marked monitoring benchmark, to separate four artifacts: stated reasoning, attempted action, execution enforcement, and monitor evidence. In trader runs with DeepSeek V4 Pro and Gemini 3.5 Flash, visible rules reduce but do not eliminate rejected actions, and incentive or persona framing shifts behavior. A bridge study shows that trader rationales can mislead an independent monitor unless enforcement evidence is shown. In monitoring, simple structured baselines either match or exceed prompt-only LLMs. The results frame financial compliance evaluation as an audit of rule-grounded actions and evidence use, rather than a single compliance score.