arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

DeFiFlowBench:自然语言DeFi工作流合成中安全可执行性的基准测试与改进

DeFiFlowBench: Benchmarking and Improving Safe Executability in Natural-Language DeFi Workflow Synthesis

Abhinav Rajeev Kumar, Harshit Arora, Varun Singh, Manikandan Nanjappan

arXiv 2609.11504首次发表:更新:

发表机构

SRM Institute of Science and Technology(SRM科学与技术学院)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

提出DeFiFlowBench基准和Koan-Safe方法,通过明确交易保护与执行评估,提升自然语言DeFi工作流合成的安全可执行性。

AI 中文摘要

一个结构上有效的DeFi工作流仍可能授权一笔代价高昂的交易。我们引入了DeFiFlowBench,一个包含207个团队编写提示词的自然语言DeFi工作流合成基准。该基准衡量图覆盖度、配置完整性和声明的安全谓词,然后在本地EVM上测试支持的交易配置。在固定的5%价格影响上限下,直接提示、约束提示和少样本提示每种配置分别产生14-19次不安全的留出执行。从报价中推导出的滑点界限并不能防止订单本身的价格影响。我们提出了Koan-Safe,它结合了仅提示的意图解析器、可替换的生成器以及带有默认安全参数的结构修复。在75个留出工作流提示上,其混合变体在静态安全代理上得分为0.67,而最佳基线为0.33。Koan-Safe在保存的基准输出上记录了零次不安全执行。一项匹配候选消融实验在禁用强制措施时产生了14-17次不安全执行。额外测试暴露了默认注入的局限性:宽松的现有阈值仍可能授权不安全的交易。一个单独评估的策略上限在36个案例的诊断网格上解决了这一失败。这些结果支持明确的交易保护和基于执行的评估,同时将声明的安全性与一般保证区分开来。

英文摘要

A structurally valid DeFi workflow can still authorize a costly trade. We introduce DeFiFlowBench, a benchmark of 207 team-authored prompts for natural-language DeFi workflow synthesis. It measures graph coverage, configuration completeness, and declared safety predicates, then tests supported trade configurations on a local EVM. Direct, constrained, and few-shot prompting produce 14-19 unsafe held-out executions per configuration under a fixed 5% price-impact cap. A slippage bound derived from a quote does not prevent the price impact of the order itself. We propose Koan-Safe, which combines a prompt-only intent parser, a replaceable generator, and structural repair with default safety parameters. On 75 held-out workflow prompts, its hybrid variant scores 0.67 on the static safety proxy, compared with 0.33 for the best baseline. Koan-Safe records no unsafe executions on the saved benchmark outputs. A matched-candidate ablation produces 14-17 unsafe executions when enforcement is disabled. Additional tests expose the limits of default injection: permissive existing thresholds can still authorize unsafe trades. A separately evaluated policy cap addresses this failure on a 36-case diagnostic grid. These results support explicit trade protections and execution-based evaluation, while distinguishing declared safety from a general guarantee.

CommentsCode and benchmark: https://github.com/Varun-2538/Koan

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑