发表机构
Lehigh University; University of California, San Diego; NYU Tandon School of Engineering; NYU Abu Dhabi(里海大学; 加利福尼亚大学圣迭戈分校; 纽约大学坦登工程学院; 纽约大学阿布扎比分校)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本文提出面向金融FPGA设计的开源基准FinHardBench,通过三类实验评估6种大语言模型,发现其在功能正确性、系统级配置优化等方面存在差异,策略级规格变更问题仍待解决。
AI 中文摘要
大型语言模型(LLMs)能否生成不仅正确,而且运行快速的硬件?本文针对金融现场可编程门阵列(FPGA)设计研究该问题,在该领域,5-10纳秒的延迟决定竞争优势,且随着协议、策略和监管规定的演变,设计需不断迭代。本文提出FinHardBench——包含33项金融计算任务的基准测试,同时开展三项模拟真实FPGA迭代周期的实验:根据规格生成新模块、在6级交易流水线中调整系统级配置、使现有模块适配规格变更。对6种大语言模型的1530余次实验轮次评估得出三项发现:(1)模型在特定任务上的功能正确率为19%-61%,时序退化最高达13.7倍;(2)在系统级设计空间探索(DSE)中,顶级大语言模型收敛到最优配置的可靠性高于随机搜索、模拟退火和贝叶斯优化基线(在相同24轮预算下,5/5随机种子达标,而基线为0-4/5);(3)策略级规格变更对大多数模型而言仍是未解决的问题。在6种模型中,代码生成与设计空间探索的排名重叠度中等:最强的代码生成器并非最快的架构优化器,而最弱的代码生成器MiniMax M2.7仍在5个随机种子中的4个上达到系统最优。在FinHardBench的任务中,难度与训练数据模式可用性的关联比与抽象级别的关联更紧密。FinHardBench作为开源基准已发布。
英文摘要
Can large language models generate not just correct, but fast hardware? This paper investigates the question in financial FPGA design, where 5-10 nanoseconds of latency determines competitive advantage and designs iterate continuously as protocols, strategies, and regulations evolve. FinHardBench, a benchmark of 33 financial computing tasks, is presented together with three experiments that mirror the real-world FPGA iteration cycle: generating new modules from specifications, tuning system-level configurations across a 6-stage trading pipeline, and adapting existing modules to specification changes. Evaluation of six LLMs on 1530+ experiment rounds yields three findings: (1) models achieve 19-61% functional correctness with timing degradation up to 13.7$\times$ on specific tasks; (2) in system-level design space exploration, top LLMs converge to the optimal configuration with higher reliability than random search, simulated annealing, and Bayesian optimization baselines (5/5 seeds vs. 0-4/5 at the same 24-round budget); (3) strategy-level specification changes remain unsolved for most models. Across the six models, generation and DSE rankings overlap moderately: the strongest code generator is not the fastest architecture optimizer, and the weakest code generator (MiniMax M2.7) still reaches the system optimum on 4 of 5 seeds. On the tasks in FinHardBench, difficulty tracks training data pattern availability more closely than abstraction level. FinHardBench is released as an open-source benchmark.
Comments16 pages (10 pages main text). Published as a conference paper at COLM 2026. Code and benchmark: https://github.com/owenfucell/FinHardBench