发表机构
Kensho Technologies(Kensho Technologies)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对生产金融数据库中不透明整数键导致通用文本到SQL性能低下的问题,提出FLINT系统,通过查找代理、嵌入检索和基于外键链的模式链接三个组件,在359个问题上超越多种基线,并已部署于生产环境。
AI 中文摘要
通用文本到SQL系统在Spider和BIRD等学术基准上取得了强劲性能,这些基准中的模式相对较浅,列值通常是人类可读的。在生产金融数据库中,概念存储为不透明的整数键而非人类可读字符串,这些方法的表现降至50%以下,因为即使是简单查询也需要多次连接,且过滤谓词引用不透明的ID。我们提出了金融链接文本到SQL(FLINT),一个领域专用的文本到SQL系统,通过三个关键组件弥合这一差距:(1)一个查找代理,动态地将自然语言概念解析为特定于问题的引用表约束;(2)基于嵌入的检索,从紧凑的专家编写模板库中检索结构相似的查询模板;(3)模式链接,通过遍历外键链将大型表模式修剪到相关子集,而非仅依赖名称相似性。我们在两个数据集上进行了评估,共包含359个问题,涉及生产金融模式。FLINT在使用相同LLM的情况下优于各种最先进的基线。该系统已作为金融数据检索服务的一部分部署到生产中。
英文摘要
General-purpose Text-to-SQL systems achieve strong performance on academic benchmarks like Spider and BIRD, where schemas are relatively shallow and column values are often human readable. In production financial databases, where concepts are stored as opaque integer keys rather than human-readable strings, these methods fall below 50%, as even simple queries require multiple joins and filter predicates reference opaque IDs. We present Financial LINking Text-to-SQL (FLINT), a domain-specialized Text-to-SQL system that closes this gap through three key components: (1) a lookup agent that dynamically resolves natural-language concepts to question-specific reference table constraints, (2) embedding-based retrieval of structurally similar query templates from a compact, expert-authored bank, and (3) schema linking that prunes a large table schema to the relevant subset by traversing foreign-key chains, rather than relying on name similarity alone. We evaluate on two datasets totaling 359 questions over production financial schemas. FLINT outperforms various state-of-the-art baselines using the same LLM. The system is deployed in production as part of a financial data retrieval service.
CommentsEMNLP Industry Track 2026