Text2Dashboard:面向企业DataBrain的自然语言仪表板生成受控代理架构
Text2Dashboard: A Governed Agent Architecture for Natural-Language Dashboard Generation over Enterprise DataBrain
查看机构详情
- University College London(伦敦大学学院)
机构由 AI 辅助整理,请以论文原文为准。
浏览论文内容
中文总结 AI 辅助
Text2Dashboard提出一种受控代理架构,将模型提议与确定性执行结合,在DataBrain上实现自然语言生成可检查仪表板,在冻结任务中达到6/8严格成功率,并验证了故障处理能力。
中文摘要 AI 辅助
Text2Dashboard是一个DataBrain专属原型,可将自然语言分析请求转化为可检查的仪表板。一个可安装的Codex插件和独立的Agent Runtime将受模式约束的模型决策与类型化工具、持久状态以及用于审批、审计、检查点、恢复和故障处理的确定性Hooks相结合。该流程解析实体、发现元数据、强制执行只读SQL、组合仪表板,并应用静态检查、动态预检和浏览器检查。模型提出操作,而确定性软件控制执行并记录状态转换。我们在冻结的真实DataBrain任务和受控Hook故障上评估了该工作流。在元数据和SQL任务上,严格成功率为6/8:元数据选择通过4/4,所有四个SQL任务均满足语义标准,2/4满足精确输出列契约。最终版本通过了4/4个单面板仪表板任务、一个双面板任务和一个现有仪表板细化任务;一个参数化任务超出了其步骤限制。所有十个故障场景均达到指定结果,且无未经批准的外部副作用。在每个报告组中,模型推理占观察运行时的97%以上。这些针对DataBrain的小规模结果并不能证明生产就绪性、通用文本到SQL的准确性,或相对于手动仪表板构建的效率优势。
英文摘要
Text2Dashboard is a DataBrain-specific prototype that turns natural-language analytic requests into inspectable dashboards. An installable Codex plugin and standalone Agent Runtime combine schema-constrained model decisions with typed tools, persistent state, and deterministic Hooks for approval, audit, checkpointing, recovery, and failure handling. The pipeline resolves entities, discovers metadata, enforces read-only SQL, composes dashboards, and applies static checks, dynamic preflight, and browser inspection. The model proposes actions while deterministic software controls execution and records state transitions. We evaluate the workflow on frozen real-DataBrain tasks and controlled Hook faults. Strict success was 6/8 on metadata and SQL tasks: metadata selection passed 4/4, all four SQL tasks met semantic criteria, and 2/4 met the exact output-column contract. The final release passed 4/4 single-panel dashboard tasks, one two-panel task, and one existing-dashboard refinement; a parameterised task exceeded its step limit. All ten fault scenarios met their specified outcomes without unapproved external side effects. Model inference accounted for over 97\% of observed runtime in every reported group. These small, DataBrain-specific results do not establish production readiness, general text-to-SQL accuracy, or an efficiency advantage over manual dashboard construction.