arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.28594cs.AI

从问题优先到分析师优先:面向主动企业分析的领域专家技能与验证知识编译

From Question-First to Analyst-First: Domain-Expert Skills and Verified Knowledge Compilation for Proactive Enterprise Analytics

Harmohit Singh, Rahul Sharma

首次发表
浏览论文内容

中文总结 AI 辅助

该研究提出了一种从问题优先转为分析师优先的主动企业分析系统,通过领域专家技能抽象与离线知识编译循环实现,无需查询日志,可生成验证过的报告与建议问题。

中文摘要 AI 辅助

对话式分析系统假设用户已形成明确问题,导致非专业用户面对陌生企业数据模式时只能空白查询。商业“主动”工具仅通过分析师精心设计的指标层检测统计异常来缩小差距,而学术领域的下一个问题推荐器依赖新数据集所不具备的查询日志。我们描述了一种将交互模型从问题优先反转为分析师优先的生产级分析系统,该系统通过两个耦合架构思路实现。第一,可插拔的领域专家“技能”抽象:一种基于文件夹、无需数据库的主题包(包含清单、各阶段提示维度、关键词路由引用、报告模板及可选计算),通过确定性模式匹配为每个(客户端、数据集)自动选择,并作为跨领域关注点嵌入智能体管道、模式探索器和报告引擎的每个阶段,缺失时则退化为无操作。由于技能是通过确定性解析的自包含文件夹,其目录可无限扩展:形成一个可扩展的领域专家市场。第二,离线知识编译循环:智能体通过 DuckDB 探查数据集的 parquet(对生产环境零负载),运行经审核的逐表收敛并带自我修复重试,通过值重叠验证连接的数据,生成持久的模式知识,用于驱动常设专家报告,其中每个发布的指标都通过重新执行其证据 SQL 进行重新验证,还生成反映报告议程的建议问题。这些形成了主动循环:报告呈现数据,数据生成问题,点击后启动已验证的深入探究,所有操作都在使用查询框之前完成。我们提供了形式化模型并报告了单租户示例证据,未提出用户研究或基准主张,贡献在于该架构及其可防御性。

英文摘要

Conversational analytics systems assume the user already has a well-formed question, leaving a non-expert facing a blank query box on an unfamiliar enterprise schema. Commercial 'proactive' tools narrow this gap only by detecting statistical anomalies over analyst-curated metric layers, and academic next-question recommenders depend on query logs that a fresh dataset lacks. We describe a production analytics system that inverts the interaction model from question-first to analyst-first through two coupled architectural ideas. First, a pluggable domain-expert 'skill' abstraction: a folder-based, database-free subject-matter pack (a manifest, per-stage prompt facets, keyword-routed references, report templates, and optional compute) auto-selected per (client, dataset) by deterministic schema matching and spliced as a cross-cutting concern into every stage of an agentic pipeline, the schema explorer, and the report engines, degrading to a strict no-op when absent. Because a skill is a self-contained folder resolved deterministically, the catalogue is open-ended: an extensible marketplace of domain experts. Second, an offline knowledge-compilation loop: an agent probes the dataset's parquet via DuckDB (zero load on production), runs critic-gated per-table convergence with self-healing retries, and data-validates joins by value overlap, producing durable schema knowledge that drives standing expert reports whose every published metric is re-verified by re-executing its evidence SQL, plus suggested questions that mirror the report agenda. These close a proactive loop: reports surface numbers, the numbers seed questions, and a click launches a verified deep dive, all before the query box is used. We give a formal model and report illustrative single-tenant evidence. We make no user-study or benchmark claims; the contribution is the architecture and its defensibility.

补充信息

↑