发表机构
Virginia Tech; The Chronicle of Higher Education(弗吉尼亚理工大学; 高等教育纪事报)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
DataWeave通过对话交互、模式锚定、分析规划和可执行查询生成,支持记者对结构化数据进行探索性分析,将LLM作为可检查、纠正和引导的交互式伙伴,并在IPEDS数据集上验证了其有效性。
AI 中文摘要
数据新闻,即利用数据分析来发掘具有新闻价值的故事的实践,日益依赖于记者和调查新闻工作者揭示趋势、差异和问责叙事的能力。在实践中,探索大型结构化数据集仍然缓慢且脆弱:记者必须在多年间跨多个数据集浏览数百个变量,理解数据编码惯例,并在假设不断演变的情况下编写复杂的分析代码。尽管LLM常被宣传为“用英语提问,获得SQL/答案”,但真实的新闻编辑室工作流程暴露了反复出现的失败,例如模式不匹配和漂移、对领域语义和单位的误读,以及隐含的假设。我们提出了DataWeave,一个通过结合对话交互、模式锚定、分析规划和可执行查询生成来支持对结构化数据进行探索性分析的系统。DataWeave并未将LLM视为自主答案引擎,而是将其定位为交互式伙伴,其输出可以在假设变化时被检查、纠正和引导。我们展示了一项案例研究,专业记者使用我们的系统分析美国教育部的高等教育综合数据系统(IPEDS),这是一个具有大量领域语义和频繁模式更新的高重要性公共数据集。我们还报告了部署经验和迭代改进如何塑造了当前DataWeave架构及其分析工作流程。我们的发现提炼了在结构化数据分析中实现可信人机协作的设计原则和部署经验。
英文摘要
Data journalism, the practice of using data analysis to surface newsworthy stories, depends increasingly on the ability of reporters and investigative journalists to uncover trends, disparities, and accountability narratives. In practice, exploring large structured datasets remains slow and brittle: journalists must navigate hundreds of variables across many datasets over years, understand data coding conventions, and write non-trivial analysis code while hypotheses evolve. Although LLMs are often touted as "ask in English, get SQL/answers," real newsroom workflows expose recurring failures, e.g., schema mismatches and drift, misread domain semantics and units, and silent assumptions. We present DataWeave, a system that addresses these needs by combining conversational interaction, schema grounding, analytical planning, and executable query generation to support exploratory analysis over structured data. Rather than treating LLMs as autonomous answer engines, DataWeave frames them as interactive partners whose outputs can be inspected, corrected, and steered as hypotheses shift. We present a case study with professional journalists using our system to analyze the U.S. Department of Education's Integrated Postsecondary Education Data System (IPEDS), a high-stakes public dataset with substantial domain semantics and frequent schema updates. We also report how deployment experience and iterative refinement shaped the current DataWeave architecture and its analytical workflow. Our findings distill design principles and deployment lessons for trustworthy human-LLM collaboration in structured data analysis.