发表机构
The New York Times; National Public Radio(《纽约时报》; 美国国家公共广播电台)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本文介绍《纽约时报》部署的爱泼斯坦文件引擎,一种将记者提问转化为SQL查询并返回可验证答案的AI智能体,通过重复匹配方法突出新信息,服务了100多名记者并助力至少20篇报道,主张新闻智能体应作为资料与知识接口而非自主写作者。
AI 中文摘要
2026年1月30日,美国司法部发布了一个关于杰弗里·爱泼斯坦的多媒体资料集,其中包括约三百万页的PDF文件。我们描述了《纽约时报》部署用于调查这些文件的人工智能智能体——爱泼斯坦文件引擎。该引擎将记者的提问转化为跨三个语料库的Google BigQuery SQL查询:与爱泼斯坦相关的发布文件、《纽约时报》的档案以及外部的爱泼斯坦相关新闻标题。它使用大语言模型来规划查询,并返回富含引用的答案,记者可以验证并信任这些答案。超过100名记者使用了该引擎,它为至少20篇已发表的文章做出了贡献。我们报告了记者如何查询它,并描述了Diff——我们的文本和视觉重复匹配方法,该方法放大了新颖性信号,使引擎能够呈现真正的新信息。我们认为,新闻编辑室的智能体最能服务于新闻编辑室的方式不是作为自主写作者,而是作为原始资料和机构知识的接口。
英文摘要
On Jan. 30, 2026, the U.S. Department of Justice released a mixed-media collection concerning Jeffrey Epstein, including about three million pages of PDFs. We describe the Epstein Files Engine, an A.I. agent The New York Times deployed to investigate the files. The Engine translated reporter questions into Google BigQuery SQL queries across three corpora: Epstein-related releases, the Times's archive and external, Epstein-related news headlines. It used an LLM to plan queries and returned citation-rich answers a reporter could verify and trust. More than 100 journalists used the Engine, and it contributed to at least 20 published stories. We report how reporters queried it and describe Diff, our text-and-visual duplicate matching method that amplified novelty signals and allowed the Engine to surface genuinely new information. We argue that newsroom agents serve newsrooms best not as autonomous writers, but as interfaces to source material and institutional knowledge.
Comments6 pages, 2 figures, 2 tables. Presented at the Computation + Journalism Symposium (C+J 2026)