arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

深度研究写作的重新设计与审计以生成忠实报告

Redesigning and Auditing Deep Research Writing for Faithful Reports

Hiroaki Hayashi, Pranav Narayanan Venkit, Prafulla Kumar Choubey, Chien-Sheng Wu

arXiv 2608.28643首次发表:更新:

发表机构

Salesforce AI Research(Salesforce AI研究院)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

针对深度研究系统的细粒度事实错误问题,提出CLAIMPROBE审计方法,设计CLAIMWRITER主张分层写作器,可显著降低幻觉、提升必要事实召回率且成本效益更高。

AI 中文摘要

基于规则的深度研究(DR)系统评估往往会掩盖生成报告中细粒度的事实错误。我们推出CLAIMPROBE,这是一种主张级审计方法,可将深度研究报告分解为各项主张,并对照检索到的证据衡量幻觉、错误归因、引用规范度以及必要事实的召回率。通过CLAIMPROBE,我们发现,即使强大的深度研究流水线的规则评分保持稳定,它们仍可能遗漏关键证据并错误归因主张。随后我们提出CLAIMWRITER,这是一种基于主张的分层写作器,它提取源事实,将其映射到查询生成的大纲,并从与源链接的主张表示中起草每个部分。在三个现有的深度研究框架中,仅将报告写作器替换为CLAIMWRITER,就能将幻觉减少2.6至4.5倍,将必要事实的召回率提高1.2至1.7倍,同时在很大程度上保持报告的整体质量。CLAIMWRITER还支持局部修订:当源发生变化时,它在所有更新方法中以最高速率将变化的源事实传播到修订后的报告中,同时也更具成本效益。

英文摘要

Rubric-based evaluations of deep-research (DR) systems often obscure fine-grained factual failures in generated reports. We introduce CLAIMPROBE, a claim-level audit that decomposes DR reports into claims and measures hallucination, misattribution, citation hygiene, and necessary-fact recall against retrieved evidence. Using CLAIMPROBE, we find that strong DR pipelines can omit key evidence and misattribute claims even when their rubric scores remain stable. We then propose CLAIMWRITER, a hierarchical claim-based writer that extracts source facts, maps them to a query-derived outline, and drafts each section from a source-linked claim representation. Across three prior DR frameworks, replacing only the report writer with CLAIMWRITER reduces hallucination by 2.6 to 4.5 times and improves necessary-fact recall by 1.2 to 1.7 times, while largely preserving overall report quality. CLAIMWRITER also enables localized revision: when sources change, it propagates changed source facts into revised reports at the highest rate among update methods, while also being more cost-effective.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑