CITECHOICE:文档呈现如何在智能体搜索中重新分配引文信用的因果审计
CITECHOICE: A Causal Audit of How Document Presentation Redistributes Citation Credit in Agentic Search
浏览论文内容
中文总结 AI 辅助
本研究通过CITECHOICE因果审计框架,揭示智能体搜索中文档呈现方式能因果性地重新分配引文信用,但非提高来源采纳率,且存在显著评估噪声。
中文摘要 AI 辅助
当多个检索到的来源支持同一主张时,答案引擎会引用其中一些而非全部。我们将这一决策称为引文分配,并引入CITECHOICE,这是对真实多轮智能体搜索的因果审计。从129份日常查询记录中,CITECHOICE选出113对同一调用中的文档对,这些文档对独立验证支持同一预先指定的事实,且不观察排名或答案结果;盲审人工复核确认了103对。它运行一种哈希验证的2×2重放实验,将配对顺序与一个目标的联合生成、保真度检查的结构化和散文式呈现交叉组合,而记录的其余部分保持不变。结果出现三项发现。首先,也是最重要的,结构化呈现集中了引文信用,而非明显提高来源采纳率。它使目标引文计数每答案增加+0.50次(95%置信区间[+0.20,+0.84];霍尔姆校正p=.033),而未增加总引文数或减少竞争者信用。预先指定的发生率效应(目标是否被引用)为+4.5个百分点,且不具结论性(95%置信区间[-1.4,+10.4];p=.168)。其次,观察性位置差异超过受控重排效应:排名1与排名5之间的引文率差距为42.3个百分点,而主重放中为+7.9个百分点,保留组中为0.0个百分点。第三,引文评估存在可测量的噪声下限。尽管聚合计数效应在30个冻结家族的重新解码下重复出现,但15%的二元决策发生变化,且解码约占单代家族效应方差的45%。综合这些发现,它们隔离了在控制下仍存续的因素:呈现能在冻结记录内因果性地重新分配可见引文信用。它们并未确立可靠的来源采纳、纯粹的格式机制或普遍的排名优势。
英文摘要
When several retrieved sources support the same claim, an answer engine cites some but not others. We call this decision citation allocation and introduce CITECHOICE, a causal audit of authentic multi-turn agentic search. From 129 everyday-query transcripts, CITECHOICE selects 113 same-call document pairs with independently verified support for the same pre-specified fact, without observing ranks or answer outcomes; blinded human review confirms 103. It runs a hash-verified 2-by-2 replay crossing pair order with jointly generated, fidelity-checked structured and prose renderings of one target while the rest of the transcript remains fixed. Three results emerge. First, and most importantly, structured rendering concentrates citation credit rather than clearly increasing source admission. It raises target citation count by +0.50 citations per answer (95 percent CI [+0.20, +0.84]; Holm-adjusted p=.033), without increasing total citations or reducing competitor credit. The pre-specified incidence effect (whether the target is cited at all) is +4.5 percentage points and inconclusive (95 percent CI [-1.4, +10.4]; p=.168). Second, observational position differences exceed controlled reordering effects: the citation-rate gap between rank 1 and rank 5 is 42.3 percentage points, compared with +7.9 percentage points in the main replay and 0.0 percentage points held out. Third, citation evaluation has a measurable noise floor. Although the aggregate count effect repeats under fresh decoding of 30 frozen families, 15 percent of binary decisions change and decoding accounts for an estimated 45 percent of single-generation family-effect variance. Together, these findings isolate what survives control: presentation can causally redistribute visible citation credit within frozen transcripts. They do not establish reliable source admission, a pure formatting mechanism, or a general rank advantage.