AI 中文总结
本研究通过对5个API模型、6国7目标的审计,发现LLM社会推理中调查国元数据可降低预测损失,但随机披露该线索无法可靠减弱其误导性,PROV-FORECAST数据集含14400组概率分布。
AI 中文摘要
当调查国元数据具有信息性时,它可以提升大型语言模型(LLM)对个体回答的预测效果,但随机分配的同一线索可能会误导预测。本研究采用记录内审计方法,测试披露随机标签的均匀、与记录无关的来源是否会降低其对国家导向的采纳度,以及经过验证的调查国是否会降低保留的Brier损失。研究通过独立人口锚点和记录的人类答案,在5个固定API模型、6个国家和7个经开发选定的目标上测量方向和结果。在主要的72条记录的审核后面板中,不透明标签和已披露随机标签各自产生0.214的国家导向偏移,配对衰减为0.0003(95%置信区间[-0.0157, 0.0166])。经过验证的国家使Brier损失降低0.040(95%置信区间[0.024, 0.056]),而随机标签的遗憾值包含零。非重叠的混合覆盖一致性面板保留了已披露随机标签的正向移动和经过验证的效用,同时衰减仍不确定。在选定目标上,经过验证的元数据在两个面板中均有用,但披露并未可靠地减弱随机标签的采纳度。PROV-FORECAST包含来自校正后面板的14400个配对项目级概率分布。
英文摘要
Survey-country metadata can improve an LLM's forecast of an individual response when informative, yet the same cue may redirect the forecast when assigned at random. A within-record audit tests whether disclosing a random label's uniform, record-independent origin reduces its country-directed uptake, and whether verified survey country lowers held-out Brier loss. Independent population anchors and recorded human answers measure direction and consequence across five fixed API models, six countries, and seven development-selected targets. In the primary post-review 72-record panel, opaque and disclosed-random labels each produced country-direction shifts of 0.214. Paired attenuation was 0.0003 (95% CI [-0.0157, 0.0166]). Verified country reduced Brier loss by 0.040 (95% CI [0.024, 0.056]), while random-label regret included zero. A non-overlapping mixed-coverage consistency panel retained positive disclosed-random movement and verified utility, while attenuation remained uncertain. On the selected targets, verified metadata was useful in both panels, but disclosure did not reliably attenuate random-label uptake. PROV-FORECAST contains 14,400 paired item-level probability distributions from the corrected panel.
Comments7 pages, 2 figures