KITE:利用稀疏旗舰校准扩展Jev群体实验
KITE: Scaling Jev Population Experiments with Sparse Flagship Calibration
- The University of Tokyo(东京大学)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
KITE通过稀疏旗舰锚点校准行为内核,以低成本扩展群体实验,显著降低误差并提升决策增益,支持大规模干预筛选与审计。
AI中文摘要:
KITE对每个唯一状态查询一次类型化行为内核,然后使用事件键控随机性和公共随机数从表中执行任意规模的群体。一个昂贵的旗舰模型被保留用于稀疏的配对锚点,以估计干预效果。测量到的人-模型差异作为共享误差传播到每个结论中。因此,群体实验成本随唯一状态和锚点数量扩展,而不确定性由关于人的证据而非蒙特卡洛噪声决定。在包含9,070名参与者的Epstein实验中,覆盖1.7%状态的锚点将效应误差降低了41%(绝对MAE减少0.0125)。在37个留出的SocSci210实验中,0.5-1.5%的锚点覆盖率将捕获的决策增益从0.27提高到0.39。该内核在16国研究的全部15个新国家中通过了内容保真度标准。共享差异在名义80%和90%下产生了93%和96%的回顾性覆盖率,而仅靠人类采样不确定性分别为29%和36%。一百万智能体在笔记本电脑上于0.9秒内执行了20个制表步骤。该架构提供了一条在人体试验前筛选候选干预措施、进行多国内容审计以及进行不确定性感知政策比较的途径,成本仅为数千次内核调用及稀疏旗舰锚点。属性特定的证据记录将每次使用与其验证范围、校正来源和不确定性联系起来,使这些应用可审计。
英文摘要:
KITE queries a typed behavioral kernel once per unique state, then executes populations of any size from the table with event-keyed randomness and common random numbers. An expensive flagship model is reserved for sparse paired anchors that estimate intervention effects. Measured human-model discrepancy is propagated as shared error into every conclusion. Population-experiment cost thus scales with unique states and anchors, while uncertainty is governed by evidence about people rather than Monte Carlo noise. On Epstein experiments with 9,070 participants, anchors covering 1.7% of states reduced effect error by 41% (absolute MAE reduction 0.0125). On 37 held-out SocSci210 experiments, 0.5-1.5% anchor coverage raised captured decision gain from 0.27 to 0.39. The kernel passed content-fidelity criteria in all 15 new countries of a 16-country study. Shared discrepancy yielded retrospective coverage of 93% and 96% at nominal 80% and 90%, versus 29% and 36% from human sampling uncertainty alone. A million agents executed 20 tabulated steps in 0.9 seconds on a laptop. This architecture offers a route to screening candidate interventions before human trials, multi-country content audits, and uncertainty-aware policy comparison at the cost of a few thousand kernel calls with sparse flagship anchors. Property-specific evidence records connect each use to its validation scope, correction provenance, and uncertainty, making these applications auditable.