arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2607.21268cs.MAcs.AIecon.GNq-fin.EC

pAI-Econ-claude:用于人工智能辅助经济理论发展的门控人在回路多智能体架构

pAI-Econ-claude: A Gated Human-in-the-Loop Multi-Agent Architecture for AI-Assisted Economic Theory Development

Chen Zhu, Xiaolu Wang, Weilong Zhang

首次发表
浏览论文内容

中文总结 AI 辅助

研究针对经济学等领域中基于大语言模型智能体输出缺乏正确性信号的可靠性问题,提出pAI-Econ-claude门控人在回路多智能体架构,经实验评估,该架构能提高可审计性,证明不可逆人类判断分配比纯智能体自主性更具信息性。

中文摘要 AI 辅助

在许多社会科学研究任务中,如经济学领域,基于大语言模型的智能体必须产出没有廉价、任务完整且机器可读的正确性信号的输出。这给多智能体系统带来了独特的可靠性问题:当没有组件能认证最终结果时,生成、批判、协调和人类判断该如何组织?我们通过pAI-Econ-claude来解决此问题,它是一种用于人工智能辅助经济理论发展的门控、人在回路多智能体架构。智能体通过可检查中间记录的共享工作区进行协调;专门的门诊断目标故障模式并推荐回退而不认证正确性;人类检查点对难以逆转的决策保留决定权。我们在五个匹配的经济理论任务上针对无门控基线评估了该架构。两位对配置不知情的评估者对所有五个成对排名达成一致,在四个任务中更喜欢门控架构,在一个任务中更喜欢基线。平均失败严重程度从1.58降至1.16,而总体有用性从2.60升至3.10。最大的观察到的收益发生在现实检查拒绝错误的市场结构前提以及证明审查促使修改错误的福利主张时。负面案例表明支架也可能过于激进地压缩一个经济上重要的机制。结果支持一个有限的主张:门控监督提高了人工智能辅助经济理论的可审计性而不替代形式验证,并且不可逆人类判断的分配是比纯智能体自主性更具信息性的设计变量。工作流程可在这个https网址公开获取。

英文摘要

In many social-science research tasks, such as economics, LLM-based agents must produce outputs for which no cheap, task-complete, machine-readable correctness signal exists. This creates a distinctive reliability problem for multi-agent systems: how should generation, critique, coordination, and human judgment be organized when no component can certify the final result? We address this problem through pAI-Econ-claude, a gated, human-in-the-loop multi-agent architecture for AI-assisted economic theory development. Agents coordinate through a shared workspace of inspectable intermediate records; specialized gates diagnose targeted failure modes and recommend loopbacks without certifying correctness; and human checkpoints retain authority over decisions that are costly to reverse. We evaluate the architecture on five matched economic-theory tasks against an ungated baseline. Two evaluators blinded to configuration agreed on all five pairwise rankings, preferring the gated architecture in four tasks and the baseline in one. Mean failure severity fell from 1.58 to 1.16, while overall usefulness rose from 2.60 to 3.10. The largest observed gain occurred when a reality check rejected a false market-structure premise and a proof review prompted revision of a false welfare claim. The negative case shows that scaffolding can also compress an economically important mechanism too aggressively. The results support a bounded claim: gated oversight improves the auditability of AI-assisted economic theory without substituting for formal verification, and the allocation of irreversible human judgment is a more informative design variable than pure agent autonomy. The workflow is publicly available at https://github.com/maxwell2732/pAI-Econ-claude.

发表机构

  • China Agricultural University(中国农业大学)
  • University of Cambridge(剑桥大学)

机构由 AI 辅助整理,请以论文原文为准。

↑