arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

共享与结构化输入削弱推理型AI智能体的集体随机选择

Shared and structured inputs undermine collective random choice by reasoning AI agents

Takahiro Ezaki, Naoto Imura, Katsuhiro Nishinari

arXiv 2610.09667首次发表:更新:

发表机构

Research Center for Advanced Science and Technology, The University of Tokyo; Department of Aeronautics and Astronautics, School of Engineering, The University of Tokyo(东京大学先端科学技术研究中心; 东京大学工学部航空宇宙工学科)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本研究通过行为测试揭示推理AI智能体在随机选择中依赖共享和结构化输入,导致集体关联与偏倚,并指出输入依赖的偏倚、相关性和可预测性应成为智能体评估的核心目标。

AI 中文摘要

随机选择在资源分配和审计中被广泛使用,因此其可靠实施对AI智能体系统至关重要。对六个推理模型的行为测试揭示了基于标识符的选择中所使用的阈值规则和整除规则。对于遵循阈值的GPT-6 Sol和Gemini 3.8 Flash,单智能体测量前瞻性地预测了在共享标识符下的关联参与,以及在具有共同时间戳位的不同标识符下的偏倚参与。更改日期、格式和标识符标签揭示了这些预测何时成立。要求独立随机化的明确指令减少了但并未消除共享输入的相关性。为测试对监督的影响,我们要求四个模型随机选择客户请求以供人工审查。GPT-6 Sol在从标识符中可预测地选择的同时接近目标比率;其他模型很少选择请求。所有四个模型都紧密遵循提供的随机抽取结果。这些发现暴露了仅凭选择率无法发现的集体和审计漏洞,使得输入依赖的偏倚、相关性和可预测性成为智能体评估的核心目标。

英文摘要

Random selection is widely used in resource allocation and auditing, making reliable implementation essential for AI-agent systems. Behavioural tests across six reasoning models uncovered threshold and divisibility rules used in identifier-based choices. For threshold-following GPT-6 Sol and Gemini 3.8 Flash, single-agent measurements prospectively predicted correlated participation under shared identifiers and biased participation under distinct identifiers with common timestamp bits. Changing dates, formats and identifier labels revealed when these predictions held. Explicit instructions to randomize independently reduced but did not eliminate shared-input correlation. To test implications for oversight, we asked four models to select customer requests randomly for human review. GPT-6 Sol approached the target rate while selecting predictably from identifiers; the others rarely selected requests. All four closely followed supplied random draws. These findings expose collective and audit vulnerabilities that selection rates alone miss, making input-dependent bias, correlation and predictability central targets for agent evaluation.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑