发表机构
Blossom AI; Blossom AI Labs(布罗瑟姆人工智能公司; 布罗瑟姆人工智能实验室)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本研究发现LLM智能体商业评估中存在构念效度失效问题,提出构念效度契约,指出原安全措施福利增益估计无效,受控研究结果不确定,需经多维度检查才能明确安全措施的实际价值。
AI 中文摘要
交互式模拟越来越多地用于评估由语言模型智能体组成的市场中的政策,其输出可能呈现出经济相关的结果——如价格、利润、消费者剩余和福利——却未体现出声明中所描述的行为。我们在一个可配置酒店交易的多轮买卖双方测试平台中审计了这种风险。初始实现报告称,在Qwen2.5 1.5B至14B的模型层级中,两项市场安全措施带来的福利增益分别为+87.4、+35.0和+28.8;同时,该实现为弃权(不执行)和未弃权的智能体提供了不同的报价模式和选择流程。在保持报价模式和买方选择器固定的情况下,配对对比结果变为+7.2、-13.9和+23.8。四个最大的14B单代效应平均值为+229;在每个配置条件下生成三代后,其平均值为+37.6(95%自助抽样区间为[-34.2, 109.3]),而生成残差占该事后探针变异的49.9%。卖方激励检查呈现非单调性:增加利润压力产生的利润低于默认卖方提示。脚本化阳性对照表明了这一问题的重要性:利润最大化的卖方已实现最优福利,因此安全措施主要是再分配并降低福利;仅当卖方被明确编程为强制低效组合时,安全措施才会创造福利。我们提出了一份构念效度契约,将激励有效性、协议隔离、随机稳定性和福利核算区分开来,并在提出实质性政策主张前返回“无效”或“不确定”。在本研究中,根据协议隔离,原始估计为“无效”;而受控研究在激励有效性和随机稳定性下仍为“不确定”。该案例并非表明安全措施无效,而是表明在模拟智能体和协议通过这些检查之前,其表面价值无法被识别。
英文摘要
Interactive simulations increasingly evaluate policies in markets populated by language-model agents. Their outputs can look economic---prices, profits, consumer surplus, and welfare---without instantiating the behavior named in the claim. We audit this risk in a multi-turn buyer--seller testbed for configurable hotel transactions. An initial implementation reported welfare gains from two marketplace guardrails of +87.4, +35.0, and +28.8 across a Qwen2.5 1.5B--14B ladder. It also gave guarded and unguarded agents different offer schemas and choice procedures. Holding the schema and buyer chooser fixed changes the paired contrasts to +7.2, -13.9, and +23.8. The four largest 14B single-generation effects averaged +229; after three generations per profile-condition, they averaged +37.6 (95% bootstrap interval [-34.2, 109.3]), while generation residuals account for 49.9% of variation in this post-hoc probe. A seller-incentive check is non-monotone: increasing profit pressure produces less profit than the default seller prompt. Scripted positive controls show why this matters. A profit-maximizing seller already attains first-best welfare, so guardrails mostly redistribute and reduce welfare; they create welfare only when the seller is explicitly programmed to force inefficient bundles. We contribute a construct-validity contract separating incentive validity, protocol isolation, stochastic stability, and welfare accounting, and returning INVALID or INCONCLUSIVE before substantive policy claims. In our case, the original estimate is INVALID under protocol isolation, while the controlled study remains INCONCLUSIVE under incentive validity and stochastic stability. The case does not show that guardrails are ineffective; it shows their apparent value is unidentified until the simulated agents and protocol pass these checks.
CommentsSubmitted to the NeurIPS 2026 Trust-AI-Eval Workshop. 7 pages, 2 figures, 3 tables. The accompanying artifact is available at https://anonymous.4open.science/r/a2a-evaluation-artifact-staging-2F25/