AI 中文总结
该研究发现LLM智能体易被伪造证据诱导对不可知问题作出错误承诺,其「行动/弃权(不执行)」闸门可通过合成案例微调优化,但对僵化响应格式敏感,部署需兼顾这一特性。
AI 中文摘要
向大语言模型(LLM)智能体展示一份专业外观的市场面板时,其对一个经证明不可预测的问题作出方向性预测的频率,远高于仅被询问该问题的情况:在12种前沿模型中,随着证据的升级,承诺率从6.5%升至54.0%。当面板上的所有数字均为虚构时,其承诺率依然相近:即便整个展示内容都是编造的,除问题本身外模型所见的一切均非真实,承诺率仍从24.5%升至36.8%,与真实市场数据产生的37.6%在统计上无差异。解锁自信行动的并非信息,而是其包装的权威性。这种失败是狭隘且可定位的。能力不足并非原因:在附于相同面板的匹配可回答问题上,相同模型几乎总能以接近完美的准确率作答。信念也非原因:在使行动率波动48个百分点的梯度上,陈述的概率几乎未发生变化,且表现逊于气候学基线。判断缺失也非原因:被要求在行动前对问题的可知性分类时,模型90%的时间称其为不可约,却仅在其中0.4%的情况下作出承诺。失败的是「行动/弃权(不执行)」的闸门,且该效应集中于少数模型而非普遍存在。由于该闸门可分离,故可被训练。在540个合成案例(主要为骰子、硬币、罐子和计时器)上对3B模型进行监督微调,使原始案例的承诺率降至0.0%,并迁移至三个未见过的领域。它并非对所有情况都有效:当响应格式留有推理空间时,闸门有效;而移除该空间的僵化格式会使模型在原本能正确回答的问题上表现自信且错误。该闸门可训练且对上下文敏感,部署需兼顾这两点。
英文摘要
An LLM agent shown a professional-looking market panel commits to a directional call on a provably unpredictable question far more often than one asked the bare question: across 12 frontier models, commitment rises from 6.5% to 54.0% as evidence is escalated. It commits just as readily when every number on the panel is invented: fabricating the entire display, so nothing the model can see is true except the question itself, still lifts commitment from 24.5% to 36.8%, statistically indistinguishable from the 37.6% produced by genuine market data. What unlocks confident action is not information but the authority of its packaging. The failure is narrow and locatable. Incapacity is not the answer: on matched answerable questions attached to the same panels, the same models answer essentially always, at near-perfect accuracy. Nor is it belief - stated probabilities barely move across the gradient that swings action by 48 points, and score worse than a climatological baseline. Missing judgment isn't it either: asked to classify a question's knowability before acting, models call it irreducible 90% of the time and then commit on just 0.4% of those. The act/don't-act gate is what fails, and the effect is concentrated in a few models rather than universal. Because the gate is separable, it can be trained. Supervised fine-tuning of a 3B model on 540 synthetic cases, predominantly dice, coins, jars and timers, drives commitment to 0.0% on the original cases and transfers to three unseen domains. It does not survive everything: the gate holds exactly when the response format leaves room to reason, and rigid formats that remove that room leave the model confident and wrong on questions it otherwise answers correctly. The gate is trainable and context-fragile, and deployment needs both halves of that sentence.
Comments27z pages, 6 figures. Code, data, pre-registration and all cached model outputs: https://github.com/Pranav-1100/confidence-calibration-evaluation . Also archived at Zenodo, DOI 10.5281/zenodo.22043517