摇摆的预言机:LLM中的选择性更新与相关失效及其对科学工作流的影响
Wavering Oracles: Selective Updating and Correlated Failures in LLMs and Their Implications for Scientific Workflows
- HSBC Business School, Peking University(北京大学汇丰商学院)
- Artificial Intelligence Research Institute, Shenzhen University of Advanced Technology(深圳先进技术研究院人工智能研究所)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
本研究通过SycoBench-600评估多模型科学工作流中的选择性更新、错误相关性与裁决机制,发现模型间错误相关显著,提出可靠性选择器可恢复多数投票至预言机差距的67.7%。
AI中文摘要:
科学工作流日益依赖重复查询、多模型和交互式智能体。因此,可靠性取决于模型能否保持正确的结论、接受有效的修正,并产生选择器能够区分的错误。我们以SycoBench-600作为受控测量基板,通过选择性更新(定义为对误导性建议的抵抗力和对正确建议的接受度)来评估这些要求。该研究覆盖十个模型和17,055条轨迹。已发布的模型在选择性上跨度从13.4到71.6个百分点。在相同的本地评估下,Qwen3-4B在选择性适应上达到45.6点,Gemma3-4B在负14.1点处失稳,而SmolLM3-3B同时跟随正确和错误的明确建议,产生零选择性。匹配的干预措施识别出模型对怀疑、权威和明确建议的特定响应。在七个已发布模型中,最佳模型达到95.3%的准确率,多数投票达到88.6%,而预言机上限为99.8%。平均错误相关性为0.285,将七个模型的有效独立数量降至2.58。一个留一茎族(stem-family)的可靠性选择器达到96.2%的准确率,恢复了多数投票到预言机差距的67.7%。这些结果确立了选择性更新、错误多样性和校准裁决作为多模型科学工作流可联合测量的设计目标。
英文摘要:
Scientific workflows increasingly use repeated queries, multiple models, and interacting agents. Reliability therefore depends on whether models preserve correct conclusions, accept valid corrections, and contribute errors that a selector can distinguish. Using SycoBench- 600 as a controlled measurement substrate, we evaluate these requirements through selective updating, defined by resistance to misleading suggestions and uptake of correct suggestions. The study covers ten models and 17,055 trajectories. Published models span 13.4 to 71.6 percentage points in selectivity. Under identical local evaluation, Qwen3-4B is selectively adaptive at 45.6 points, Gemma3-4B is destabilized at minus 14.1 points, and SmolLM3-3B follows both correct and wrong explicit suggestions, producing zero selectivity. Matched interventions identify model specific responses to doubt, authority, and explicit advice. Among seven published models, the best reaches 95.3 percent accuracy, plurality reaches 88.6 percent, and the oracle ceiling is 99.8 percent. Mean error correlation of 0.285 reduces seven models to an effective independent count of 2.58. A leave-one-stem-family-out reliability selector reaches 96.2 percent, recovering 67.7 percent of the plurality-to-oracle gap. These results establish selective updating, error diversity, and calibrated adjudication as jointly measurable design targets for multi-model scientific workflows.