联合近乎免费,推理并非如此:AI 联合科学家在蛋白质表征工作流中的权衡
Federation Is Nearly Free, Reasoning Is Not: Tradeoffs for AI Co-Scientists in Protein Characterization Workflows
浏览论文内容
中文总结 AI 辅助
该研究在生产级科学智能体平台上,对比联合拓扑、PPO策略及不同LLM等,发现LLM选择对蛋白质表征预测质量影响最大,PPO策略成本低且一致性好,联合对性能影响可忽略,为科学工作流智能体部署提供了指导。
中文摘要 AI 辅助
自然语言驱动的自主联合科学家工作流存在一个基本权衡:灵活性与推理能力以牺牲确定性、可复现性和可观测性为代价。这类智能体日益需要跨机构边界通信,而联合拓扑会影响延迟与成本。我们在一个生产级科学智能体平台上通过受控消融实验系统评估了这些权衡。我们采用一项可验证任务:给定一个蛋白质序列,要求智能体通过路由常用工具来可靠表征其功能。我们对比了联合拓扑、经典强化学习(RL)与大语言模型(LLM)驱动的框架、语言模型以及提示专业度,还按蛋白质新颖性对结果分层。我们发现,大语言模型的选择对预测质量的影响远大于拓扑或提示(Opus 准确率约 92%-94%,o4-mini 约 40%-50%)。近端策略优化(PPO)策略在零令牌成本、最快延迟和完美一致性方面,准确率几乎与最优大语言模型相当(88%),但无推理轨迹。经专业提示的大语言模型达到最高准确率,但成本高且一致性较差;任务难度最大时,提示依赖性最强。联合对性能的影响可忽略不计。这些结果为部署科学工作流智能体提供了可行指导:对于常规可验证任务,低成本确定性策略可提供接近前沿的准确率与完全可复现性,而灵活的大语言模型推理最适合开放性发现。
英文摘要
Natural language driven autonomous co-scientist workflows involve a fundamental trade-off between flexibility and reasoning at the expense of determinism, reproducibility, and observability. Such agents increasingly must communicate across institutional boundaries, where federation topology can shape latency and cost. We systematically evaluated these tradeoffs using a controlled ablation on a production agentic platform for science. We use a verifiable task: given a protein sequence, we ask an agent to confidently characterize its function by routing across common tools. We compare federation topology, classic RL vs LLM-driven harnesses, language model, and prompt expertise. We also stratify results by protein novelty. We find that the choice of LLM dominated prediction quality far more than topology or prompting (Opus ~92%-94% vs o4-mini ~40%-50%). The PPO policy was nearly as accurate as the best LLM (88%) at zero token cost, fastest latency, and perfect consistency, but yields no reasoning trace. Expert prompted LLMs reached the highest accuracy but were high-cost and less consistent; prompt dependence was largest when the task was hardest. Federation imposed a negligible penalty on performance. These results offer actionable guidance for deploying agents for scientific workflows: for routine, verifiable tasks, a cheap deterministic policy delivers near-frontier accuracy with complete reproducibility, while flexible LLM reasoning is best reserved for open-ended discovery.