发表机构
Icahn School of Medicine at Mount Sinai; University of Georgia; Carnegie Mellon University; Boston University School of Public Health; University of Central Florida(西奈山伊坎医学院; 佐治亚大学; 卡内基梅隆大学; 波士顿大学公共卫生学院; 中佛罗里达大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本综述将LLM复杂问题求解建模为序贯估计-决策问题,提出五组件框架与问题-控制拟合诊断假设,区分错误与不确定性类型,指出验证、恢复、校准和评估为开放挑战。
AI 中文摘要
使用大规模语言模型(LLMs)进行复杂问题求解(CPS)通常被归结为更强的推理能力或更长的生成过程。然而,早期步骤错误放大、提示词脆弱性以及未能修正错误承诺等问题,难以仅通过缺失知识或表达能力不足来解释。本综述将CPS解释为在潜在解状态上的一个序贯估计与决策问题。一个控制器维护关于未观测解轨迹的信念,随着噪声中间证据的到来更新该信念,并决定是承诺、验证、分支、回滚还是弃权(不执行),以最小化期望损失。推理提供候选转换和解释,而过程控制则塑造和评估这些提议,并调节后续转换和观测。在此框架内,我们将现有方法围绕五个组成部分进行组织:显式状态表示、转换结构、验证与约束执行、搜索与回滚,以及不确定性管理。我们还根据评估指标所估计的统计量来解读这些指标。该框架进一步产生了一个诊断假设:当干预措施针对观测失败所涉及的错误或不确定性组成部分时,它们应最为有效。我们区分了系统性、随机性和不可约错误,以及认知性和偶然性不确定性,并将这种对齐称为问题-控制拟合,其失败称为控制不匹配。例如,额外采样可以减少采样变异性,同时使共享的系统性错误保持不变。这一视角阐明了当前方法估计和控制了什么,什么仍然未受控制,以及为什么可靠的验证、有针对性的恢复、校准的不确定性和匹配预算的评估是核心的开放问题。
英文摘要
Complex problem solving (CPS) with large language models (LLMs) is often framed as a matter of stronger reasoning or longer generation. Yet early-step error amplification, prompt brittleness, and failures to revise incorrect commitments are difficult to explain by missing knowledge or expressive capacity alone. This survey interprets CPS as a sequential estimation-and-decision problem over a latent solution state. A controller maintains a belief about an unobserved solution trajectory, updates it as noisy intermediate evidence arrives, and decides whether to commit, verify, branch, roll back, or abstain to minimize expected loss. Reasoning supplies candidate transitions and interpretations, whereas process control shapes and evaluates those proposals and regulates subsequent transitions and observations. Within this framework, we organize existing methods around five components: explicit state representation, transition structuring, validation and constraint enforcement, search and rollback, and uncertainty management. We also interpret evaluation metrics according to the statistical quantities they estimate. The framework further yields a diagnostic hypothesis: interventions should be most effective when they target the error or uncertainty component implicated by an observed failure. We distinguish systematic, stochastic, and irreducible error together with epistemic and aleatoric uncertainty, and call this alignment problem-control fit and its failure control mismatch. For example, additional sampling may reduce sampling variability while leaving a shared systematic error unchanged. This perspective clarifies what current methods estimate and control, what remains uncontrolled, and why reliable validation, targeted recovery, calibrated uncertainty, and matched-budget evaluation are central open problems.
Comments82 pages, 7 figures. Submitted to Artificial Intelligence Review