arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2610.04156cs.CLcs.AIcs.LG

轨迹推导置信度:用于可靠、资源感知的临床文本到SQL智能体

Trajectory-Derived Confidence for Reliable, Resource-Aware Clinical Text-to-SQL Agents

Mincheol Daniel Song, Joshua Ward, Jake Jung, Guang Cheng

首次发表
浏览论文内容

中文总结 AI 辅助

提出轨迹推导的置信度层Sentinel,通过分类器在运行前、中、后三节点决策停止,提升临床文本到SQL智能体的可靠性得分并削减计算开销。

中文摘要 AI 辅助

用于临床文本到SQL应用的LLM智能体会在多个步骤中自主推理,但无法评估其自身推理或输出是否可信。在医疗保健等高杠杆应用中,这构成了重大风险,系统错误可能代价高昂。这些可靠性失败同时也是资源失败:错误的推理轨迹将计算预算花费在必须丢弃的输出上。我们引入了Sentinel,一个基于轨迹的、分类器驱动的置信度层,它分析智能体的推理、代码和数据库输出,并在三个节点决定是否停止:在智能体运行前拒绝无法回答的问题,在运行中途终止注定失败的轨迹,以及在交付时扣留不可信的答案。在此,利用Chow规则,我们根据EHRSQL共享任务的可靠性得分优化决策,该得分在给定效用权重下惩罚错误答案。在基准EHRSQL上,我们发现当错误具有低效用权重时,Sentinel将该得分从+0.08提升至+0.24,远高于拒绝所有问题所得的+0.03,同时交付答案的准确率从54%上升至69%,覆盖率从84%降至55%。在更高风险情况下,我们测试的智能体很少能可靠地作答以交付,而Sentinel自行检测到这一点,弃权(不执行)至同样的+0.03,而未受监控的智能体得分则为-3.02。相同的停止决策削减了计算量:在低风险情况下,系统仍会作答,Sentinel以极低或无可靠性代价消除了13-28%的智能体步骤。

英文摘要

LLM agents for clinical text-to-SQL applications reason autonomously over multiple steps but cannot assess whether their own reasoning or outputs can be trusted. In high leverage applications such as healthcare, this presents a critical risk where system mistakes can be costly. These reliability failures are also resource failures: an incorrect reasoning trajectory spends computation budget on outputs that must be discarded. We introduce Sentinel, a trajectory-derived, classifier-based confidence layer that analyzes an agent's reasoning, code and database outputs to decide at three points whether to stop: refusing unanswerable questions before the agent runs, halting doomed trajectories mid-run, and withholding untrustworthy answers at delivery. Here, utilizing Chow's rule, we optimize decisions under the EHRSQL shared task's Reliability Score, which penalizes incorrect answers given a utility weighting, and find on the benchmark EHRSQL that Sentinel raises this score from +0.08 to +0.24 when mistakes have a low utility weighting, well above the +0.03 earned by refusing every question, with delivered-answer accuracy rising from 54% to 69% as coverage falls from 84% to 55%. At higher stakes the agents we test rarely answer reliably enough to deliver, and Sentinel detects this on its own, abstaining to that same +0.03 where the unmonitored agent scores -3.02. The same stopping decisions cut computation: at low stakes, where the system still answers, Sentinel eliminates 13-28% of agent steps at little or no reliability cost.

发表机构

  • University of California, Los Angeles(加利福尼亚大学洛杉矶分校)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑