arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

流系统中事件触发大语言模型调用的不确定性感知序贯决策规则

Uncertainty-Aware Sequential Decision Rules for Event-Triggered LLM Invocation in Streaming Systems

Zhaohui Wang

arXiv 2607.13048首次发表:更新:

发表机构

Viterbi School of Engineering, University of Southern California(南加州大学维特比工程学院)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

研究流系统中何时调用大语言模型的问题,将其转化为基于风险的序贯停止问题,证明了相关理论结果。通过在涡轮风扇退化数据上的实验,验证假设、比较基线,得出次线性遗憾、高诊断质量等结论,表明异常分数驱动的风险函数更优。

AI 中文摘要

流推理管道越来越多地将轻量级快速模型与成本高昂但能提供丰富语义理解的大语言模型(LLM)相结合。何时调用LLM这一核心问题尚未得到充分的形式化处理。我们将其视为基于风险的序贯停止问题,当观测历史上的风险函数超过阈值时触发策略。在此框架下,我们证明了六个结果,包括排除触发抖动的最小事件间隔时间界限、阈值策略的最优性等。一些经典触发族可表示为此框架的特殊情况。在有实际LLM调用的涡轮风扇退化数据(CMAPSS)上,我们验证了理论假设,消融了风险函数设计,与六个基线进行比较,并分析了成本敏感性和LLM故障模式。结果证实了次线性遗憾,所有原则性触发的α<1;高诊断质量,在我们的标准下,1600次LLM诊断中有92.9%达到接地分数>=0.75;以及异常分数驱动的风险函数在帕累托AUC上比其他方法占优约一个数量级。

英文摘要

Streaming inference pipelines increasingly pair lightweight fast models with Large Language Models (LLMs) that provide rich semantic understanding at substantial cost. The central question of when to invoke the LLM has received limited formal treatment. We cast this as a risk-based sequential stopping problem, where a trigger policy fires when a risk functional over the observation history exceeds a threshold. Within this framework, we prove six results: a minimum inter-event time bound excluding trigger chattering; optimality of threshold policies via smooth pasting; approximate SPRT guarantees under estimated parameters; O(sqrt(T log T)) regret for stationary streams, extending to O(sqrt((C_T + 1) T log T)) under C_T changepoints; O(1/sqrt(T)) convergence of online gradient descent for adaptive thresholds; and a calibration-to-miss-rate transfer inequality. Several classical trigger families, including event-triggered, optimal stopping, SPRT, CUSUM, and Bayesian triggers, can be expressed as special cases of this framework. On turbofan degradation data (CMAPSS) with real LLM calls, we empirically verify the theoretical assumptions, ablate the risk function design, compare against six baselines including a RouteLLM-style router and contextual bandits, and analyze cost sensitivity and LLM failure modes. The results confirm sublinear regret, with alpha < 1 for all principled triggers; high diagnostic quality, with 92.9 percent of 1600 LLM diagnoses reaching grounding score >= 0.75 under our rubric; and that anomaly-score-driven risk functions dominate alternatives by roughly an order of magnitude on the Pareto AUC.

Comments18 pages, 5 figures. Accepted to the ECML PKDD 2026 Research Track

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑