arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.38612cs.CL

StreamDecisionBench:评估演化语言流上的生效决策

StreamDecisionBench: Evaluating Decisions in Force on Evolving Language Streams

Jhen-Ke Lin, Chung Chun Wang

首次发表
浏览论文内容

中文总结 AI 辅助

针对语言模型在流式证据下决策可能过时的问题,提出StreamDecisionBench基准,按时刻评估生效决策并归因错误,发现推理提升判断但增加延迟代价,权衡随时间尺度变化。

中文摘要 AI 辅助

随着自然语言驱动越来越多的应用,语言模型日益作为决策组件在程序内部运行:程序将当前状态发送给模型,并依据模型返回的决策执行操作,直到新的决策到来。当推理过程中证据发生变化时,一个对自身状态正确的决策可能在状态过去后仍然生效,例如当客户开始读出卡号时,通话录音器仍在运行;未计时的(离线)准确率会将此类错误计为正确。我们提出StreamDecisionBench(SDB),它评估每个时刻的生效决策,并将每个错误时刻归因于判断、延迟或两者兼有。其场景在四个应用家族中流式提供证据,参考决策由可执行代码根据公开规则计算得出。我们通过在对数时间轴上的归一化曲线下面积,汇总更新间隔为1-5秒内的生效准确率,对相等的乘法范围赋予相等的权重。在四个托管模型的六种设置中,该得分与场景级未计时准确率和预言机综合时间得分的乘积相差在2.9个百分点以内。推理改善了判断,但在低努力水平下,延迟代价使Luna和Terra(我们同时评估的不进行推理的两个GPT模型)相对于未计时准确率分别损失38.8和42.8个百分点;一个判断较弱但速度更快的组件在无推理情况下达到了与Terra相似的综合得分。总体和家族曲线显示了这些权衡在何处发生变化,使评估的时间尺度依赖性可见。

英文摘要

Language models increasingly make real-time decisions in applications that apply the latest answer until a newer one arrives. A late answer can prolong an outdated decision, such as a call recorder still running while a customer reads out card details, an error offline accuracy misses. We make three contributions. First, we release StreamDecisionBench (SDB), a dataset of eight streaming scenarios in four application families, with executable reference decisions derived from public rules. Second, we propose an evaluation protocol and a metric, in-force accuracy: the share of time the applied decision is correct across update intervals of 0.5-8 s. It reflects accuracy and latency jointly, attributing each error to judgment, latency or both. Third, we evaluate thirteen single-model settings, and this attribution separates speed-limited from judgment-limited models: slower, more accurate models lose 42-51% of the time to outdated answers, a fast model 34% to wrong ones. We therefore test hybrids in which a slow model corrects a fast one; with the right pairing and configuration, a hybrid outperforms every single model. However, even the best evaluated system keeps a correct decision in force only about two-thirds of the time, leaving a substantial gap for real-time use.

发表机构

  • National Yang Ming Chiao Tung University(国立阳明交通大学)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑