StreamReason-Bench:大型语言模型能否推理事件时间流处理语义?
Can Large Language Models Reason about Event-Time Stream-Processing Semantics?
查看机构详情
- Northeastern University(东北大学)
机构由 AI 辅助整理,请以论文原文为准。
浏览论文内容
中文总结 AI 辅助
该研究构建了StreamReason-Bench基准测试,发现LLM在事件时间流处理语义推理任务上表现不佳,思维链可提升其性能,难点在于事件时间与延迟数据处理。
中文摘要 AI 辅助
流处理系统越来越多地将工作交给大型语言模型(LLM),包括编写管道、分类警报、读取日志等,而所有这些工作都假设模型知晓事件时间流处理的行为。我们直接验证这一假设。StreamReason-Bench要求模型充当事件时间流处理器:给定一个窗口查询和一组乱序事件,它必须报告哪些窗口触发(及其聚合结果),以及哪些事件因延迟被丢弃。答案键来自Dataflow模型语义的小型参考实现,因此我们可以通过部分信用行F1进行精确评分,无需运行引擎。在涵盖滚动窗口、跳跃窗口、会话窗口和处理时间窗口的600个生成条目上,模型在事件时间任务上表现不佳。被要求直接回答时,没有任何真正遵循指令的模型能达到34%的精确匹配;思维链(CoT)方法使其中几个模型的精确匹配率大致翻倍(GPT-4o从0.34升至0.48),而只有一个默认进行推理的前沿模型接近解决该任务(精确匹配率达0.85)。作为对照的处理时间窗口(无水印且无延迟事件)几乎被所有有能力的模型解决。这一差距表明,难点在于事件时间和延迟数据处理,而非窗口或算术。按窗口类型对错误进行排序也得出了相同结论:延迟数据错误在事件时间窗口中占主导,在对照任务中消失,且会话窗口大多在会话边界的确定上失败。
英文摘要
Streaming systems increasingly hand work to large language models (LLMs): writing pipelines, triaging alerts, reading logs. All of it assumes the model knows how event-time stream processing behaves, and we test that assumption directly. StreamReason-Bench asks a model to stand in for an event-time stream processor. Given a windowed query and a stream of out-of-order events, it reports which windows fire, with their aggregates, and which events are dropped as late. The answer key comes from a small reference implementation of Dataflow-model semantics, so we can grade exactly, and with a partial-credit row-F1, without running an engine. On 600 generated items covering tumbling, hopping, session, and processing-time windows, the models do poorly on event time. When told to answer directly, no model that actually follows the instruction clears 34% exact match; chain-of-thought (CoT) roughly doubles that for several of them (GPT-4o goes from 0.34 to 0.48), and only one frontier model that reasons by default comes near solving the set (0.85). A processing-time control, with no watermarks and nothing late, is almost solved by every capable model. The gap points to event-time and late-data handling, not windowing or arithmetic, as the hard part. Sorting errors by window type tells the same story: late-data mistakes dominate the event-time windows and vanish on the control, while session windows mostly fail on where the session boundaries fall.