arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

预测工作流基准:评估语言模型在预算化预测工具下的决策能力

Forecast Workflow Bench: Evaluating Language-Model Decisions with Budgeted Forecast Tools

Shunya Nagashima

arXiv 2609.27385首次发表:更新:

发表机构

Neurogica Inc.(Neurogica 公司)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

FWBench通过1,251个案例评估语言模型在预算约束下选择和使用时间序列预测工具进行决策的能力,发现GPT-6 Astra以2.5%预算实现优于固定策略的决策质量。

AI 中文摘要

时间序列基础模型(TSFMs)为运营决策提供预测,但仅凭准确性并不能决定其价值。评估使用这些模型的智能体需要衡量决策质量和预测成本。FWBench在1,251个电力和自行车租赁案例上,使用固定预测工具和模拟容量合同评估了这一能力。智能体选择模型、历史数据和预测范围,然后提交容量以最小化规定的损失-成本目标。我们评估了两种托管配置和八种本地配置,包括小型语言模型,并测试了有无TSFMs的本地模型。GPT-6 Astra选择性地购买廉价的短视界预测,仅使用预算的2.5%,在三种损失-成本权重下对节省的决策进行评分时,其表现优于固定策略。FWBench能够可复现地评估语言模型如何在成本约束下选择和利用时间序列预测来做出决策。

英文摘要

Time-series foundation models (TSFMs) provide forecasts for operational decisions, but accuracy alone does not determine their value. Evaluating agents that use these models requires measuring decision quality and forecast cost. FWBench evaluates this capability on 1,251 electricity and cycle-hire cases using fixed forecast tools and simulated capacity contracts. Agents select models, histories and horizons, then submit capacities to minimize a stated loss-cost objective. We evaluated two hosted and eight local configurations, including small language models, and tested local models with and without TSFMs. GPT-6 Astra bought inexpensive short-horizon forecasts selectively, using 2.5% of the budget, and outperformed fixed policies when the saved decisions were scored with three loss-cost weightings. FWBench enables reproducible evaluation of how language models select and use time-series forecasts to make decisions under cost constraints.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑