arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.16607cs.IR

衡量工具增强型大语言模型中的决策尺度使用:一个对比性城市基准

Measuring Decision-Scale Use in Tool-Augmented LLMs: A Contrastive Urban Benchmark

发表机构佛罗里达大学
查看机构详情
  • University of Florida(佛罗里达大学)

机构由 AI 辅助整理,请以论文原文为准。

Ray Chen, Vivian Wong, Christan Grant

首次发表
浏览论文内容

中文总结 AI 辅助

本研究提出URBANCONTRASTIVEQA基准,测试工具增强型大语言模型能否基于历史基线比较城市活动异常程度,发现仅原始计数易误导,需暴露本地基线,并发布相关数据与脚本。

中文摘要 AI 辅助

城市决策支持常常询问某个特定地点的活动是否异常高或低,而不是哪个地点具有更大的原始计数。在一个安静社区中的二十次接送可能比机场的一百八十次更异常。我们引入了URBANCONTRASTIVEQA,一个基准测试,询问工具增强型语言模型是否能够进行这种基于基线的相对比较。每个项目将来自纽约市、芝加哥和西雅图的公共出行数据中的两个城市情境配对,并根据当前活动偏离该地点历史基线的程度进行标记。我们评估了六种指令调整模型在五种工具输出格式下的表现。仅使用原始计数时,模型常常选择较大的数字,即使该数字在其区域内较不异常。服务器计算的基线分数和序数标签提高了准确性,但收益因模型而异。对于异构的城市数据流,工具接口需要暴露本地基线,而不仅仅是活动量。我们发布了配对库、标签、评分脚本和数据卡。

英文摘要

Urban decision-support often asks whether activity is unusually high or low for a specific place, not which place has the larger raw count. Twenty pickups in a quiet neighborhood can be more abnormal than 180 at an airport. We introduce URBANCONTRASTIVEQA, a benchmark that asks whether tool-augmented language models can make this baseline-relative comparison. Each item pairs two urban situations from public mobility data in NYC, Chicago, and Seattle, labeled by how far current activity deviates from that place's historical baseline. We evaluate six instruction-tuned models under five tool-output formats. With only raw counts, models often pick the larger number even when it is less abnormal for its zone. Server-computed baseline scores and ordinal labels raise accuracy, but gains vary by model. For heterogeneous urban feeds, tool interfaces need to expose local baselines, not just activity volumes. We release the pair bank, labels, scoring scripts, and data card.

↑