arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

FloodReasonBench:面向边缘端具身洪水响应的VLM推理分割基准

FloodReasonBench: Benchmarking VLM Reasoning Segmentation for Embodied Flood Response at the Edge

Rajat Bhattacharjya, Yoomee Jung, Minwoo Kim, Sing-Yao Wu, Eli Bozorgzadeh, Nalini Venkatasubramanian, Nikil Dutt

arXiv 2608.15410首次发表:更新:

发表机构

University of California, Irvine; Kookmin University(加州大学欧文分校; 国民大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

FloodReasonBench针对边缘端具身洪水响应,构建专用数据集并刻画推理分割流水线的性能权衡,为资源受限场景提供任务与系统层面的基准支持。

AI 中文摘要

推理分割使视觉语言模型(VLMs)能够将与任务相关的语言请求转换为像素级视觉 grounding,为具身智能体提供自然的感知接口。然而,现有基准大多聚焦于通用视觉场景,忽略了洪水响应平台面临的领域与资源约束。我们提出FloodReasonBench,这是一个面向边缘端具身洪水响应的VLM推理分割基准。其核心是FloodResponseSeg,一个由真实场景和响应相关目标构建的洪水专用推理分割数据集。除任务准确率外,该基准还刻画了轻量视觉编码、分层分割推理和压缩中间表示下的推理分割流水线。我们观察到,在通用预适配设置中存在强烈的分区依赖准确率变化,而经洪水适配的目标工作负载设计空间在各分区间呈现出显著更紧凑的准确率范围。在NVIDIA Jetson AGX Xavier上的评估进一步揭示了推理分割准确率、边缘端延迟、能耗与通信开销之间的权衡,支持在质量约束下选择边缘操作点。这些结果共同为资源受限的边缘端具身洪水响应的推理分割提供了任务与系统层面的刻画。

英文摘要

Reasoning segmentation enables vision-language models (VLMs) to translate mission-relevant language requests into pixel-level visual grounding, offering a natural perception interface for embodied agents. However, existing benchmarks largely focus on generic visual scenes and overlook the domain and resource constraints encountered in flood-response platforms. We present FloodReasonBench, a benchmark for VLM reasoning segmentation for embodied flood response at the edge. At its core, FloodReasonBench introduces FloodResponseSeg, a flood-specific reasoning-segmentation dataset constructed from real-world scenes and response-relevant targets. Beyond task accuracy, the benchmark characterizes reasoning-segmentation pipelines under lightweight visual encoding, hierarchical split inference, and compressed intermediate representations. We observe strong partition-dependent accuracy variation in the generic pre-adaptation setting, while the flood-adapted target-workload design space exhibits a substantially more compact accuracy range across partitions. Evaluation on an NVIDIA Jetson AGX Xavier further exposes the tradeoffs among reasoning-segmentation accuracy, edge-side latency, energy, and communication footprint, enabling quality-constrained selection of edge operating points. Together, these results provide a task- and system-level characterization of reasoning segmentation for resource-constrained embodied flood response at the edge.

CommentsPaper is currently under review. The code and dataset will be made public upon acceptance

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑