arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

MAP:面向现实场所的多模态可访问性规划基准

MAP: A Benchmark on Multimodal Accessibility Planning for Real World Places

Jason Armitage, Ioannis Tsochantaridis, Linda Mazzone, Chuqiao Yan, Srini Narayanan, Sarah Ebling

arXiv 2608.28384首次发表:更新:

发表机构

University of Zurich; Google DeepMind(苏黎世大学; 谷歌DeepMind)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

研究推出首个多模态可访问性规划基准MAP,含声明验证与视觉证据检索两项评估,支持动态信息下的AI系统对比,为服务可访问性需求用户的多模态AI提供评估框架。

AI 中文摘要

我们推出MAP,这是首个用于评估多模态AI系统作为助手的基准,这些助手服务有可访问性需求的用户规划现实场所的访问行程。在评估中,系统会收到验证或推荐满足可访问性要求的兴趣点的请求。MAP包含两项新颖评估:可访问性规划的声明验证,评估场所信息与所述可访问性特征是否匹配,并识别满足请求的可访问性特征的场所;可访问性规划的视觉证据检索,检查多模态AI系统能否为请求的场所和可访问性特征选择视觉证据。我们的方法支持在场所信息和可访问性信息可能随时间变化的场景中,通过定期评估系统并刷新基准的真实数据,来对比不同AI系统的性能。该基准基于自动评分及部分响应的人工评分构建。

英文摘要

We introduce MAP, the first benchmark to evaluate multimodal AI systems as assistants for users with accessibility requirements when planning visits to places in the real world. In our evaluation, systems are presented with requests to verify or recommend a point of interest meeting an accessibility requirement. MAP contains two novel assessments: Claim verification for accessibility planning assesses if information on places and stated accessibility features is supported and identifies places that satisfy requested accessibility features. Visual evidence retrieval for accessibility planning checks if a multimodal AI system can select visual evidence for the requested place and accessibility feature. Our methodology supports comparison of AI systems in a setting where place information and accessibility information can change over time by evaluating systems and refreshing ground truth data at scheduled times. The benchmark is based on automatic rating and human rating for a proportion of responses.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑