arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

从图表到文本:模型感知的多模态解释作为可访问非视觉交互的基础

From Plots to Words: Model-Aware Multimodal Explanations as a Foundation for Accessible, Non-Visual Interaction

Nur Keleşoğlu, Łukasz Sobczak, Joanna Domańska

arXiv 2608.24910首次发表:更新:

发表机构

Institute of Theoretical and Applied Informatics, PAS(波兰科学院理论与应用信息研究所)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

该研究提出多智能体框架,将时间序列预测的视觉输出转为模型感知文本解释,可提升非视觉交互的可访问性,其可解释配置使解释质量最高提升32%。

AI 中文摘要

多模态大语言模型正越来越多地被用于交互式系统中,但确保跨异构模态的一致、可信推理仍然具有挑战性。我们提出了一种上下文感知的多智能体框架,该框架整合了文本查询、数值数据、视觉表示和模型衍生信号,用于可解释的时间序列预测。其一个显著特征是,它将主要为视觉形式的预测输出(例如趋势图)转换为结构化的、模型感知的文本解释。我们认为,这使得该方法成为非视觉、可访问交互的天然基础,对盲人和视力受损用户尤为重要,因为以图表为中心的界面对这些用户基本无法访问。该框架支持三种逐步丰富的流程(基线、可解释、可解释性),能够系统比较单模态、感知驱动和模型感知的响应。在使用基于大语言模型的评判者作为人类评估的早期代理的探索性评估中,可解释配置的整体解释质量比数值基线提高了多达32%,在可信度和模型感知方面也有显著提升。我们认为,针对目标用户(包括屏幕阅读器和语音界面用户)的以用户为中心的验证是必要的下一步,而非在此处已确立的主张。

英文摘要

Multimodal large language models are increasingly used in interactive systems, yet ensuring consistent, trustworthy reasoning across heterogeneous modalities remains challenging. We present a context-aware, multi-agent framework that integrates textual queries, numerical data, visual representations, and model-derived signals for explainable time-series forecasting. A distinctive feature is that it turns predominantly visual forecasting outputs (e.g., trend plots) into structured, model-aware textual explanations. We argue that this makes the approach a natural foundation for non-visual, accessible interaction of particular relevance to blind and visually impaired users, for whom plot-centric interfaces are largely inaccessible. The framework supports three progressively richer pipelines (baseline, interpretable, explainable), enabling systematic comparison of unimodal, perception-driven, and model-aware responses. In an exploratory evaluation using an LLM-based judge as an early-stage proxy for human assessment, the explainable configuration improves overall explanation quality by up to 32% over a numerical baseline, with notable gains in trustworthiness and model awareness. We position user-centered validation with target users, including screen-reader and speech-interface users, as the essential next step rather than a claim established here.

DOI:10.1145/3776591.3833865

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑