基于MCP的AI驱动屏幕阅读器智能体的可访问性树标准化
MCP-Driven Accessibility Tree Standardization for AI-Powered Screen Reader Agents
浏览论文内容
中文总结 AI 辅助
本文提出基于MCP的统一可访问性层架构,解决LLM智能体跨平台可访问性感知的集成复杂性问题,为可访问AI智能体提供语义信息基础。
中文摘要 AI 辅助
与图形用户界面交互的大型语言模型(LLM)智能体越来越依赖原始截图或特定平台的可访问性应用程序编程接口(API)来感知界面状态。这两种方法对于辅助应用都存在局限性:基于截图的感知缺少屏幕阅读器所需的语义角色和关系,而Windows UI Automation、macOS Accessibility、Android AccessibilityService以及Web ARIA等特定平台API需要针对每个平台进行单独集成。本文提出一种架构,使用模型上下文协议(MCP)作为异构可访问性框架与基于LLM的辅助智能体之间的统一传输和模式层。MCP可访问性服务器通过与平台无关的表示形式公开与ARIA对齐的角色、标签、状态以及可聚焦元素层次结构,支持跨操作系统和应用的一致交互。该框架还引入一种MCP资源模型,用于在不同会话间持久存储用户的可访问性偏好。本文针对三个研究问题对该架构进行分析:可访问性树表示的协议可扩展性、可访问性树与基于截图的感知之间的延迟和语义保真度权衡,以及通过MCP资源实现持久可访问性配置文件的支持。本文未提供实证实现,而是提供一个概念框架,该框架通过对可访问性API、GUI智能体架构和MCP规范的比较分析得到支撑。分析表明,标准化的MCP可访问性层可降低特定平台的集成复杂性,同时保留可访问AI智能体所需的语义信息,为未来的实现和评估奠定基础。
英文摘要
Large language model (LLM) agents that interact with graphical user interfaces increasingly rely on either raw screenshots or platform-specific accessibility application programming interfaces (APIs) to perceive interface state. Both approaches have limitations for assistive applications: screenshot-based perception lacks the semantic roles and relationships required by screen readers, while platform-specific APIs such as Windows UI Automation, macOS Accessibility, Android AccessibilityService, and web ARIA require separate integrations for each platform. This paper proposes an architecture that uses the Model Context Protocol (MCP) as a unified transport and schema layer between heterogeneous accessibility frameworks and LLM-based assistive agents. An MCP accessibility server exposes ARIA-aligned roles, labels, states, and focusable-element hierarchies through a platform-independent representation, enabling consistent interaction across operating systems and applications. The framework also introduces an MCP resource model for persisting user accessibility preferences across sessions. The architecture is analyzed with respect to three research questions: protocol extensibility for accessibility-tree representation, latency and semantic fidelity trade-offs between accessibility trees and screenshot-based perception, and support for persistent accessibility profiles through MCP resources. Rather than presenting an empirical implementation, this work contributes a conceptual framework supported by comparative analysis of accessibility APIs, GUI agent architectures, and the MCP specification. The analysis suggests that a standardized MCP accessibility layer can reduce platform-specific integration complexity while preserving the semantic information required for accessible AI agents, providing a foundation for future implementation and evaluation.
发表机构
- University of Texas(得克萨斯大学)
机构由 AI 辅助整理,请以论文原文为准。