发表机构
AI Center, Faculty of Computer and Information Systems, Islamic University of Madinah; AI V&V Lab, King Fahd University of Petroleum and Minerals; Faculty of Computing and Information Technology, University of the Punjab(麦地那伊斯兰大学计算机与信息系统学院AI中心; 法赫德国王石油矿产大学AI验证与确认实验室; 旁遮普大学计算与信息技术学院)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对虚拟世界对盲人视障用户的可访问性挑战,提出MetaBlind多智能体架构,通过编排器筛选相关信息输出,以减少听觉过载并提升交互体验。
AI 中文摘要
虚拟世界如今承载着课堂、会议、研讨会、商店和社交场所,而它们所呈现的几乎每一次交互都假设用户能够扫描三维场景、跟随虚拟化身并阅读浮动面板。盲人和视障(BVI)用户只能依赖各自孤立解决单一任务的辅助工具:识别物体、读取文本、描述场景或规划路线。一个实时虚拟房间打破了这种模式:障碍物、发言者、手势、聊天、幻灯片和通知同时涌来,而一个将所有信息都叙述出来的工具只是将视觉障碍换成了听觉障碍。本文提出了MetaBlind,一种将非视觉访问分布在八个专门智能体上的架构,涵盖感知、导航、社交与物体交互、通信、安全与信任、记忆以及个性化,并在这些智能体与用户之间设置了一个可访问性编排器。智能体将候选信息发布到共享的可访问性上下文中,而不是直接向用户说话。编排器根据安全相关性、目标相关性、紧迫性、置信度、用户相关性以及预估的聆听负担对每个候选信息进行评分,然后仅通过语音、结构化音频或触觉输出发布其判定为当时相关的项目。我们为选择步骤给出了正式表述,将编排周期指定为一种算法,并定义了针对单智能体助手的评估协议。MetaBlind目前处于设计阶段,尚无原型测量或用户研究,该协议说明了哪些结果将支持该设计,哪些结果将反驳它。
英文摘要
Virtual worlds now host classrooms, meetings, conferences, shops, and social venues, and nearly every interaction they expose assumes a user who can scan a three-dimensional scene, follow avatars, and read floating panels. Blind and visually impaired (BVI) users are left with assistive tools that each solve one task in isolation: naming an object, reading text, describing a scene, or planning a route. A live virtual room defeats that model: obstacles, speakers, gestures, chat, slides, and notifications arrive together, and a tool that narrates all of them trades a visual barrier for an auditory one. This paper presents MetaBlind, an architecture that distributes nonvisual access across eight specialized agents, spanning perception, navigation, social and object interaction, communication, safety and trust, memory, and personalization, and that places an Accessibility Orchestrator between those agents and the user. Agents publish candidate information into a shared accessibility context instead of speaking to the user directly. The orchestrator scores each candidate on safety relevance, goal relevance, urgency, confidence, user relevance, and estimated listening load, then releases only the items it judges relevant at that moment through speech, structured audio, or haptic output. We give the selection step a formal statement, specify the orchestration cycle as an algorithm, and define an evaluation protocol against a single-agent assistant. MetaBlind is reported at the design stage, with no prototype measurement or user study, and the protocol states which outcomes would support the design and which would refute it.
CommentsPaper submitted in ICAAD Conference 2026