BrailleBench:探究大型语言模型的多准则盲文理解能力
BrailleBench: Investigating Multi-Criteria Braille Comprehension in Large Language Models
浏览论文内容
中文总结 AI 辅助
本研究推出盲文理解基准BrailleBench,评估6款LLMs的盲文处理能力,发现其印刷体英语能力与盲文可及性存在差距,为盲文AI系统开发提供指导。
中文摘要 AI 辅助
尽管大型语言模型(LLMs)为知识获取和计算辅助提供了媒介,但其能力应同样惠及弱势群体。然而,现有AI系统是否足够包容,让盲人和双盲用户通过盲文获得相同功能尚不清楚,盲文的符号、缩约形式和数字表示对模型理解提出了独特要求。为此,我们推出BrailleBench,这是一个从不同准则评估LLMs盲文理解能力的基准。BrailleBench整合了来自五个数据集的5570个实例,涵盖数学、常识和多跳问答,涉及英语以及1级和2级盲文。我们设计了不同的配置,以了解这些系统能否理解盲文撰写的内容、用盲文表达答案以及完成端到端的盲文交互。为确保质量并防止评估偏差,该基准通过确定性的、经专家审核的流程构建,采用自主创建的Braille Toolkit,未使用任何由LLMs生成的数据实例。我们从多个方面评估了六个具有代表性的LLMs。结果显示,印刷体英语能力与盲文可及性之间存在持续差距;盲文理解和表达呈不对称性,其中2级盲文在输入端相比1级盲文尤其脆弱,而全盲文请求进一步降低了性能。这些实验观察为未来盲文AI系统的开发提供了宝贵指导,BrailleBench中的所有相关资源均公开供未来研究使用。
英文摘要
Although Large language models (LLMs) mediate access to knowledge and computational assistance, their capabilities should benefit vulnerable groups in the same way. However, it is unclear whether existing AI systems are inclusive enough for blind and deafblind users to access the same functionality through Braille, whose indicators, contractions, and digital representations introduce distinct requirements for model comprehension. To this end, we introduce BrailleBench, a benchmark for evaluating LLMs in Braille comprehension from different Criteria. BrailleBench aligns 5,570 instances from five datasets, including mathematics, commonsense, and multi-hop question answering across English and Braille Grades 1 and 2. Different configurations are designed to understand whether the systems can comprehend Braille-authored content, express answers in Braille, and complete end-to-end Braille interaction. To ensure the quality and prevent evaluation bias, the benchmark is built through a deterministic, expert-reviewed pipeline via a self-created Braille Toolkit without using any data instances generated by LLMs. We evaluate six representative LLMs from various aspects. The results reveal a persistent gap between print-English capability and Braille accessibility. Braille understanding and expression are asymmetric, where Grade 2 is especially fragile on the input side compared to Grade 1, and fully Braille requests further reduce performance. The experimental observations provide valuable guidance for the development of future Braille AI systems. All related resources in BrailleBench are publicly available for future research.
发表机构
- Rochester Institute of Technology(罗切斯特理工学院)
- Amazon.com, Inc.(亚马逊公司)
- Virginia Tech(弗吉尼亚理工大学)
机构由 AI 辅助整理,请以论文原文为准。