arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

从盲编辑到经验证的修复:构建可信的用户端大语言模型智能体以实现网页可访问性

From Blind Edits to Verified Repair: Building Trustworthy User-Side LLM Agents for Web Accessibility

Lily Bundgaard Wanscher, Markus Heidemann Lorensen, Mohammed Ammad Shafiq, Mahyar Tourchi Moghaddam, Mina Alipour

arXiv 2608.24913首次发表:更新:

AI 中文总结

本研究构建了用户端LLM智能体的三个核心模块,通过双条件协议诊断盲编辑的缺陷,再结合验证修复工具实现了网页可访问性的可信修复,相关资源已公开。

AI 中文摘要

在浏览时对网页进行适配的用户端辅助智能体,可解决网站作者未修复的可访问性故障,而大语言模型使这类智能体成为可能。本文为此贡献了三个构建模块:其一为完整的隐私保护型浏览器智能体,即一款Chrome扩展程序,它会提取网页的样式表,将其精简以适配本地模型的上下文窗口,向模型请求用于解决WCAG及W3C认知可访问性指南中18项指标的附加CSS,并将结果可逆地注入当前网页;其二为双条件协议,该协议对收益与危害的测量同样谨慎,应用于6个小型开放权重模型(7B至14B),涉及10个违规密集型网站和10个高可访问性的实时网站。诊断结果虽令人警醒但十分精准:未经验证的生成在100次试验中(5个可生成可注入CSS的模型),改进与退化网页的比例相近(24次改进对20次退化),修复了排版却破坏了依赖感知的属性;其三针对上述诊断,提出了经验证的修复工具,将三语 seeded 违规基准与审计-注入-验证循环相结合,仅当违规严格减少时才接受变更,从结构上杜绝了自动化检查中的退化情况。在真实浏览器中,该工具检测到57个 seeded 违规且无假阳性,拒绝了126个对抗性有害候选变更,所有代码、提示、基准材料、汇总数据及验证日志均已公开。

英文摘要

Assistive agents that adapt web pages on the user's side, at the moment of browsing, could reach the accessibility failures that site authors leave unfixed, and large language models make such agents newly plausible. We contribute three building blocks toward that goal. The first is a complete, privacy-preserving browser agent: a Chrome extension that extracts a page's style sheets, condenses them to fit a local model's context window, asks the model for additive CSS addressing 18 metrics from WCAG and the W3C cognitive accessibility guidance, and injects the result reversibly into the live page. The second is a dual-condition protocol that measures harm as carefully as benefit, applied to six small open-weight models (7B to 14B) on ten violation-rich and ten highly accessible live sites. The diagnosis is sobering but precise: unverified generation improved and regressed pages at similar rates (24 improvements against 20 regressions across the 100 trials of the five models that produced injectable CSS), fixing typography while breaking perception-dependent properties. The third answers the diagnosis: a verified repair instrument pairing a trilingual seeded-violation benchmark with an audit-inject-verify loop that accepts a change only if violations strictly decrease, so regression on the automated checks is impossible by construction. In a real browser the instrument detects 57 of 57 seeded violations with no false positives and rejects 126 of 126 adversarially harmful candidates. All code, prompts, benchmark materials, aggregate data, and validation logs are released.

CommentsAccepted to be presented at ICME 2026 (5-9 October, Napoli, Italy) and to be published in ICMI companion proceedings by ACM

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑