arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.23417cs.SEcs.CV

Omni2Web:音视频网站开发基准测试

Omni2Web: Benchmarking Audiovisual Website Development

Minghao Han, Zhenghao Xing, Xize Cheng, Yuxuan Wang, Junming Lin, Ling Wang, Yinsong Yan, Yunfei Chu, Qize Yang, Jin Xu

首次发表
浏览论文内容

中文总结 AI 辅助

针对屏幕录制网页编辑请求中弱指示表达导致的意图恢复难题,提出双语基准Omni2Web,含918实例与三轨道评估,发现现有模型在意图恢复与代码执行上仍有巨大提升空间。

中文摘要 AI 辅助

屏幕录制的网页编辑请求包含诸如“这个”和“那里”等弱指示表达,其指代对象依赖于语音、光标轨迹、页面状态和编辑历史。此类请求需要超越许多现有网页编辑基准所假设的显式规范的意图恢复。我们引入了Omni2Web,一个包含918个实例、跨越13,907个编辑步骤的双语基准。它定义了三个互补的轨道:直接编辑(Direct Editing)从录制中评估网页编辑,指令恢复(Instruction Recovery)衡量显式意图恢复,指令效用(Instruction Utility)测试恢复的指令能否驱动固定的代码执行器。我们评估了17个开源和闭源模型。最佳模型在直接编辑的编辑保真度评分(EFS)上达到51.17,在指令恢复评分(IRS)上达到49.14;在固定执行器下,最强的恢复指令达到51.08 EFS,仍远低于使用oracle指令获得的89.69 EFS。逐步分析表明,正确的接地并不能保证成功的编辑,而一些Omni模型恢复的指令被固定编码模型执行的效果显著优于其直接编辑。受控消融进一步证明了时间对齐的音视频证据的价值,而替代评判器保持了领先者和大致排序。这些发现共同揭示了多模态意图恢复和代码执行中的巨大提升空间,并强调了将Omni重写器与编码模型配对的前景。

英文摘要

Screen-recorded web editing requests contain weak deictic expressions such as ``this'' and ``there,'' whose referents depend on speech, cursor trajectories, page state, and edit history. Such requests require intent recovery beyond the explicit specifications assumed by many existing web-editing benchmarks. We introduce Omni2Web, a bilingual benchmark of 918 instances spanning 13,907 edit steps. It defines three complementary tracks: Direct Editing evaluates webpage editing from recordings, Instruction Recovery measures explicit intent recovery, and Instruction Utility tests whether recovered instructions can drive a fixed code executor. We evaluate 17 open- and closed-source models. The best models attain 51.17 on the Edit Fidelity Score (EFS) for Direct Editing and 49.14 on the Instruction Recovery Score (IRS); under the fixed executor, the strongest recovered instructions reach 51.08 EFS, still far below the 89.69 EFS obtained with oracle instructions. Step-level analyses show that correct grounding does not guarantee successful edits, while some Omni models recover instructions that the fixed coding model executes substantially better than their direct edits. Controlled ablations further demonstrate the value of temporally aligned audiovisual evidence, while alternative judges preserve the leader and broad ordering. Together, these findings reveal substantial headroom in multimodal intent recovery and code execution and highlight the promise of pairing Omni rewriters with coding models.

发表机构

  • FDU(复旦大学)
  • Alibaba Group(阿里巴巴集团)
  • CUHK(香港中文大学)
  • THU(清华大学)
  • PolyU(香港理工大学)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑