arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

TeXFix-Bench:面向基于大语言模型的文档源代码修复的经验驱动多格式基准

TeXFix-Bench: An Empirically Grounded Multi-Format Benchmark for LLM-Based Document Source Repair

Prajwal S. Venkateshmurthy

arXiv 2608.07617首次发表:更新:

AI 中文总结

该研究构建了基于经验故障分类法的多格式文档修复基准TeXFix-Bench,通过实验发现仅编译成功率会高估修复质量,Typst修复难度高于LaTeX和Markdown,并发布了相关分类法与工具。

AI 中文摘要

科学技术写作依赖于必须可编译的标记源代码:LaTeX、Typst和Markdown流程会因缺失分隔符、不匹配的环境、损坏的导入或包冲突而失败。现有的文档修复评估通过临时编辑注入故障,缺乏经验故障模型。我们提出TeXFix-Bench,这是一个基于挖掘的故障分类法的、面向基于大语言模型的全源代码修复的多格式基准。对来自TeX Stack Exchange、GitHub提交和包文档的本地化硬崩溃LaTeX故障的扎根理论研究(168个经验证的故障,双开放编码κ=0.34)产生了一个18类分类法,实例化为DocMut:跨三种格式的48个AST感知算子。三模型跨基准测试显示,DocMut故障比相同种子上基于模式的突变难修复5.6-9.2个百分点,而真实错误案例研究(88个挖掘的人为崩溃,修复成功率67.0%)则将合成集从下方框定。我们从743个开源种子构建了10,437个实例,并在固定的零样本协议和提供商固定路由下评估了七个大语言模型,收集了48,651次尝试,总推理成本约为200美元。完整的6,613实例×7模型平衡矩阵确认了所有排名。固定引擎门产生了27.5分的意向治疗编译差距(56.7%-84.2%)。Typst比LaTeX和Markdown明显更难。对28,129次可编译修复的恢复预言机显示,13.6%-18.5%的可编译修复会实质性地修改文档文本,且恢复排名与编译排名存在差异:编译率最低的模型在其成功案例中恢复内容最佳。仅编译成功率会高估修复质量。我们发布了该分类法、DocMut和所有实验产物。

英文摘要

Scientific and technical writing depends on markup sources that must compile: LaTeX, Typst, and Markdown pipelines fail on missing delimiters, mismatched environments, broken imports, or package conflicts. Existing document-repair evaluations inject faults with ad-hoc edits that lack an empirical fault model. We present TeXFix-Bench, a multi-format benchmark for LLM-based full-source document repair grounded in a mined fault taxonomy. A Grounded-Theory study of localized hard-crash LaTeX faults from TeX Stack Exchange, GitHub commits, and package documentation (168 verified faults, dual open coding at $κ$=0.34) yields an 18-category taxonomy instantiated as DocMut: 48 AST-aware operators across three formats. A three-model cross-benchmark shows DocMut faults are 5.6-9.2 pp harder to repair than pattern-based mutations on the same seeds, and a real-error case study (88 mined human crashes, 67.0% repair success) brackets both synthetic sets from below. We construct 10,437 instances from 743 openly licensed seeds and evaluate seven LLMs under a fixed zero-shot protocol with provider-pinned routing, collecting 48,651 attempts at about USD 200 total inference cost. A complete 6,613-instance x 7-model balanced matrix confirms all rankings. A pinned engine gate yields a 27.5-point intention-to-treat compile spread (56.7-84.2%). Typst is markedly harder than LaTeX and Markdown. A restoration oracle over 28,129 compiling repairs shows that 13.6-18.5% of compiling repairs materially alter document text, and restoration rank diverges from compile rank: the model with the lowest compile rate restores content best among its successes. Compile success alone overstates repair quality. We release the taxonomy, DocMut, and all campaign artifacts.

Comments9 pages, 4 figures, 9 tables. Artifacts: https://doi.org/10.5281/zenodo.21831797

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑