KhatianDoc:诊断多模态大语言模型在孟加拉语法律土地记录上失败的人工验证基准
KhatianDoc: A Human-Verified Benchmark Diagnosing Multimodal LLM Failure on Bengali Legal Land Records
查看机构详情
- North South University(北南大学)
机构由 AI 辅助整理,请以论文原文为准。
浏览论文内容
中文总结 AI 辅助
本研究推出KhatianDoc基准,评估多模态大语言模型读取孟加拉语法律土地记录的能力,发现多数模型在关键任务上表现极差,为相关系统提供了验证数据。
中文摘要 AI 辅助
孟加拉国的土地所有权记录采用Ana-Ganda-Kora-Kranti-Til,这是一种十六进制位置分数系统,具有专用Unicode字形,无主流字体,且未被任何OCR流程或分词器覆盖。承载这些分数的手写记录RS Khatian是数百万块土地的权威权属记录,也是民事诉讼的常见对象,但尚无基准测试机器能否读取此类记录。我们推出KhatianDoc,这是一个由来自孟加拉国蒙希甘杰Vumi(土地)办公室的107份真实RS Khatian记录构建的四任务基准:符号识别、十六进制转十进制转换、结构化字段提取以及基于1634个问答对的法律文档问答。真实标注由人工转录,经土地法律从业者验证完全一致,并通过位置标记进行匿名化处理,保留了多跳问答所需的指代区分度。我们在固定零样本协议下评估了六个多模态大语言模型(8B至72B+,开源与闭源)。五个问答类别(占我们分层数据集的39.3%)中,所有模型均未返回正确答案;在算术任务上,所有输出数字的模型表现均逊于常数均值基线,精确匹配与近似匹配得分一致:存在去相关而非近似的情况。对我们自身指标的审计发现了两个相反方向的人为偏差:我们修正了一个拒绝评分错误,并报告了修正后的得分与原始得分,同时标记了一个被夸大的元数据指标作为上界。KhatianDoc所记录的并非性能差距,而是能力缺失,为未来系统提供了经过验证的真实标注。带有编辑后图像版本的代码和数据可公开获取。
英文摘要
Land ownership in Bangladesh is recorded in Ana-Ganda-Kora-Kranti-Til, a base-16 positional fraction system with dedicated Unicode glyphs, no mainstream font, and no coverage in any OCR pipeline or tokenizer. The handwritten records that carry these fractions, RS Khatians, are the authoritative title record for millions of parcels and a frequent subject of civil litigation, yet no benchmark has asked whether a machine can read one. We introduce KhatianDoc, a four-task benchmark built from 107 real RS Khatian records from the Vumi (land) Office of Munshiganj, Bangladesh: symbol recognition, base-16-to-decimal conversion, structured field extraction, and legal document question answering over 1,634 QA pairs. Ground truth was transcribed by hand, verified by a land-law practitioner to full agreement, and anonymized through positional tokens that keep the referential distinctions multi-hop questions depend on. We evaluate six multimodal LLMs (8B to 72B+, open and closed) under a fixed zero-shot protocol. Five QA categories, 39.3% of our stratified set, return zero correct answers from every model; on the arithmetic task, every model that emits a number does worse than a constant-mean baseline, with exact- and near-match scores coinciding: decorrelation, not approximation. Auditing our own metrics surfaced two artifacts in opposite directions: we correct a refusal-scoring bug and report the fixed scores beside the originals, and flag an inflated metadata metric as an upper bound. KhatianDoc documents not a performance gap but the absence of a capability, with verified ground truth for future systems. Code and data, with a redacted image release, are publicly available.