arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.27376cs.CL

跨语言法律问答:越南劳动法的检索、翻译与验证器引导修正

Cross-Lingual Legal QA for Vietnamese Labour Law: Retrieval, Translation, and Verifier-Guided Correction

发表机构维新大学 · 兰卡斯特大学 · 卡迪夫大学
查看机构详情
  • VinUniversity(维新大学)
  • Lancaster University(兰卡斯特大学)
  • Cardiff University(卡迪夫大学)

机构由 AI 辅助整理,请以论文原文为准。

Nguyen Minh Chi, Mo El-Haj, Nguyen Ha Thanh, Dawn Knight, Paul Rayson

首次发表
浏览论文内容

中文总结 AI 辅助

针对越南劳动法的跨语言法律问答,构建双语评估套件,提出验证器引导的流水线,发现密集检索优于稀疏检索,验证器修正提升引用保留,但自动诊断与人类判断不完全一致。

中文摘要 AI 辅助

跨语言法律问答必须在跨语言检索法规的同时防止无依据的法律主张。我们引入了一个包含231对越南语-英语问答对的双语评估套件,这些问答对源自越南劳动法。其中,75对额外标注了五种具有挑战性的法律推理现象。我们评估了一个验证器引导的流水线,该流水线将答案分解为主张,检查引用的可达性和蕴含关系,并纠正引用失败和矛盾。我们还引入了六个自动诊断指标,用于评估对检索证据的忠实性,涵盖引用、情态、例外、程序、结论和证据支持。实验表明,学习型稀疏检索在英语到越南语的检索中表现不佳(R@5≈0.032),而密集检索达到0.358,并略优于混合检索。在我们的受控比较和支持性敏感性分析中,翻译位置对这些自动诊断指标没有统计上可检测的影响。验证器引导的修正将系统层面的引用保留提高了0.022-0.034,但在其余维度上没有产生可靠的增益。人工评估进一步表明,自动诊断指标与人类对答案质量的判断并不完全一致。

英文摘要

Cross-lingual legal question answering must retrieve statutes across languages while preventing unsupported legal claims. We introduce a bilingual evaluation suite of 231 Vietnamese--English question--answer pairs from Vietnamese labour law. Of these, 75 are additionally annotated for five challenging legal reasoning phenomena. We evaluate a verifier-guided pipeline that decomposes answers into claims, checks citation reachability and entailment, and corrects citation failures and contradictions. We also introduce six automatic diagnostics for faithfulness to retrieved evidence, covering citations, modality, exceptions, procedures, conclusions, and evidential support. Experiments show that learned-sparse retrieval performs poorly for English-to-Vietnamese retrieval (R@5~=~0.032), whereas dense retrieval reaches 0.358 and slightly outperforms hybrid retrieval. Translation placement has no statistically detectable effect on these automatic diagnostics in our controlled comparison and supporting sensitivity analyses. Verifier-guided correction improves citation preservation by $0.022$--$0.034$ at the system level but produces no reliable gains in the remaining dimensions. Human evaluation further shows that the automatic diagnostics do not fully align with human judgements of answer quality.

补充信息

↑