arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

图卢法律理解中的跨语言迁移:依赖文字脚本的性能提升与检索增强生成(RAG)引发的知识冲突

Cross Lingual Transfer in Tulu Legal Comprehension: Script-Dependent Improvement and RAG-Induced Knowledge Conflict

Sindhu Shetty, Spurthi Setty, Natan Vidra

arXiv 2608.28645首次发表:更新:

发表机构

Stanford Legal Design Lab(斯坦福法律设计实验室)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本文针对低资源语言图卢语的法律理解任务,测试了Llama3等三种模型的跨语言迁移效果,发现其性能依赖文字脚本,RAG框架会引发知识冲突,相关评估技术可推广至低资源多语言RAG场景。

AI 中文摘要

缺乏充足训练语料的低资源语言常借助相关的高资源语言作为理解的支架,但仍需开发严格的评估方法以识别跨语言低资源环境中模型的失效情况。以法律领域为背景,本文测试了Llama3、Hex-1、Sarvam三种模型对低资源达罗毗荼语图卢语(Tulu)撰写的法律投诉的分类能力。在达罗毗荼文字脚本间转写查询使模型无需大规模训练即可初步理解投诉内容,但理解程度高度依赖文字脚本(其中卡纳达语(Kannada,另一种相对低资源语言)呈现出最强的正向趋势)。基于卡纳达语法律论文语料库通过RAG框架进行检索产生了混合结果:部分模型在特定条件下理解能力呈现弱正向趋势,但模型失效时通常源于两个维度:事实替换(聚焦于特定段落摘录而扭曲推理)和虚构(无查询或语料依据的幻觉)。在低资源领域内,结果表明模型的信息解析及后续推理是推理失效的来源,而非语料内容。依赖脚本的理解能力与RAG鲁棒性似乎相伴而生,推理轨迹分析与部署的统计诚实框架进一步支持了这一点——这些技术更广泛适用于低资源多语言RAG评估。

英文摘要

Low-resource languages without an adequate training corpus often use a related, higher-resource language as a scaffold for comprehension. Still, there is a need to develop rigorous evaluation methods to identify when models fail in cross lingual low-resource environments. Using the legal domain as a backdrop, three models (Llama3, Hex-1, Sarvam) were tested on the ability to classify legal complaints written in a low resource Dravidian language (Tulu). Transliterating queries across Dravidian scripts allowed models to gain a preliminary understanding of speakers' complaints without the use of wide scale training, though the level of comprehension was heavily script dependent (with Kannada - another relatively low-resource language - producing the strongest positive trend). Retrieving from a corpus of Kannada legal papers across a RAG framework caused mixed results. Some models had a weak positive trend in comprehension under certain conditions, but when models failed, it was often across two axes: fact substitution (fixating on specific passage excerpts that skewed reasoning) and confabulation (hallucination that had no basis in either query or corpus). Within low resource domains, results identify the model's parsing of information and subsequent reasoning as the source of reasoning failure, rather than corpus contents. Script-dependent comprehension and RAG robustness also seem to travel together. This is further supported by the reasoning-trace analysis and a statistical-honesty framework deployed - techniques that are more broadly applicable to low-resource multilingual RAG evaluation.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑