逐词阅读法律问题:2144条越南法律标题的嵌入轨迹
Reading a Legal Question Word by Word: Embedding Trajectories of 2,144 Vietnamese Legal Headlines
浏览论文内容
中文总结 AI 辅助
本研究逐词追踪法律标题嵌入轨迹,发现黄金文章在6-7个内容词后锁定,数字影响最大,并归纳出六种锁定原型,提出上下文调制的加法行走模型。
中文摘要 AI 辅助
密集检索器将一个问题编码为一个向量,但问题是逐词到达的。我们使用Nemotron-3-Embed 8B/1B和Qwen3-Embedding 8B/0.6B,逐词阅读来自Thu Vien Phap Luat(越南法律图书馆)的2144条保留标题,对65,444个前缀与20,034篇文章进行编码,此外还对来自1,112个多问题标题的3,438个子问题的每个前缀以及168个答案的每个前缀进行了编码。(i) 在每个编码器中,黄金文章在读取疑问框架之前,经过中位数6-7个内容词后即达到排名第1,并在78-85%的情况下保持到结束。(ii) 在多问题标题中,锁定在94-98%的情况下出现在第一个子问题内部;第二个子问题在89-95%的情况下使排名保持不变;单独编码时,第二个子问题在42-58%的情况下达到排名第1,而第一个子问题为91-96%,且锁定词相同(95-97%相同)。(iii) 数字、日期和工具标识符使嵌入移动的距离是内容词的两倍,是疑问词的四倍;72-78%的步骤朝向黄金文章移动,而结尾的疑问框架在95-99%的标题中朝相反方向移动。(iv) 排名/余弦聚类产生六种原型(即时、典型、不稳定、延迟、永不锁定),这些原型因法律领域和形式而异(卡方p < 1e-8):房地产和诉讼标题从不锁定在数字上;环境和会计标题有三分之一的时间锁定在数字上。(v) 逐词阅读的答案在8-16个词后检索到其文章,并在83-89%的情况下按提问顺序回答子问题。(vi) 一个词的步骤在标题间保持一致的方向(余弦0.25-0.33;数字为0.44-0.60);前面的问题将该步骤旋转约60度,问候语旋转约30度;步骤随i^{-0.8}缩小;双问题标题在其两个问题的线性混合的12-17度范围内。我们称此为上下文调制的加法行走。
英文摘要
A dense retriever encodes a question as one vector, but the question arrives one word at a time. We read 2,144 held-out headlines from Thu Vien Phap Luat (Vietnamese legal library) word by word with Nemotron-3-Embed 8B/1B and Qwen3-Embedding 8B/0.6B, encoding 65,444 prefixes against 20,034 articles, plus every prefix of 3,438 sub-questions from 1,112 multi-question headlines and of 168 answers. (i) The gold article becomes rank 1 after a median of 6-7 content words in every encoder, before the interrogative frame is read, and stays there to the end in 78-85% of cases. (ii) In a multi-question headline the lock is inside the first sub-question 94-98% of the time; the second leaves rank unchanged in 89-95%; encoded alone, the second reaches rank 1 in 42-58% vs 91-96% for the first, at the same lock word (95-97% identical). (iii) Numbers, dates and instrument identifiers move the embedding twice as far as content words and four times as far as interrogative words; 72-78% of steps move toward the gold article, and the closing interrogative frame moves against that direction in 95-99% of headlines. (iv) Rank/cosine clustering yields six archetypes (instant, typical, unstable, late, never-locking) that differ by legal area and form (chi-squared p < 1e-8): real-estate and litigation headlines never lock on a number; environmental and accounting headlines do so a third of the time. (v) An answer read word by word retrieves its article after 8-16 words and addresses the sub-questions in order asked in 83-89% of cases. (vi) A word's step keeps a consistent direction across headlines (cosine 0.25-0.33; 0.44-0.60 for numbers); a preceding question rotates that step by about 60 degrees and a greeting by about 30 degrees; steps shrink as i^{-0.8}; and a two-question headline is within 12-17 degrees of a linear mix of its two questions. We call this a context-modulated additive walk.