发表机构
The University of Manchester(曼彻斯特大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
研究针对法医作者身份验证中似然比估计问题,提出平方根校正和单字校正两种归一化技术,无需校准模型,经实验评估,性能与逻辑回归校准相当,单字校正更优,还具有减少数据需求等优势,支持在法医环境采用。
AI 中文摘要
作者身份验证(AV)是确定两篇文本是否由同一作者撰写的任务。在法医背景下,AV证据的强度可用似然比来量化。大多数AV方法基于分数,从这些分数中得出校准良好的似然比需要单独的校准模型,这需要额外的相关案例数据,获取和准备这些数据通常很耗时。本研究提出了两种新的归一化技术,即平方根校正和单字校正,用于从AV方法LambdaG中得出似然比,而无需校准模型。这些校正旨在减轻因文本过长或重复性高而导致的证据强度高估问题。通过对数似然比成本(Cllr),在15个语料库和一系列文本长度(100 - 9500个词元)上针对逻辑回归校准评估性能。所提出的方法实现了与逻辑回归校准相当的性能,单字校正在约45%的测试中表现更优。消除训练校准模型的需求减少了数据需求、时间和复杂性,提高了法医文本比较的可及性和透明度。这种实证性能和实际优势的结合支持了所提出方法在法医环境中的采用。
英文摘要
Authorship verification (AV) is the task of determining whether two texts were written by the same author. In a forensic context, the strength of AV evidence can be quantified using likelihood ratios. Most AV methods are score-based and deriving well-calibrated likelihood ratios from these scores requires a separate calibration model. This, in turn, requires additional amounts of case-relevant data, which is often time-consuming to obtain and prepare. This study proposes two novel normalisation techniques, the Square Root Correction and the Hapax Correction, for deriving likelihood ratios from the AV method LambdaG without the need of a calibration model (Nini et al. 2026). These corrections are designed to mitigate the overestimation of evidential strength that may result from long or highly repetitive texts. Performance is evaluated against logistic regression calibration across fifteen corpora and a range of text lengths (100-9,500 tokens), using the log-likelihood ratio cost (Cllr). The proposed methods achieve performance comparable to logistic regression calibration, with the Hapax Correction outperforming it in approximately 45% of tests (weighted by corpora). Furthermore, performance was more frequently close (within 5%) when the Hapax Correction was outperformed by logistic regression calibration, compared with the reverse comparison. Eliminating the need to train a calibration model reduces data-requirements, time and complexity, thereby increasing the accessibility and transparency of forensic text comparison. This combination of empirical performance and practical advantages supports the adoption of the proposed methods in forensic settings.