arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

CHiPS:罗马尼亚语文本中用于轻量级作者归属的字符直方图和位置信号

CHiPS: Character Histograms and Positional Signals for Lightweight Authorship Attribution in Romanian Texts

Sanda-Maria Avram, George C. Ţurcaş

arXiv 2607.22884首次发表:更新:

发表机构

Faculty of Mathematics and Computer Science, Babeş-Bolyai University(巴比什-波雅依大学数学与计算机科学学院)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

该研究针对罗马尼亚语文本提出CHiPS轻量级字符级作者归属方法,研究两种互补写作风格指纹,包括字符直方图分类器和位置信号分类器,还介绍了融合变体等,通过实验探讨受限字符证据在严格泄漏控制下的效果。

AI 中文摘要

我们提出了CHiPS,一种用于罗马尼亚语文本的轻量级字符级作者归属方法。所有实验均为封闭集:真实作者是训练数据中的候选作者之一。CHiPS研究了两种互补的写作风格指纹:CH-SVM,一种基于单字符边际分布的字符直方图分类器;以及FFT12-LR,一种位置信号分类器,它将选定的字符和标点类别表示为脉冲序列,并提取傅里叶/韦尔奇频谱描述符。我们还报告了CHiPS-F,一种泄漏安全的决策级融合变体,以及一个仅在折叠外预测上训练的可选前5列表重排器。该方法不需要分词、句法分析、预训练语言模型或变压器微调,并且在直方图组件中避免了n≥2的字符n元特征。在一个包含来自392个源文本组、由10位作者撰写的400个文件的锁定分组ROST分割上,进行源文本级评估和分组五折模型选择,CHiPS-F达到了0.9310的准确率和0.9341的宏F1。一个匹配但无限制的字符2-5克TF-IDF SVM比较器在相同的保留组上达到了1.0000的准确率和宏F1,因此该贡献并非声称具有最佳的分类准确率。相反,实验探讨了在严格的泄漏控制下,受限的、透明的字符证据能走多远。在ROSTories-cleaned上,一个包含来自1240个源文本组、由19位作者撰写的1248个文件的二级ROST重叠语料库上,相同的协议给出了CHiPS-R的0.8919准确率和0.8708宏F1。

英文摘要

We propose CHiPS, a lightweight character-level authorship attribution method for Romanian texts. All reported experiments are closed-set: the true author is one of the candidate authors in the training data. CHiPS studies two complementary fingerprints of writing style: CH-SVM, a character-histogram classifier based on one-character marginal distributions, and FFT12-LR, a positional-signal classifier that represents selected characters and punctuation classes as impulse trains (binary indicator sequences over character positions) and extracts Fourier/Welch spectral descriptors. We also report CHiPS-F, a leakage-safe decision-level fusion variant, and an optional top-5 listwise reranker trained only on out-of-fold predictions. The method requires no tokenization, syntactic analysis, pretrained language model, or transformer fine-tuning, and it avoids character $n$-gram features with $n \geq 2$ in the histogram component. On a locked grouped ROST split comprising 400 files from 392 source-text groups, written by 10 authors, with source-text-level evaluation and grouped five-fold model selection, CHiPS-F reaches 0.9310 accuracy and 0.9341 macro-F1. A matched but unrestricted character 2--5-gram TF--IDF SVM comparator reaches 1.0000 accuracy and macro-F1 on the same held-out groups, so the contribution is not a claim of best possible classification accuracy. Instead, the experiments ask how far restricted, transparent character evidence can go under strict leakage control. On ROSTories-cleaned, a secondary ROST-overlapping corpus comprising 1,248 files from 1,240 source-text groups, written by 19 authors, the same protocol gives 0.8919 accuracy and 0.8708 macro-F1 for CHiPS-R.

Comments17 pages, 12 tables

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑