发表机构
Luleå University of Technology(吕勒奥理工大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本研究比较了基于LLM和嵌入的作者表示策略,提出两阶段LISA框架,在零样本作者归属中取得最优性能,凸显作者表示的关键作用。
AI 中文摘要
作者归属(AA)需要捕捉细粒度的文体特征,这使得在零样本(ZS)设置下(即没有任务特定监督可用时)尤其具有挑战性。在本工作中,我们通过评估一个仅标签提示基线以及三种作者表示策略:代表性写作样本、LLM生成的描述和风格嵌入(LISA),来研究作者表示对零样本作者归属的影响。前三种方法使用LLM提示进行归属,而基于嵌入的方法使用风格嵌入和余弦相似度。我们研究了提示设计的影响,并提出了一种两阶段基于嵌入的归属框架,该框架结合了候选空间缩减和嵌入维度选择。结果表明,仅标签的零样本作者归属是无效的,而纳入作者特定表示则持续提高归属性能。在评估的方法中,所提出的两阶段LISA框架实现了最强的整体性能,而LLM生成的风格描述提供了更紧凑的作者风格表示,但以牺牲部分归属性能为代价。这些发现证明了作者表示在零样本作者归属中的重要性,同时表明当前开源LLM在没有更有效的表示学习的情况下,仍不足以进行稳健的归属。
英文摘要
Authorship Attribution (AA) requires capturing fine-grained stylistic characteristics, making it particularly challenging in zero-shot (ZS) settings where no task-specific supervision is available. In this work, we investigate the effect of author representations on ZS AA by evaluating a label-only prompting baseline together with three author representation strategies: representative writing samples, LLM-generated descriptions, and style embeddings (LISA). The first three approaches perform attribution using LLM prompting, while the embedding-based approach uses style embeddings with cosine similarity. We investigate the influence of prompt design and propose a two-stage embedding-based attribution framework that combines candidate space reduction with embedding-dimension selection. The results show that label-only ZS AA is ineffective, while incorporating author-specific representations consistently improves attribution performance. Among the evaluated approaches, the proposed two-stage LISA framework achieves the strongest overall performance, whereas LLM-generated style descriptions provide a substantially more compact representation of author style at the cost of some attribution performance. These findings demonstrate the importance of author representation in ZS AA, while indicating that current open-source LLMs remain insufficient for robust attribution without more effective representation learning.