发表机构
The University of Hong Kong; University of Science and Technology of China(香港大学; 中国科学技术大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本文针对通用文本嵌入在任务区分上的不足,提出PGCL方法,通过多流结合的正则化模块生成任务适配表示,在三个代理任务数据集上均优于原始冻结嵌入,明确了方法适用边界并修订了理论分析。
AI 中文摘要
可靠的事后评估需判断已生成文本是否在生成后满足目标标准,本文在冻结嵌入的聚焦场景下,采用毒性检测、细粒度情感分类和有序评论评分等原则评估代理任务开展研究。通用文本嵌入被广泛用于此类任务,但宽泛的语义相似性会将语义相似但任务不同的示例置于表示空间的重叠区域。我们提出原型引导的对比学习(Prototype-Guided Contrastive Learning, PGCL),这是一种构建于冻结文本嵌入之上的原型引导几何正则化模块,该模块结合语义流、原型锚点注意力流、监督对比学习、基于偏移的原型间隔正则化以及流正则化,生成紧凑的任务适配表示,无需更新基础编码器。控制实验表明,PGCL在所有三个数据集上均优于原始冻结嵌入,且在AmazonReviews数据集上与直接基线的差距最明显,在GoEmotions和ToxicComment数据集上与强大的直接冻结度量学习基线具有竞争力。我们还添加了监督残差适配器、编码器-LoRA、全微调、目标消融、敏感性分析以及完全记录的少样本大语言模型协议诊断,以明确该方法的适用边界。理论分析被修订为在原型映射空间的明确假设下对原型间隔行为的充分条件说明,而非无条件训练或最终嵌入分离的保证。
英文摘要
Reliable post-hoc evaluation asks whether already generated text satisfies a target criterion after generation. In this paper we study a focused frozen-embedding setting using principle-evaluation proxy tasks: toxicity detection, fine-grained emotion categorization, and ordinal review rating. General-purpose text embeddings are widely deployed for such tasks, but broad semantic similarity can place semantically similar yet task-distinct examples in overlapping regions of the representation space. We introduce Prototype-Guided Contrastive Learning (PGCL), a prototype-guided geometric regularization module built on top of frozen text embeddings. The module combines a semantic stream, a prototype-anchor attention stream, supervised contrastive learning, offset-based prototype-margin regularization, and stream regularization to produce a compact task-adapted representation without updating the base encoder. Controlled experiments show that PGCL improves over raw frozen embeddings on all three datasets and gives the clearest direct-baseline margin on AmazonReviews, while remaining competitive with strong direct frozen metric-learning baselines on GoEmotions and ToxicComment. We also add supervised residual-adapter, encoder-LoRA, full fine-tuning, objective ablation, sensitivity, and fully logged few-shot LLM protocol diagnostics to define the boundary of the claim. The theoretical analysis is revised as a sufficient-condition account for prototype-margin behavior under explicit assumptions in the prototype-mapping space, rather than as an unconditional training or final-embedding separation guarantee.
CommentsAccepted for publication in Transactions on Machine Learning Research (TMLR). 27 pages