IREA:基于中间表示的嵌入对齐用于规范性RAG
IREA: Intermediate Representation-based Embedding Alignment for Normative RAG
浏览论文内容
中文总结 AI 辅助
针对LLM伦理判断不足,提出IREA双向对齐方法,将查询与规范映射到共享情境-行为表示,提升规范性检索与伦理判断效果。
中文摘要 AI 辅助
大型语言模型(LLMs)在各种任务中表现出强大的性能,但在涉及伦理判断的问题上仍然存在困难。以往的研究尝试在伦理标准上训练LLMs,但伦理规范的多样性和相对性使其难以完全内化到模型参数中。作为替代方案,我们引入了规范性RAG,一种利用外部规范性知识支持伦理判断的检索增强方法。规范性检索涉及上下文丰富的叙事性查询与概括性规范陈述之间明显的非对称性。现有的事实性检索方法依赖于仅对查询进行扩展以形成类似文档的形式,不足以解决这种非对称性。因此,我们提出了基于中间表示的嵌入对齐(IREA),一种双向对齐方法,将两种文本类型映射到共享的情境-行为表示中。该表示以规范化形式捕获伦理上显著的情境和行为信息,减少表面层面的差异,并改善嵌入空间中的对齐。实验结果表明,IREA在多种设置下提高了规范性检索和下游伦理判断的性能,证明了双向对齐对于规范性RAG的有效性。
英文摘要
Large language models (LLMs) have shown strong performance across various tasks, but they still struggle with questions involving ethical judgment. Previous studies have attempted to train LLMs on ethical standards, but the diversity and relativity of ethical norms make them difficult to fully internalize in model parameters. As an alternative, we introduce normative RAG, a retrieval-augmented approach that supports ethical judgment using external normative knowledge. Normative retrieval involves a distinct asymmetry between context rich narrative queries and generalized normative statements. Existing factual retrieval methods rely on query-only expansion into a document-like form, making them insufficient for resolving this asymmetry. Therefore, we propose Intermediate Representation-based Embedding Alignment (IREA), a bidirectional alignment method that maps both text types into a shared situation-behavior representation. This representation captures ethically salient contextual and behavioral information in a normalized form, reducing surface-level discrepancies and improving alignment in the embedding space. Experimental results show that IREA improves normative retrieval and downstream ethical judgment across multiple settings, demonstrating the effectiveness of bidirectional alignment for normative RAG.
发表机构
- Konkuk University(建国大学)
- NAVER Cloud(NAVER云)
机构由 AI 辅助整理,请以论文原文为准。