arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.20393cs.CLcs.AI

面向智能体对话AI的风格可控且事实保留生成的知识图谱门控去事实化方法

Knowledge-Graph-Gated Defactualization for Style-Controllable and Fact-Preserving Generation in Agentic Conversational AI

Tanmay Kumar Shrivastava, Darsh Rohit Nandu, Rajesh Kumar Mundotiya

首次发表
浏览论文内容

中文总结 AI 辅助

针对智能体对话AI的事实保留与风格可控生成需求,提出DSR知识工程框架,结合KG与激活引导,在LLaMA系列模型上验证其可提升实体恢复率并保持风格控制,无需微调。

中文摘要 AI 辅助

部署在客服等事实敏感应用中的智能体大语言模型(LLM)必须同时兼顾事实正确性与生成响应的风格可控性。激活引导技术通过扰动隐藏表征实现无需微调的风格控制,但它缺乏区分可验证事实与风格内容的显式机制,会导致语义泄漏。针对该挑战,本文提出了\textit{去事实化-引导-再水化}(Defactualize-Steer-Rehydrate, DSR)框架,这是一种知识工程框架,将带类型、显著性加权的知识图谱(KG)与激活引导相结合。DSR通过分层正则表达式、命名实体识别(NER)或词汇分类器流水线提取显著性实体,在引导前用带类型占位符替换这些实体,生成后通过显著性引导的再水化过程确定性恢复已验证的实体值。本文在6款LLaMA系列模型(参数规模1B至13B)、600个A2A生成的客服案例(共1200次生成结果)上对DSR进行评估,并开展了专门的KG ablation研究。结果显示,DSR相比仅用引导的基线方法,已验证实体恢复率显著提升(Cohen's d=0.225,经Bonferroni校正的p值为1.0×10⁻⁴),尽管绝对恢复率仍较低,但它在不同模型家族中均能保持有效的风格控制。分层可分性与引导强度诊断进一步揭示了表征级引导与事实基础间此前未被探索的交互作用。这些结果表明,显式知识工程可在无需模型微调的情况下,系统性提升生成式AI的可信性、可控性与可复现性。相关代码、缓存引导向量及评估脚本已公开发布,以支持可复现性研究。

英文摘要

Agentic large language models (LLMs) deployed in fact-sensitive applications such as customer support must simultaneously preserve factual correctness and generate responses in a controllable stylistic register. Activation steering enables fine-tuning-free style control by perturbing hidden representations, but it lacks an explicit mechanism for distinguishing verifiable facts from stylistic content, leading to semantic leakage. We address this challenge through \emph{Defactualize-Steer-Rehydrate} (DSR), a knowledge-engineering framework that integrates a typed, salience-weighted knowledge graph (KG) with activation steering. DSR extracts salient entities using a layered regex or NER or lexical-classifier pipeline, replaces them with typed placeholders prior to steering, and deterministically restores verified values through salience-guided rehydration after generation. DSR is evaluated across six LLaMA-family models (1B--13B parameters) on 600 A2A-generated customer-support cases (1,200 generations), with a dedicated KG ablation study. DSR significantly increases verified-entity recovery relative to a steering-only baseline (Cohen's $d=0.225$, $p_{\text{Bonf}}=1.0\times10^{-4}$), though the absolute recovery rate remains modest, while preserving effective style control across diverse model families. Layer-wise separability and steering-strength diagnostics further show previously unexplored interactions between representation-level steering and factual grounding. hese results demonstrate that explicit knowledge engineering can systematically enhance trustworthy, controllable, and reproducible generative AI without requiring model fine-tuning. Code, cached steering vectors, and evaluation scripts are publicly released to support reproducibility.\footnote{https://github.com/Tanmay-IITDSAI/KG-Gated-Defactualization}

发表机构

  • Indian Institute of Technology (IIT) Bhilai(印度比莱印度理工学院)

机构由 AI 辅助整理,请以论文原文为准。

↑