发表机构
CNRS/LORIA; Université de Lorraine(法国国家科学研究中心/洛林计算机科学与应用实验室; 洛林大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本研究提出数据驱动知识蒸馏(DDKD)及保结构增强方法,在无域内训练数据的跨域多任务数据到文本生成任务中,17亿参数的DDKD性能优于零样本推理与微调,且构建了QUINTD-5数据集。
AI 中文摘要
结构化数据以多种形式存在(表格、知识图谱、图表、时间序列),将其转换为文本可能涉及不同的生成任务。然而,大多数现有的数据到文本(D2T)生成研究聚焦于特定任务和数据集,要么依赖特定任务的训练数据,要么依赖大语言模型的零样本能力。本研究在既无域内训练文本也无测试参考、且域、生成目标和输入结构差异显著的场景下,研究跨域D2T生成。我们对比了数据驱动知识蒸馏(DDKD)与零样本推理、域外D2T数据微调的效果,并引入了通过结构子采样和扰动实现的保结构增强。在五个基准上的实验表明,在模型规模固定为17亿参数的情况下,DDKD的性能始终优于微调与零样本推理;此外,得到的小型模型在五个域中的两个上优于大得多的微调模型,在其余三个域上取得了可比的性能。我们进一步构建了QUINTD-5(QUINTD-1的五倍扩展),并表明仅缩放真实目标域输入仅能带来有限提升,而我们的增强策略对于跨域蒸馏而言更有效且成本效益更高。
英文摘要
Structured data exists in many forms (tables, knowledge graphs, charts, and time series), and converting it into text may involve different generation tasks. However, most prior work on data-to-text (D2T) generation has focused on specific tasks and datasets, relying either on task-specific training data or on the zero-shot capabilities of large language models. We study cross-domain D2T generation in a setting where neither in-domain training text nor test references are available, and where domains, generation goals, and input structures vary substantially. We compare data-driven knowledge distillation (DDKD) against zero-shot inference and fine-tuning on out-of-domain D2T data, and introduce structure-preserving augmentation via structural subsampling and perturbation. Experiments on five benchmarks show that, at constant model size (1.7B parameters), DDKD consistently outperforms both fine-tuning and zero-shot inference. Moreover, the resulting small models outperform a much larger finetuned model on two of the five domains, achieving comparable performance on the remaining three. We further construct QUINTD-5, a fivefold extension of QUINTD-1, and show that simply scaling real target-domain inputs yields only modest gains, whereas our augmentation strategy remains more effective and more cost-efficient for cross-domain distillation.
CommentsAccepted by EMNLP Findings 2026