arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

HSS-Synth:面向大语言模型的人文社科数据合成

HSS-Synth: Humanities and Social Sciences Data Synthesis for LLMs

Ru Peng, Tianyu Zhao, Xijun Gu, Zhiting Fan, Haokai Xu, Jinyang Zhang, Yawen Zeng, Yihong Zhuang, Kexin Yang, Junyang Lin, Dayiheng Liu, Junbo Zhao

arXiv 2607.27379首次发表:更新:

发表机构

Zhejiang University; Inclusion AI, Ant Group; Peking University; Qwen Team, Alibaba Group(浙江大学; 蚂蚁集团Inclusion AI; 北京大学; 阿里巴巴集团通义千问团队)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本文针对LLM的HSS数据稀缺问题,提出以学科为中心的HSS-Synth流水线,生成23.7万指令微调样本,使Qwen3-8B-Base达SOTA并提升相关能力。

AI 中文摘要

高质量、多样化的数据对大语言模型(LLM)至关重要,但仍稀缺且成本高昂。数据合成是可行的替代方案,在封闭任务中表现出色,但人文社科(HSS)领域被忽视,其开放性特质使合成工作颇具挑战。不同于以往以能力为核心、零散的尝试,本文采用以学科为中心的范式,定义了首个涵盖14个主流领域的HSS领域体系,并推出首个面向HSS的数据合成流水线HSS-Synth。HSS-Synth包含三部分:(1)通过多步骤过滤和由评判者评估的文本优化,从网络语料库构建种子文档;(2)指定“需求+人设”,将种子文档回译为多样化但忠实的指令,并进行严格的问答对齐检查;(3)通过教师强制回答突破LLM响应限制,在响应生成过程中输入种子文档以锚定语义、减少幻觉、保留语气与完整性。HSS-Synth生成了23.7万高质量、多样化的指令微调样本,在16个基准测试中优于14个领先基线。微调后的Qwen3-8B-Base达到新的SOTA,接近官方Qwen3-8B,在无性能波动的情况下同时提升了人类偏好和知识能力。大量实验验证了HSS-Synth的鲁棒性与可迁移性,其代码公开于此https URL。

英文摘要

High-quality, diverse data are vital for large language models (LLMs) but remain scarce and costly. Data synthesis is a viable alternative and succeeds on closed tasks, yet the humanities and social sciences (HSS) are overlooked, and their open-ended nature makes synthesis challenging. Moving beyond prior capability-centric, fragmented attempts, we adopt a subject-centric paradigm, define the first HSS domain system covering 14 mainstream fields, and introduce HSS-Synth, the first data synthesis pipeline for HSS. HSS-Synth comprises: (1) constructing seed documents from web corpora via multi-step filtering and text refinement evaluated by a judge; (2) specifying "requirements + persona" to backtranslate seed documents into diverse yet faithful instructions with a strict Q&A alignment check; and (3) breaking LLM response limits via teacher-forced Answering that feeds seed documents during response generation to anchor semantics, reduce hallucinations, and preserve tone and integrity. HSS-Synth yields 237k high-quality, diverse instruction-tuning samples that outperform 14 leading baselines on 16 benchmarks. The fine-tuned Qwen3-8B-Base sets a new SOTA and approaches the official Qwen3-8B, improving both human preference and knowledge capabilities without performance seesaws. Extensive experiments demonstrate HSS-Synth's robustness and transferability. Our code is publicly available at https://github.com/pengr/HSS-Synth.

CommentsACL Findings 2026 Paper

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑