发表机构
Fudan University; National University of Singapore; University of Chinese Academy of Sciences(复旦大学; 新加坡国立大学; 中国科学院大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本文提出 Chinese-Jev,一种通过统一数据处理和训练流程提升中文决策准确性的 System One 模型,在通用和医学等领域超越 Jev,并实现显著加速和低延迟。
AI 中文摘要
System One 模型(如 Jev)为需要决策而非开放式回答的任务提供了一种生成式语言模型的高效替代方案。然而,现有的 Jev 模型在中文决策准确性方面表现有限,限制了其在通用和专业场景中的实用性。在本文中,我们介绍了 Chinese-Jev,一种通过统一数据处理和训练流程解决这一差距的 System One 模型。我们的数据处理协议将异构的中文标注转换为候选选项上的概率目标,从而在跨领域和问题格式之间实现共享的训练公式。为了实现高效推理,Chinese-Jev 采用轻量级的仅编码器骨干网络进行文本编码,并通过面向决策的训练学习对候选答案进行评分。为了解决预训练分布与下游中文场景之间的错配问题,我们首先在包含 1000 万个示例的通用语料库上训练模型,然后分别针对医学、法律和金融领域进行微调。为了评估通用和特定领域中文环境中的决策准确性和校准,我们引入了 Chinese-Jev Bench(CJ-Bench)。在第一阶段预训练之后,Chinese-Jev 在通用领域任务上的准确率超过了闭源的 Jev 模型 1.24%,同时实现了 20.3 倍的加速。随后的特定领域微调在医学领域相比 Jev 实现了 4.0% 的准确率提升,并在专业领域达到了 Jev 平均准确率的 92%,同时实现了 17 倍的加速和每示例平均仅 15 毫秒的延迟。我们进一步展示了在移动设备上部署 INT8 量化模型,实现每次决策约 1.0 秒的推理延迟。该项目可在以下 https URL 获取。
英文摘要
System One models such as Jev offer an efficient alternative to generative language models for tasks that require decisions rather than open-ended responses. However, existing Jev models exhibit limited Chinese-language decision accuracy, restricting their utility in both general and specialized settings. In this paper, we introduce Chinese-Jev, a System One model that addresses this gap through a unified data processing and training pipeline. Our data processing protocol converts heterogeneous Chinese-language annotations into probability targets over candidate options, enabling a shared training formulation across domains and question formats. To enable efficient inference, Chinese-Jev adopts a lightweight encoder-only backbone for text encoding and learns to score candidate answers through decision-oriented training. To address the misalignment between the pre-training distribution and downstream Chinese-language scenarios, we first train the model on a general-purpose corpus of 10 million examples, then fine-tune it separately for the medical, legal, and financial domains. To evaluate decision accuracy and calibration in both general and domain-specific Chinese-language settings, we introduce Chinese-Jev Bench (CJ-Bench). After first-stage pre-training, Chinese-Jev exceeds the accuracy of the closed-source Jev model by 1.24% on general-domain tasks while achieving a 20.3x speedup. Subsequent domain-specific fine-tuning yields a 4.0% accuracy improvement over Jev in medicine and achieves 92% of Jev's average accuracy across specialized domains, with a 17x speedup and an average latency of only 15 ms per example. We further demonstrate on-device deployment of an INT8-quantized model on mobile devices, achieving an inference latency of approximately 1.0 second per decision. The project is available at https://gulucaptain.github.io/Chinese-Jev/.
Comments10 pages, 6 figures