当机器说话:将机器原生符号整合到预训练大语言模型的统一生成框架
When Machines Speak: A Unified Generative Framework for Integrating Machine-Native Symbols into Pretrained Large Language Models
浏览论文内容
中文总结 AI 辅助
研究针对预训练大语言模型无法处理机器原生符号的问题,提出UniLang框架将机器原生符号与自然语言 token 同等建模,在序列推荐、法律先例预测任务上均优于基线,拓展了LLMs的应用范围。
中文摘要 AI 辅助
许多现实世界的AI系统使用离散的机器原生符号而非自然语言来表示实体、行为和结构化信息。这些表示形式紧凑且保留了任务相关的结构,但它们处于预训练大语言模型(LLMs)的语言 token 空间之外,在语言建模与结构化预测之间造成了根本性的鸿沟。我们提出了UniLang,一种统一生成框架,通过扩展预训练LLMs,使其将机器原生符号与自然语言 token 同等视为一级生成单元,从而弥合这一鸿沟。UniLang利用有根基的机器原生表示扩展了LLMs的词汇表和嵌入空间,使文本和符号 token 能够在单一自回归目标下被联合建模与生成。这种统一接口让预训练LLMs能够直接对机器原生表示进行操作,无需将其表述为自然语言或依赖特定任务架构。我们在两个结构截然不同的任务上评估UniLang:序列推荐和法律先例预测,涵盖不同领域和类型的结构化预测。在这两个任务中,UniLang始终优于强大的基线,为将预训练LLMs扩展到语言之外、将其用作异构机器原生表示的通用生成建模主干提供了可行路径。
英文摘要
Many real-world AI systems represent entities, behaviors, and structured information using discrete machine-native symbols rather than natural language. While these representations are compact and preserve task-relevant structure, they lie outside the linguistic token space of pretrained large language models (LLMs), creating a fundamental divide between language modeling and structured prediction. We introduce UniLang, a unified generative framework that bridges this divide by extending pretrained LLMs to treat machine-native symbols as first-class generative units alongside natural-language tokens. UniLang expands the LLM's vocabulary and embedding space with grounded machine-native representations, enabling textual and symbolic tokens to be jointly modeled and generated under a single autoregressive objective. This unified interface allows pretrained LLMs to directly operate on machine-native representations without requiring them to be verbalized as natural language or relying on task-specific architectures. We evaluate UniLang on two structurally distinct tasks, sequential recommendation and legal precedent prediction, spanning different domains and types of structured prediction. Across both tasks, UniLang consistently outperforms strong baselines, demonstrating a path toward extending pretrained LLMs beyond language and using them as a common generative modeling backbone for heterogeneous machine-native representations.