发表机构
Adobe Media and Data Science Research (MDSR); IIIT-Delhi; IIT Kanpur; SUNY at Buffalo(Adobe媒体与数据科学研究(MDSR); 德里印度信息技术学院; 坎普尔印度理工学院; 纽约州立大学布法罗分校)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本研究针对LLM与非语言智能体协作的 verbalization 瓶颈,提出潜在状态内化方法,构建LLAMIA-Bench基准,发现该方法优于 verbalization 集成,140亿参数的LLAMIA模型性能超越相关模型且分布外泛化能力强。
AI 中文摘要
大型语言模型(LLMs)正越来越多地被部署为协调器,通过自然语言协调专门的子智能体解决复杂任务。然而,在游戏、机器人等诸多重要领域,性能最强的可用智能体并非语言模型。将非语言智能体与LLMs集成需要「 verbalization( verbalization )」:在每一个交互步骤中,将非语言智能体丰富的连续表示压缩为稀疏的文本摘要。为研究 verbalization 是否构成协作瓶颈,我们推出了 \textsc{LLAMIA-Bench},这是一套包含六个不同协作国际象棋任务的基准测试,涵盖三个方面:行为模仿、状态评估和自然语言解释。每个任务都实例化了一个成熟的国际象棋问题,该问题无法由LLM或国际象棋引擎单独解决。为解决LLM与非语言智能体的协作问题,我们提出了「 latent state internalization(潜在状态内化)」,该方法将子智能体的连续表示直接投影到LLM的 token 流中,作为学习到的状态 token,随着动作推进环境状态而进行动态重新编码。将潜在状态内化与 verbalization 集成进行比较,我们的实验揭示了持续存在的「 verbalization debt( verbalization 债务)」:性能差距在整个训练过程中不断扩大,并且在LLM从40亿参数扩展到140亿参数时仍然存在。一个使用潜在状态内化训练的140亿参数模型 \textsc{LLAMIA},在所有基准任务上的性能与任务专家和前沿模型(包括具有工具访问权限的GPT-5.1)相当或更优,并且在分布外泛化方面表现出色,而特定任务的微调模型则在此方面出现性能崩溃。
英文摘要
LLMs are increasingly deployed as orchestrators that coordinate specialized subagents to solve complex tasks through natural language. However, in many important domains like game playing and robotics, the strongest available agents are not language models. Integrating non-language agents with LLMs would require \emph{verbalization}: compressing their rich continuous representations into sparse textual summaries at each interaction step. To study whether verbalization constitutes a bottleneck, we introduce \textsc{LLAMIA-Bench}, a suite of six diverse collaborative chess tasks spanning three facets: behavioral imitation, state assessment, and natural-language explanation. Each task instantiates a well-established chess problem that neither the LLM nor the chess engine can solve alone. To solve LLM collaboration with non-language agents, we introduce \emph{latent state internalization}, which projects the subagent's continuous representations directly into the LLM's token stream as learned state tokens, with dynamic re-encoding as actions advance the environment state. Comparing internalization to verbalized integration, our experiments reveal a consistent \emph{verbalization debt}: the performance gap widens throughout training and persists as the LLM scales from 4B to 14B parameters. A single 14B model, \textsc{LLAMIA}, trained with latent state internalization, matches or exceeds task specialists and frontier models including GPT-5.1 with tool access across all benchmark tasks, and generalizes out-of-distribution where task-specific finetunes collapse
CommentsAccepted at EMNLP 2026