arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

BirdsongChat:一种用于多模态具身行为模拟的混合多智能体框架

BirdsongChat: A Hybrid Multi-Agent Framework for Multimodal Embodied Behavior Simulation

Callie C. Liao, Duoduo Liao, Ellie L. Zhang

arXiv 2609.20887首次发表:更新:

发表机构

Stanford University; George Mason University; IntelliSky(斯坦福大学; 乔治梅森大学; IntelliSky)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

提出混合多智能体框架BirdsongChat,通过统一参数表示桥接语义推理与物理执行,实现多模态具身行为模拟,在鸟类交互场景中达到高一致性。

AI 中文摘要

多模态具身系统需要将人类意图转化为跨异构模态的可解释且协调的行为。然而,现有的多模态智能体通常依赖隐式表示,限制了可控性和跨模态一致性。我们提出了一种用于交互式多模态行为模拟的混合多智能体框架,该框架通过统一参数表示(UPR)桥接语义推理与物理执行。基于LLM的推理智能体将多模态输入转化为UPR,UPR编码行为状态和可解释的控制参数,供模拟智能体生成同步的3D运动、空间化声景和环境行为。我们开发了BirdsongChat作为所提框架的原型实现,以交互式鸟类行为模拟作为测试平台,紧密耦合运动、发声和环境上下文。BirdsongChat在涉及物种、行为、情感状态、环境和多鸟交互的文本和图像引导场景中进行了评估。该系统在跨模态一致性上获得94.4%的归一化分数,情感一致性达到100%,生成一致性为92.6%。这些结果表明,显式的中间表示有效地桥接了语义推理与物理执行,提高了可控性和多模态同步性。因此,所提框架为需要跨模态可解释的语义到物理协调的具身AI系统提供了一种可泛化的设计原则,在仿生生态声学、群体机器人、虚拟环境和创意多媒体中具有潜在应用。

英文摘要

Multimodal embodied systems require translating human intentions into interpretable and coordinated behaviors across heterogeneous modalities. However, existing multimodal agents often rely on implicit representations, limiting controllability and cross-modal consistency. We present a hybrid multi-agent framework for interactive multimodal behavior simulation that bridges semantic reasoning and physical execution through a Unified Parameter Representation (UPR). LLM-based reasoning agents transform multimodal inputs into UPR, which encodes behavioral states and interpretable control parameters for simulation agents generating synchronized 3D motion, spatialized soundscapes, and environmental behaviors. We develop BirdsongChat as a prototype implementation of the proposed framework, using interactive avian behavior simulation as a testbed that tightly couples motion, vocalization, and environmental context. BirdsongChat is evaluated on text- and image-guided scenarios involving species, behaviors, affective states, environments, and multi-bird interactions. The system achieves normalized scores of 94.4\% for cross-modal coherence, 100% for affective consistency, and 92.6% for generation consistency. These results demonstrate that an explicit intermediate representation effectively bridges semantic reasoning and physical execution, improving controllability and multimodal synchronization. The proposed framework thus offers a generalizable design principle for embodied AI systems requiring interpretable semantic-to-physical coordination across modalities, with potential applications in bio-inspired ecoacoustics, swarm robotics, virtual environments, and creative multimedia.

CommentsPaper contents accepted by EMNLP 2026 REALM

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑