arXivDaily arXiv每日学术速递 周一至周五更新

大厂专区

Moonshot AI(月之暗面)

至 收录 18
2601.06521 2026-07-08 cs.CV cs.CL 版本更新

BabyVision: Visual Reasoning Beyond Language

BabyVision:超越语言的视觉推理

Liang Chen, Weichu Xie, Yiyan Liang, Hongfeng He, Hans Zhao, Zhibo Yang, Zhiqi Huang, Haoning Wu, Haoyu Lu, Y. charles, Yiping Bao, Yuantao Fan, Guopeng Li, Haiyang Shen, Xuanzhong Chen, Wendong Xu, Shuzheng Si, Zefan Cai, Wenhao Chai, Ziqi Huang, Fangfu Liu, Tianyu Liu, Baobao Chang, Ming Wu, Xiaobo Hu, Kaiyuan Chen, Yixin Ren, Yang Liu, Yuan Gong, Kuan Li

机构 * UniPat AI xbench Alibaba Group(阿里巴巴集团) MoonShot AI StepFun Peking University(北京大学) Tsinghua University(清华大学) University of Wisconsin–Madison(威斯康星大学麦迪逊分校) Princeton University(普林斯顿大学) Nanyang Technological University(南洋理工大学) G Labs Equal Core Contributors(0G Labs 等核心贡献者)

AI总结 研究发现当代多模态语言模型依赖语言先验,在基本视觉任务上表现差。为此引入BabyVision基准评估其核心视觉能力,涵盖多任务。结果显示模型缺乏基本视觉原语,BabyVision的进展迈向人类水平视觉能力,还探索了用生成模型解决视觉推理的方法。

Comments 26 pages, Homepage at https://unipat.ai/blog/BabyVision

详情
AI中文摘要

虽然人类在获得语言之前很久就发展出了核心视觉技能,但当代多模态语言模型(MLLMs)仍然严重依赖语言先验来弥补其脆弱的视觉理解能力。我们发现了一个关键事实:最先进的MLLMs在人类(甚至3岁儿童)能轻松解决的基本视觉任务上持续失败。为系统研究这一差距,我们引入了BabyVision,一个旨在评估MLLMs独立于语言知识的核心视觉能力的基准。BabyVision涵盖广泛任务,有388个项目分为四个关键类别的22个子类。实证结果和人工评估表明,领先的MLLMs表现显著低于人类基线。Gemini3 - Pro - Preview得分为49.7,落后于6岁儿童,远低于成人平均得分94.1。这些结果表明,尽管在知识密集型评估中表现出色,但当前MLLMs仍缺乏基本视觉原语。BabyVision的进展代表了向人类水平视觉感知和推理能力迈出的一步。我们还通过提出BabyVision - Gen和自动评估工具包探索用生成模型解决视觉推理问题。我们的代码和基准数据在该https网址发布以供复制。

英文摘要

While humans develop core visual skills long before acquiring language, contemporary Multimodal LLMs (MLLMs) still rely heavily on linguistic priors to compensate for their fragile visual understanding. We uncovered a crucial fact: state-of-the-art MLLMs consistently fail on basic visual tasks that humans, even 3-year-olds, can solve effortlessly. To systematically investigate this gap, we introduce BabyVision, a benchmark designed to assess core visual abilities independent of linguistic knowledge for MLLMs. BabyVision spans a wide range of tasks, with 388 items divided into 22 subclasses across four key categories. Empirical results and human evaluation reveal that leading MLLMs perform significantly below human baselines. Gemini3-Pro-Preview scores 49.7, lagging behind 6-year-old humans and falling well behind the average adult score of 94.1. These results show despite excelling in knowledge-heavy evaluations, current MLLMs still lack fundamental visual primitives. Progress in BabyVision represents a step toward human-level visual perception and reasoning capabilities. We also explore solving visual reasoning with generation models by proposing BabyVision-Gen and automatic evaluation toolkit. Our code and benchmark data are released at https://github.com/UniPat-AI/BabyVision for reproduction.

URL PDF HTML 收藏
2605.19811 2026-06-23 cs.LG 版本更新

LionMuon: Alternating Spectral and Sign Descent for Efficient Training

LionMuon: 交替频谱和符号下降用于高效训练

Arman Bolatov, Artem Riabinin, Nikita Kornilov, Andrey Veprikov, Samuel Horváth, Martin Takáč, Aleksandr Beznosikov

机构 * DeepSeek-AI Essential AI Kimi Team

AI总结 本文提出LionMuon优化器,通过交替使用Lion和Muon的更新步骤,在保持有效性的同时显著降低平均迭代成本,同时证明了在重尾噪声下的复杂性界限,展示了其在不同模型规模下的优势。

Comments 38 pages, 13 figures, 4 tables

详情
AI中文摘要

在大规模优化中,更新步骤的廉价性和有效性是成功优化器的关键因素。基于符号的优化器如Lion或Signum产生廉价的每步更新,而Muon的谱矩阵-符号更新则在显著更高的每步成本下提供更强的方向。在本文中,我们提出LionMuon,它保留了Muon步骤的有效性,同时显著降低了平均迭代成本,类似于基于符号的方法。它在固定周期P内交替使用Lion和Muon的更新,共享一个单一的双EMA动量缓冲区。因此,优化器状态内存与Lion相同,恰好是AdamW的一半。一个更简单的单EMA变体SignMuon本身已经优于纯Muon。在P=2时,LionMuon在我们测试的124M模型大小的每个数据集和架构上都优于Muon、Lion、Signum和AdamW,在更低的计算下达到更低的验证损失,这一优势在355M和720M规模上仍然存在。在理论方面,我们证明了在重尾噪声下的严格复杂性界限,这些界限由周期平均平滑度和介于Muon和Lion之间的噪声所决定。这些界限预测了计算最优的周期以及LionMuon超越Muon和Lion的条件。代码:https://github.com/brain-lab-research/lion-muon

英文摘要

In large-scale optimization, the cheapness and effectiveness of update steps are the most crucial factors for a successful optimizer. Sign-based optimizers like Lion or Signum produce cheap per-step updates, whereas Muon's spectral matrix-sign update gives a much stronger direction at a substantially higher per-step cost. In this work, we propose LionMuon, which retains the effectiveness of Muon steps while considerably cutting the averaged iteration cost, similar to sign-based methods. It alternates between Lion's and Muon's updates on a fixed period P, sharing a single dual-EMA momentum buffer between them. The optimizer state memory therefore matches Lion and is exactly half of AdamW's. A simpler single-EMA variant, SignMuon, by itself already outperforms pure Muon. At P = 2, LionMuon Pareto-dominates Muon, Lion, Signum, and AdamW on every dataset and architecture we tested at 124M model size, reaching lower validation loss at lower compute, and the same advantage persists at 355M and 720M scale. On the theory side, we prove sharp complexity bounds under heavy-tailed noise which are governed by period-averaged smoothness and noise that interpolate between Muon's and Lion's constants. These bounds predict the compute-optimal period and the conditions under which LionMuon outruns Muon and Lion. Code: https://github.com/brain-lab-research/lion-muon

URL PDF HTML 收藏
2604.16804 2026-04-21 cs.LG cs.AI

AutoOR: Scalably Post-training LLMs to Autoformalize Operations Research Problems

AutoOR: 一种可扩展的后训练语言模型用于自动形式化运筹学问题

Sumeet Ramesh Motwani, Chuan Du, Aleksander Petrov, Christopher Davis, Philip Torr, Antonio Papania-Davis, Weishi Yan

机构 * The Moonshot Factory(谷歌登月工厂) University of Oxford(牛津大学)

AI总结 AutoOR通过合成数据生成与强化学习管道,训练语言模型自动形式化运筹学问题,实现线性、混合整数和非线性优化问题的高效解决,取得六个经典运筹学基准测试的领先或竞争性结果。

详情
AI中文摘要

优化问题在制造、物流、调度等工业领域决策中至关重要。将复杂描述转化为求解器准备的公式需要专门的运筹学(OR)专业知识,难以扩展。我们提出了AutoOR,一种可扩展的合成数据生成和强化学习管道,训练语言模型在自然语言中自动形式化跨线性、混合整数和非线性类别的优化问题。AutoOR从标准优化形式生成验证训练数据,并使用求解器执行反馈作为RL后训练的奖励信号。在8B模型上应用AutoOR,在六个经典OR基准测试中取得最先进的或竞争性结果,显著匹配更大的前沿模型。对于涉及物理动态的非线性问题类别,前沿模型得分接近0%,我们引入了课程RL策略,从有限的初始训练数据出发,使该类别在后训练中可处理。我们相信,此类方法如AutoOR可以显著加速工业决策中的AI应用。

英文摘要

Optimization problems are central to decision-making in manufacturing, logistics, scheduling, and other industrial settings. Translating complicated descriptions of these problems into solver-ready formulations requires specialized operations research (OR) expertise, making it hard to scale. We present AutoOR, a scalable synthetic data generation and reinforcement learning pipeline that trains LLMs to autoformalize optimization problems specified in natural language across linear, mixed-integer, and non-linear categories. AutoOR generates verified training data from standard optimization forms and uses solver execution feedback as the reward signal for RL post-training. AutoOR applied to an 8B model achieves state-of-the-art or competitive results across six established OR benchmarks, matching significantly larger frontier models. For a non-linear problem class involving physical dynamics, where frontier models score near 0%, we introduce a curriculum RL strategy that bootstraps from limited initial training data to make this class tractable for post-training. We believe that methods such as AutoOR can significantly accelerate industrial decision-making with AI.

URL PDF HTML 收藏
2604.02349 2026-04-06 cs.LG cs.AI

OPRIDE: Offline Preference-based Reinforcement Learning via In-Dataset Exploration

OPRIDE: 通过数据集内探索实现离线偏好驱动强化学习

Yiqin Yang, Hao Hu, Yihuan Mao, Jin Zhang, Chengjie Wu, Yuhua Jiang, Xu Yang, Runpeng Xie, Yi Fan, Bo Liu, Yang Gao, Bo Xu, Chongjie Zhang

机构 * The Key Laboratory of Cognition and Decision Intelligence for Complex Systems, Institute of Automation, Chinese Academy of Sciences(中国科学院自动化研究所复杂系统认知与决策智能重点实验室) Tsinghua University(清华大学) Moonshot AI(月之暗面) Amazon(亚马逊) University of Arizona(亚利桑那大学) Washington University in St. Louis(圣路易斯华盛顿大学)

AI总结 本文提出OPRIDE算法,通过改进探索策略和折扣调度机制提升离线偏好驱动强化学习的查询效率,实验表明其在多种任务中表现优异。

Journal ref ICLR-2026

详情
AI中文摘要

偏好驱动强化学习(PbRL)能够避免复杂的奖励设计并更好地对齐人类意图,在各种现实应用中展现出巨大潜力。然而,获取人类偏好反馈成本高且耗时,成为PbRL发展的主要障碍。本文针对离线PbRL查询效率低的问题,指出两个主要原因:探索效率低下和学习奖励函数的过度优化。为此,我们提出一种新的算法OPRIDE,旨在提高离线PbRL的查询效率。OPRIDE包含两个关键特性:一种原理性的探索策略,最大化查询信息量,以及一种折扣调度机制,旨在缓解学习奖励函数的过度优化。通过实证评估,我们证明OPRIDE显著优于先前方法,在较少查询的情况下实现了强劲表现。此外,我们提供了算法效率的理论保证。在各种运动、操作和导航任务中的实验结果突显了我们方法的有效性和通用性。

英文摘要

Preference-based reinforcement learning (PbRL) can help avoid sophisticated reward designs and align better with human intentions, showing great promise in various real-world applications. However, obtaining human feedback for preferences can be expensive and time-consuming, which forms a strong barrier for PbRL. In this work, we address the problem of low query efficiency in offline PbRL, pinpointing two primary reasons: inefficient exploration and overoptimization of learned reward functions. In response to these challenges, we propose a novel algorithm, \textbf{O}ffline \textbf{P}b\textbf{R}L via \textbf{I}n-\textbf{D}ataset \textbf{E}xploration (OPRIDE), designed to enhance the query efficiency of offline PbRL. OPRIDE consists of two key features: a principled exploration strategy that maximizes the informativeness of the queries and a discount scheduling mechanism aimed at mitigating overoptimization of the learned reward functions. Through empirical evaluations, we demonstrate that OPRIDE significantly outperforms prior methods, achieving strong performance with notably fewer queries. Moreover, we provide theoretical guarantees of the algorithm's efficiency. Experimental results across various locomotion, manipulation, and navigation tasks underscore the efficacy and versatility of our approach.

URL PDF HTML 收藏
2511.14617 2026-04-06 cs.DC cs.LG

Seer: Online Context Learning for Fast Synchronous LLM Reinforcement Learning

Seer:在线上下文学习用于快速同步LLM强化学习

Ruoyu Qin, Weiran He, Weixiao Huang, Yangkun Zhang, Yikai Zhao, Bo Pang, Xinran Xu, Yingdi Shan, Yongwei Wu, Mingxing Zhang

机构 * Moonshot AI(月之暗面) Tsinghua University(清华大学)

AI总结 Seer通过动态负载均衡、上下文感知调度和自适应分组推测解码技术,显著降低长尾延迟并提升资源利用率,实现2.04倍的端到端效率提升。

详情
AI中文摘要

强化学习(RL)已成为推动现代大语言模型(LLMs)发展的重要技术,但现有同步RL系统面临严重性能瓶颈。rollout阶段占端到端迭代时间的主导地位,由于固有的工作负载不平衡,导致显著的长尾延迟和资源利用率低下。我们提出了Seer,一种新的上下文学习RL系统,通过关键观察:共享相同提示的请求在输出长度和响应模式上具有强相似性,来解决这些问题。利用这一洞察,Seer引入了三种协调技术:(1)分 rollout 实现动态负载均衡,(2)上下文感知调度以缓解长尾请求延迟,(3)自适应分组推测解码以加速生成。这些机制协同工作,显著减少rollout期间的长尾延迟并提高资源效率。在生产级RL工作负载上的评估表明,Seer相比现有同步RL系统实现了高达2.04倍的端到端rollout吞吐量提升,同时显著降低了长尾延迟,减少72-94%。

英文摘要

Reinforcement Learning (RL) has emerged as a critical technique for advancing modern Large Language Models (LLMs), yet existing synchronous RL systems face severe performance bottlenecks. The rollout phase, which dominates end-to-end iteration time, suffers from substantial long-tail latency and poor resource utilization due to inherent workload imbalance. We present Seer, a novel context learning RL system that addresses these challenges through a key observation: requests sharing the same prompt exhibit strong similarities in output lengths and response patterns. Leveraging this insight, Seer introduces three coordinated techniques: (1) divided rollout for dynamic load balancing, (2) context-aware scheduling to mitigate long-tail request delays, and (3) adaptive grouped speculative decoding to accelerate generation. These mechanisms work in concert to markedly reduce long-tail latency and improve resource efficiency during rollout. Evaluations on production-grade RL workloads demonstrate that Seer achieves up to 2.04$\times$ end-to-end rollout throughput improvement compared to the state-of-the-art synchronous RL systems, while notably reducing long-tail latency by 72-94%.

URL PDF HTML 收藏
2505.23359 2026-03-18 cs.CV

VideoReasonBench: Can MLLMs Perform Vision-Centric Complex Video Reasoning?

VideoReasonBench: 大规模语言模型能否进行以视觉为中心的复杂视频推理?

Yuanxin Liu, Kun Ouyang, Haoning Wu, Yi Liu, Lin Sui, Xinhao Li, Yan Zhong, Y. Charles, Xinyu Zhou, Xu Sun

机构 * State Key Laboratory of Multimedia Information Processing, School of Computer Science, Peking University(多媒体信息处理国家重点实验室,计算机科学学院,北京大学) Moonshot AI Nanjing University(南京大学) School of Mathematical Sciences, Peking University(数学科学学院,北京大学)

AI总结 VideoReasonBench评估了基于视觉的复杂视频推理能力,发现大多数先进多模态语言模型在复杂视频推理任务上表现不佳,而增强推理能力的Gemini-2.5-Pro表现突出。

Comments Project Page: https://llyx97.github.io/video_reason_bench/

详情
AI中文摘要

近期研究表明,长链推理(CoT)能显著提升大规模语言模型(LLMs)在复杂任务上的性能。然而,这一优势在视频理解领域尚未得到验证,因为现有基准缺乏展示扩展CoT链优势的推理深度。尽管近期有研究提出视频推理基准,但这些任务通常依赖知识而非视觉内容。为弥合这一差距,我们引入VideoReasonBench,该基准旨在评估以视觉为中心的复杂视频推理能力。为确保视觉丰富性和高推理复杂性,VideoReasonBench中的每个视频均展示一系列细粒度操作在潜状态上的序列,该状态仅在视频部分可见。问题评估三个递增级别的视频推理技能:回忆观察到的视觉信息、推断潜状态内容以及预测视频外的信息。在这样的任务设置下,模型必须精确回忆视频中的多个操作,并逐步推理以获得正确最终答案。使用VideoReasonBench,我们全面评估了18种最先进的多模态语言模型(MLLMs),发现大多数在复杂视频推理任务上表现不佳——例如,GPT-4o仅达到6.9%的准确率,而增强推理能力的Gemini-2.5-Pro显著优于其他模型,准确率达56.0%。我们对“测试时扩展”进行的调查进一步揭示,扩展的推理预算在现有视频基准上提供无或极小的收益,但在VideoReasonBench上是提升性能的关键。

英文摘要

Recent studies have shown that long chain-of-thought (CoT) reasoning can significantly enhance the performance of large language models (LLMs) on complex tasks. However, this benefit is yet to be demonstrated in the domain of video understanding, since most existing benchmarks lack the reasoning depth required to demonstrate the advantages of extended CoT chains. While recent efforts have proposed benchmarks aimed at video reasoning, the tasks are often knowledge-driven and do not rely heavily on visual content. To bridge this gap, we introduce VideoReasonBench, a benchmark designed to evaluate vision-centric, complex video reasoning. To ensure visual richness and high reasoning complexity, each video in VideoReasonBench depicts a sequence of fine-grained operations on a latent state that is only visible in part of the video. The questions evaluate three escalating levels of video reasoning skills: recalling observed visual information, inferring the content of latent states, and predicting information beyond the video. Under such task setting, models have to precisely recall multiple operations in the video, and perform step-by-step reasoning to get correct final answers for these questions. Using VideoReasonBench, we comprehensively evaluate 18 state-of-the-art multimodal LLMs (MLLMs), finding that most perform poorly on complex video reasoning -- e.g., GPT-4o achieves only 6.9% accuracy -- while the thinking-enhanced Gemini-2.5-Pro significantly outperforms others with 56.0% accuracy. Our investigations on "test-time scaling" further reveal that extended thinking budget, while offering none or minimal benefits on existing video benchmarks, is essential for improving the performance on VideoReasonBench.

URL PDF HTML 收藏
2506.16395 2026-03-03 cs.CL

OJBench: A Competition Level Code Benchmark For Large Language Models

OJBench: 一个面向大语言模型的竞赛级代码基准

Zhexu Wang, Yiping Liu, Yejie Wang, Wenyang He, Bofei Gao, Muxi Diao, Yanxu Chen, Kelin Fu, Flood Sung, Zhilin Yang, Tianyu Liu, Weiran Xu

机构 * Beijing University of Posts and Telecommunications(北京邮电大学) Tsinghua University(清华大学) University of Chinese Academy of Sciences(中国科学院大学) Peking University(北京大学) Moonshot AI

AI总结 OJBench是一个用于评估大语言模型竞赛级代码推理能力的基准,通过232个编程竞赛问题揭示了现有模型在复杂推理任务中的局限性。

Comments 9 pages, 5 figures

详情
AI中文摘要

近年来,大语言模型(LLMs)在数学和代码推理能力方面取得了显著进展。然而,现有代码基准在评估这些能力的全谱方面存在局限,尤其是在竞赛级别。为弥合这一差距,我们引入了OJBench,一个新型且具有挑战性的基准,旨在评估LLMs的竞赛级代码推理能力。OJBench包含来自NOI和ICPC的232个编程竞赛问题,为模型的推理能力提供了更严格的测试。我们使用OJBench对37个模型进行了全面评估,包括闭源和开源模型,以及以推理为导向和非推理为导向的模型。我们的结果表明,即使是最先进的以推理为导向的模型,如o4-mini和Gemini-2.5-pro-exp,也难以应对高度具有挑战性的竞赛级问题。这突显了模型在竞赛级代码推理中的重大挑战。

英文摘要

Recent advancements in large language models (LLMs) have demonstrated significant progress in math and code reasoning capabilities. However, existing code benchmark are limited in their ability to evaluate the full spectrum of these capabilities, particularly at the competitive level. To bridge this gap, we introduce OJBench, a novel and challenging benchmark designed to assess the competitive-level code reasoning abilities of LLMs. OJBench comprises 232 programming competition problems from NOI and ICPC, providing a more rigorous test of models' reasoning skills. We conducted a comprehensive evaluation using OJBench on 37 models, including both closed-source and open-source models, reasoning-oriented and non-reasoning-oriented models. Our results indicate that even state-of-the-art reasoning-oriented models, such as o4-mini and Gemini-2.5-pro-exp, struggle with highly challenging competition-level problems. This highlights the significant challenges that models face in competitive-level code reasoning.

URL PDF HTML 收藏
2601.16694 2026-02-25 cs.CV

Affinity Contrastive Learning for Skeleton-based Human Activity Understanding

基于骨架的人体活动理解的亲和对比学习

Hongda Liu, Yunfan Liu, Min Ren, Lin Sui, Yunlong Wang, Zhenan Sun

机构 * Institute of Automation, Chinese Academy of Sciences(中国科学院自动化研究所) School of Artificial Intelligence, University of Chinese Academy of Sciences(中国科学院大学人工智能学院) School of Electronic, Electrical and Communication Engineering, Univeristy of Chinese Academy of Sciences(中国科学院大学电子电气与通信工程学院) School of Artificial Intelligence, Beijing University of Posts and Telecommunications(北京邮电大学人工智能学院) Moonshot AI

AI总结 ACLNet通过引入亲和度度量和动态温度调度,提升基于骨架的人体活动理解的特征判别能力。

Comments Accepted by TBIOM

Journal ref IEEE Transactions on Biometrics, Behavior, and Identity Science (2026)

详情
AI中文摘要

在基于骨架的人体活动理解中,现有方法往往采用对比学习范式来构建判别性特征空间。然而,许多这些方法未能利用类间结构的相似性,并忽视了异常正样本的影响。在本研究中,我们引入了ACLNet,一种亲和对比学习网络,该网络探索人类活动类之间的复杂聚类关系以提高特征判别能力。具体来说,我们提出了一种亲和度度量来细化相似性测量,从而形成活动超类,提供更有信息量的对比信号。此外,我们还引入了动态温度调度,以自适应地调整各种超类的惩罚强度。此外,我们采用基于边界的对比策略来提高类内硬正样本和负样本的分离度。在NTU RGB+D 60、NTU RGB+D 120、Kinetics-Skeleton、PKU-MMD、FineGYM和CASIA-B等数据集上的大量实验证明了我们的方法在基于骨架的动作识别、步态识别和人重识别中的优越性。源代码可在https://github.com/firework8/ACLNet获取。

英文摘要

In skeleton-based human activity understanding, existing methods often adopt the contrastive learning paradigm to construct a discriminative feature space. However, many of these approaches fail to exploit the structural inter-class similarities and overlook the impact of anomalous positive samples. In this study, we introduce ACLNet, an Affinity Contrastive Learning Network that explores the intricate clustering relationships among human activity classes to improve feature discrimination. Specifically, we propose an affinity metric to refine similarity measurements, thereby forming activity superclasses that provide more informative contrastive signals. A dynamic temperature schedule is also introduced to adaptively adjust the penalty strength for various superclasses. In addition, we employ a margin-based contrastive strategy to improve the separation of hard positive and negative samples within classes. Extensive experiments on NTU RGB+D 60, NTU RGB+D 120, Kinetics-Skeleton, PKU-MMD, FineGYM, and CASIA-B demonstrate the superiority of our method in skeleton-based action recognition, gait recognition, and person re-identification. The source code is available at https://github.com/firework8/ACLNet.

URL PDF HTML 收藏
2602.15776 2026-02-18 cs.AI

GlobeDiff: State Diffusion Process for Partial Observability in Multi-Agent Systems

GlobeDiff:多智能体系统中部分可观测性的状态扩散过程

Yiqin Yang, Xu Yang, Yuhua Jiang, Ni Mu, Hao Hu, Runpeng Xie, Ziyou Zhang, Siyuan Li, Yuan-Hua Ni, Qianchuan Zhao, Bo Xu

机构 * The Key Laboratory of Cognition and Decision Intelligence for Complex Systems, Institute of Automation, Chinese Academy of Sciences(认知与决策智能复杂系统重点实验室,自动化研究所,中国科学院) Tsinghua University(清华大学) Moonshot AI Nankai University(南开大学) Faculty of Computing, Harbin Institute of Technology(哈尔滨工业大学计算机学院)

AI总结 GlobeDiff通过多模态扩散过程解决多智能体系统中部分可观测性问题,实现高保真度的全局状态推断。

Journal ref ICLR-2026

详情
AI中文摘要

在多智能体系统领域,部分可观测性问题是一个影响有效协调和决策的关键障碍。现有方法,如信念状态估计和智能体间通信,往往效果有限。基于信念的方法受限于其仅关注过去经验而无法充分利用全局信息,而通信方法通常缺乏有效利用辅助信息的稳健模型。为了解决这一问题,我们提出全局状态扩散算法(GlobeDiff),以局部观测推断全局状态。通过将状态推断过程建模为多模态扩散过程,GlobeDiff在克服状态估计模糊性的同时,能够以高保真度推断全局状态。我们证明了在单模态和多模态分布下,GlobeDiff的估计误差可以被界。大量实验结果表明,GlobeDiff在性能上优于现有方法,并能准确推断全局状态。

英文摘要

In the realm of multi-agent systems, the challenge of \emph{partial observability} is a critical barrier to effective coordination and decision-making. Existing approaches, such as belief state estimation and inter-agent communication, often fall short. Belief-based methods are limited by their focus on past experiences without fully leveraging global information, while communication methods often lack a robust model to effectively utilize the auxiliary information they provide. To solve this issue, we propose Global State Diffusion Algorithm~(GlobeDiff) to infer the global state based on the local observations. By formulating the state inference process as a multi-modal diffusion process, GlobeDiff overcomes ambiguities in state estimation while simultaneously inferring the global state with high fidelity. We prove that the estimation error of GlobeDiff under both unimodal and multi-modal distributions can be bounded. Extensive experimental results demonstrate that GlobeDiff achieves superior performance and is capable of accurately inferring the global state.

URL PDF HTML 收藏
2602.02537 2026-02-04 cs.CV cs.LG

WorldVQA: Measuring Atomic World Knowledge in Multimodal Large Language Models

WorldVQA:评估多模态大语言模型的原子视觉世界知识

Runjie Zhou, Youbo Shao, Haoyu Lu, Bowei Xing, Tongtong Bai, Yujie Chen, Jie Zhao, Lin Sui, Haotian Yao, Zijia Zhao, Hao Yang, Haoning Wu, Zaida Zhou, Jinguo Zhu, Zhiqi Huang, Yiping Bao, Yangyang Liu, Y. Charles, Xinyu Zhou

机构 * Moonshot AI

AI总结 WorldVQA通过评估多模态大语言模型的原子视觉知识,建立衡量其事实性和百科全书广度的基准测试。

详情
AI中文摘要

我们介绍了WorldVQA,一个旨在评估多模态大语言模型(MLLMs)原子视觉世界知识的基准测试。与当前评估不同,后者往往将视觉知识检索与推理混为一谈,WorldVQA将这些能力分离,严格测量模型‘记忆’的内容。该基准测试评估了在分层分类学中对视觉实体进行定位和命名的原子能力,从常见的头类对象到长尾稀有对象。我们期望WorldVQA能够作为视觉事实性的严格测试,从而建立评估当前及下一代前沿模型百科全书广度和幻觉率的标准。

英文摘要

We introduce WorldVQA, a benchmark designed to evaluate the atomic visual world knowledge of Multimodal Large Language Models (MLLMs). Unlike current evaluations, which often conflate visual knowledge retrieval with reasoning, WorldVQA decouples these capabilities to strictly measure "what the model memorizes." The benchmark assesses the atomic capability of grounding and naming visual entities across a stratified taxonomy, spanning from common head-class objects to long-tail rarities. We expect WorldVQA to serve as a rigorous test for visual factuality, thereby establishing a standard for assessing the encyclopedic breadth and hallucination rates of current and next-generation frontier models.

URL PDF HTML 收藏
2602.02276 2026-02-04 cs.CL cs.AI cs.LG

Kimi K2.5: Visual Agentic Intelligence

Kimi K2.5:视觉代理智能

Kimi Team, Tongtong Bai, Yifan Bai, Yiping Bao, S. H. Cai, Yuan Cao, Y. Charles, H. S. Che, Cheng Chen, Guanduo Chen, Huarong Chen, Jia Chen, Jiahao Chen, Jianlong Chen, Jun Chen, Kefan Chen, Liang Chen, Ruijue Chen, Xinhao Chen, Yanru Chen, Yanxu Chen, Yicun Chen, Yimin Chen, Yingjiang Chen, Yuankun Chen, Yujie Chen, Yutian Chen, Zhirong Chen, Ziwei Chen, Dazhi Cheng, Minghan Chu, Jialei Cui, Jiaqi Deng, Muxi Diao, Hao Ding, Mengfan Dong, Mengnan Dong, Yuxin Dong, Yuhao Dong, Angang Du, Chenzhuang Du, Dikang Du, Lingxiao Du, Yulun Du, Yu Fan, Shengjun Fang, Qiulin Feng, Yichen Feng, Garimugai Fu, Kelin Fu, Hongcheng Gao, Tong Gao, Yuyao Ge, Shangyi Geng, Chengyang Gong, Xiaochen Gong, Zhuoma Gongque, Qizheng Gu, Xinran Gu, Yicheng Gu, Longyu Guan, Yuanying Guo, Xiaoru Hao, Weiran He, Wenyang He, Yunjia He, Chao Hong, Hao Hu, Jiaxi Hu, Yangyang Hu, Zhenxing Hu, Ke Huang, Ruiyuan Huang, Weixiao Huang, Zhiqi Huang, Tao Jiang, Zhejun Jiang, Xinyi Jin, Yu Jing, Guokun Lai, Aidi Li, C. Li, Cheng Li, Fang Li, Guanghe Li, Guanyu Li, Haitao Li, Haoyang Li, Jia Li, Jingwei Li, Junxiong Li, Lincan Li, Mo Li, Weihong Li, Wentao Li, Xinhang Li, Xinhao Li, Yang Li, Yanhao Li, Yiwei Li, Yuxiao Li, Zhaowei Li, Zheming Li, Weilong Liao, Jiawei Lin, Xiaohan Lin, Zhishan Lin, Zichao Lin, Cheng Liu, Chenyu Liu, Hongzhang Liu, Liang Liu, Shaowei Liu, Shudong Liu, Shuran Liu, Tianwei Liu, Tianyu Liu, Weizhou Liu, Xiangyan Liu, Yangyang Liu, Yanming Liu, Yibo Liu, Yuanxin Liu, Yue Liu, Zhengying Liu, Zhongnuo Liu, Enzhe Lu, Haoyu Lu, Zhiyuan Lu, Junyu Luo, Tongxu Luo, Yashuo Luo, Long Ma, Yingwei Ma, Shaoguang Mao, Yuan Mei, Xin Men, Fanqing Meng, Zhiyong Meng, Yibo Miao, Minqing Ni, Kun Ouyang, Siyuan Pan, Bo Pang, Yuchao Qian, Ruoyu Qin, Zeyu Qin, Jiezhong Qiu, Bowen Qu, Zeyu Shang, Youbo Shao, Tianxiao Shen, Zhennan Shen, Juanfeng Shi, Lidong Shi, Shengyuan Shi, Feifan Song, Pengwei Song, Tianhui Song, Xiaoxi Song, Hongjin Su, Jianlin Su, Zhaochen Su, Lin Sui, Jinsong Sun, Junyao Sun, Tongyu Sun, Flood Sung, Yunpeng Tai, Chuning Tang, Heyi Tang, Xiaojuan Tang, Zhengyang Tang, Jiawen Tao, Shiyuan Teng, Chaoran Tian, Pengfei Tian, Ao Wang, Bowen Wang, Chensi Wang, Chuang Wang, Congcong Wang, Dingkun Wang, Dinglu Wang, Dongliang Wang, Feng Wang, Hailong Wang, Haiming Wang, Hengzhi Wang, Huaqing Wang, Hui Wang, Jiahao Wang, Jinhong Wang, Jiuzheng Wang, Kaixin Wang, Linian Wang, Qibin Wang, Shengjie Wang, Shuyi Wang, Si Wang, Wei Wang, Xiaochen Wang, Xinyuan Wang, Yao Wang, Yejie Wang, Yipu Wang, Yiqin Wang, Yucheng Wang, Yuzhi Wang, Zhaoji Wang, Zhaowei Wang, Zhengtao Wang, Zhexu Wang, Zihan Wang, Zizhe Wang, Chu Wei, Ming Wei, Chuan Wen, Zichen Wen, Chengjie Wu, Haoning Wu, Junyan Wu, Rucong Wu, Wenhao Wu, Yuefeng Wu, Yuhao Wu, Yuxin Wu, Zijian Wu, Chenjun Xiao, Jin Xie, Xiaotong Xie, Yuchong Xie, Yifei Xin, Bowei Xing, Boyu Xu, Jianfan Xu, Jing Xu, Jinjing Xu, L. H. Xu, Lin Xu, Suting Xu, Weixin Xu, Xinbo Xu, Xinran Xu, Yangchuan Xu, Yichang Xu, Yuemeng Xu, Zelai Xu, Ziyao Xu, Junjie Yan, Yuzi Yan, Guangyao Yang, Hao Yang, Junwei Yang, Kai Yang, Ningyuan Yang, Ruihan Yang, Xiaofei Yang, Xinlong Yang, Ying Yang, Yi Yang, Yi Yang, Zhen Yang, Zhilin Yang, Zonghan Yang, Haotian Yao, Dan Ye, Wenjie Ye, Zhuorui Ye, Bohong Yin, Chengzhen Yu, Longhui Yu, Tao Yu, Tianxiang Yu, Enming Yuan, Mengjie Yuan, Xiaokun Yuan, Yang Yue, Weihao Zeng, Dunyuan Zha, Haobing Zhan, Dehao Zhang, Hao Zhang, Jin Zhang, Puqi Zhang, Qiao Zhang, Rui Zhang, Xiaobin Zhang, Y. Zhang, Yadong Zhang, Yangkun Zhang, Yichi Zhang, Yizhi Zhang, Yongting Zhang, Yu Zhang, Yushun Zhang, Yutao Zhang, Yutong Zhang, Zheng Zhang, Chenguang Zhao, Feifan Zhao, Jinxiang Zhao, Shuai Zhao, Xiangyu Zhao, Yikai Zhao, Zijia Zhao, Huabin Zheng, Ruihan Zheng, Shaojie Zheng, Tengyang Zheng, Junfeng Zhong, Longguang Zhong, Weiming Zhong, M. Zhou, Runjie Zhou, Xinyu Zhou, Zaida Zhou, Jinguo Zhu, Liya Zhu, Xinhao Zhu, Yuxuan Zhu, Zhen Zhu, Jingze Zhuang, Weiyu Zhuang, Ying Zou, Xinxing Zu

机构 * Kimi Team(Kimi 团队)

AI总结 Kimi K2.5通过联合优化文本和视觉模态,提出Agent Swarm框架,实现多模态代理智能的先进性能和高效任务处理。

Comments Kimi K2.5 tech report

详情
AI中文摘要

我们介绍了Kimi K2.5,一个开源的多模态代理模型,旨在推进通用代理智能。K2.5强调文本和视觉的联合优化,使两种模态相互增强。这包括一系列技术,如联合文本-视觉预训练、零视觉SFT和联合文本-视觉强化学习。在此多模态基础上,K2.5引入了Agent Swarm,一种自我导向的并行代理 orchestration 框架,能够动态地将复杂任务分解为异构子问题并并发执行。广泛的评估表明,Kimi K2.5在包括编程、视觉、推理和代理任务在内的各种领域均取得了最先进的结果。Agent Swarm还使延迟降低了高达4.5倍。我们发布了后训练的Kimi K2.5模型检查点,以促进未来代理智能的研究和实际应用。

英文摘要

We introduce Kimi K2.5, an open-source multimodal agentic model designed to advance general agentic intelligence. K2.5 emphasizes the joint optimization of text and vision so that two modalities enhance each other. This includes a series of techniques such as joint text-vision pre-training, zero-vision SFT, and joint text-vision reinforcement learning. Building on this multimodal foundation, K2.5 introduces Agent Swarm, a self-directed parallel agent orchestration framework that dynamically decomposes complex tasks into heterogeneous sub-problems and executes them concurrently. Extensive evaluations show that Kimi K2.5 achieves state-of-the-art results across various domains including coding, vision, reasoning, and agentic tasks. Agent Swarm also reduces latency by up to $4.5\times$ over single-agent baselines. We release the post-trained Kimi K2.5 model checkpoint to facilitate future research and real-world applications of agentic intelligence.

URL PDF HTML 收藏
2507.20534 2026-02-04 cs.LG cs.AI cs.CL

Kimi K2: Open Agentic Intelligence

Kimi K2:开放代理智能

Kimi Team, Yifan Bai, Yiping Bao, Y. Charles, Cheng Chen, Guanduo Chen, Haiting Chen, Huarong Chen, Jiahao Chen, Ningxin Chen, Ruijue Chen, Yanru Chen, Yuankun Chen, Yutian Chen, Zhuofu Chen, Jialei Cui, Hao Ding, Mengnan Dong, Angang Du, Chenzhuang Du, Dikang Du, Yulun Du, Yu Fan, Yichen Feng, Kelin Fu, Bofei Gao, Chenxiao Gao, Hongcheng Gao, Peizhong Gao, Tong Gao, Yuyao Ge, Shangyi Geng, Qizheng Gu, Xinran Gu, Longyu Guan, Haiqing Guo, Jianhang Guo, Xiaoru Hao, Tianhong He, Weiran He, Wenyang He, Yunjia He, Chao Hong, Hao Hu, Yangyang Hu, Zhenxing Hu, Weixiao Huang, Zhiqi Huang, Zihao Huang, Tao Jiang, Zhejun Jiang, Xinyi Jin, Yongsheng Kang, Guokun Lai, Cheng Li, Fang Li, Haoyang Li, Ming Li, Wentao Li, Yang Li, Yanhao Li, Yiwei Li, Zhaowei Li, Zheming Li, Hongzhan Lin, Xiaohan Lin, Zongyu Lin, Chengyin Liu, Chenyu Liu, Hongzhang Liu, Jingyuan Liu, Junqi Liu, Liang Liu, Shaowei Liu, T. Y. Liu, Tianwei Liu, Weizhou Liu, Yangyang Liu, Yibo Liu, Yiping Liu, Yue Liu, Zhengying Liu, Enzhe Lu, Haoyu Lu, Lijun Lu, Yashuo Luo, Shengling Ma, Xinyu Ma, Yingwei Ma, Shaoguang Mao, Jie Mei, Xin Men, Yibo Miao, Siyuan Pan, Yebo Peng, Ruoyu Qin, Zeyu Qin, Bowen Qu, Zeyu Shang, Lidong Shi, Shengyuan Shi, Feifan Song, Jianlin Su, Zhengyuan Su, Lin Sui, Xinjie Sun, Flood Sung, Yunpeng Tai, Heyi Tang, Jiawen Tao, Qifeng Teng, Chaoran Tian, Chensi Wang, Dinglu Wang, Feng Wang, Hailong Wang, Haiming Wang, Jianzhou Wang, Jiaxing Wang, Jinhong Wang, Shengjie Wang, Shuyi Wang, Si Wang, Xinyuan Wang, Yao Wang, Yejie Wang, Yiqin Wang, Yuxin Wang, Yuzhi Wang, Zhaoji Wang, Zhengtao Wang, Zhengtao Wang, Zhexu Wang, Chu Wei, Qianqian Wei, Haoning Wu, Wenhao Wu, Xingzhe Wu, Yuxin Wu, Chenjun Xiao, Jin Xie, Xiaotong Xie, Weimin Xiong, Boyu Xu, Jinjing Xu, L. H. Xu, Lin Xu, Suting Xu, Weixin Xu, Xinran Xu, Yangchuan Xu, Ziyao Xu, Jing Xu, Jing Xu, Junjie Yan, Yuzi Yan, Hao Yang, Xiaofei Yang, Yi Yang, Ying Yang, Zhen Yang, Zhilin Yang, Zonghan Yang, Haotian Yao, Xingcheng Yao, Wenjie Ye, Zhuorui Ye, Bohong Yin, Longhui Yu, Enming Yuan, Hongbang Yuan, Mengjie Yuan, Siyu Yuan, Haobing Zhan, Dehao Zhang, Hao Zhang, Wanlu Zhang, Xiaobin Zhang, Yadong Zhang, Yangkun Zhang, Yichi Zhang, Yizhi Zhang, Yongting Zhang, Yu Zhang, Yutao Zhang, Yutong Zhang, Zheng Zhang, Haotian Zhao, Yikai Zhao, Zijia Zhao, Huabin Zheng, Shaojie Zheng, Longguang Zhong, Jianren Zhou, Xinyu Zhou, Zaida Zhou, Jinguo Zhu, Zhen Zhu, Weiyu Zhuang, Xinxing Zu

机构 * Kimi Team(Kimi 团队)

AI总结 Kimi K2是一款开源大语言模型,通过混合专家架构和MuonClip优化器实现先进代理能力,表现优于现有非思考模型。

Comments tech report of Kimi K2, with minor updates

详情
AI中文摘要

我们介绍了Kimi K2,一个具有320亿激活参数和1万亿总参数的混合专家(MoE)大语言模型。我们提出了MuonClip优化器,通过一种新的QK-clip技术改进Muon,以解决训练不稳定问题,同时享受Muon的先进token效率。基于MuonClip,K2在15.5万亿个token上进行了预训练,零损失峰值。在后训练过程中,K2经历了一个多阶段的后训练过程,突出特点是大规模的代理数据合成管道和联合强化学习(RL)阶段,其中模型通过与真实和合成环境的交互来提升其能力。Kimi K2在开源非思考模型中实现了最先进的性能,其在代理能力方面表现突出。值得注意的是,K2在Tau2-Bench上获得66.1分,在ACEBench(En)上获得76.5分,在SWE-Bench Verified上获得65.8分,在SWE-Bench Multilingual上获得47.3分,超过了大多数开源和闭源基础的非思考设置。它在编码、数学和推理任务中也表现出强大的能力,得分分别为53.7分(LiveCodeBench v6)、49.5分(AIME 2025)、75.1分(GPQA-Diamond)和27.1分(OJBench),均无扩展思考。这些结果使Kimi K2成为迄今为止最能干的开源大语言模型之一,特别是在软件工程和代理任务中。我们发布了基础和后训练模型检查点,以促进未来代理智能的研究和应用。

英文摘要

We introduce Kimi K2, a Mixture-of-Experts (MoE) large language model with 32 billion activated parameters and 1 trillion total parameters. We propose the MuonClip optimizer, which improves upon Muon with a novel QK-clip technique to address training instability while enjoying the advanced token efficiency of Muon. Based on MuonClip, K2 was pre-trained on 15.5 trillion tokens with zero loss spike. During post-training, K2 undergoes a multi-stage post-training process, highlighted by a large-scale agentic data synthesis pipeline and a joint reinforcement learning (RL) stage, where the model improves its capabilities through interactions with real and synthetic environments. Kimi K2 achieves state-of-the-art performance among open-source non-thinking models, with strengths in agentic capabilities. Notably, K2 obtains 66.1 on Tau2-Bench, 76.5 on ACEBench (En), 65.8 on SWE-Bench Verified, and 47.3 on SWE-Bench Multilingual -- surpassing most open and closed-sourced baselines in non-thinking settings. It also exhibits strong capabilities in coding, mathematics, and reasoning tasks, with a score of 53.7 on LiveCodeBench v6, 49.5 on AIME 2025, 75.1 on GPQA-Diamond, and 27.1 on OJBench, all without extended thinking. These results position Kimi K2 as one of the most capable open-source large language models to date, particularly in software engineering and agentic tasks. We release our base and post-trained model checkpoints to facilitate future research and applications of agentic intelligence.

URL PDF HTML 收藏
2509.23045 2025-12-09 cs.AI cs.CL cs.SE

Kimi-Dev: Agentless Training as Skill Prior for SWE-Agents

Zonghan Yang, Shengjie Wang, Kelin Fu, Wenyang He, Weimin Xiong, Yibo Liu, Yibo Miao, Bofei Gao, Yejie Wang, Yingwei Ma, Yanhao Li, Yue Liu, Zhenxing Hu, Kaitai Zhang, Shuyi Wang, Huarong Chen, Flood Sung, Yang Liu, Yang Gao, Zhilin Yang, Tianyu Liu

机构 * Moonshot AI THU(清华大学) PKU(北京大学) UCAS(中国科学技术大学) BUPT(北京邮电大学) NUS(新加坡国立大学)

Comments 68 pages. GitHub repo at https://github.com/MoonshotAI/Kimi-Dev

详情
英文摘要

Large Language Models (LLMs) are increasingly applied to software engineering (SWE), with SWE-bench as a key benchmark. Solutions are split into SWE-Agent frameworks with multi-turn interactions and workflow-based Agentless methods with single-turn verifiable steps. We argue these paradigms are not mutually exclusive: reasoning-intensive Agentless training induces skill priors, including localization, code edit, and self-reflection that enable efficient and effective SWE-Agent adaptation. In this work, we first curate the Agentless training recipe and present Kimi-Dev, an open-source SWE LLM achieving 60.4\% on SWE-bench Verified, the best among workflow approaches. With additional SFT adaptation on 5k publicly-available trajectories, Kimi-Dev powers SWE-Agents to 48.6\% pass@1, on par with that of Claude 3.5 Sonnet (241022 version). These results show that structured skill priors from Agentless training can bridge workflow and agentic frameworks for transferable coding agents.

URL PDF HTML 收藏
2510.19562 2025-10-24 cs.AI

DAIL: Beyond Task Ambiguity for Language-Conditioned Reinforcement Learning

Runpeng Xie, Quanwei Wang, Hao Hu, Zherui Zhou, Ni Mu, Xiyun Li, Yiqin Yang, Shuang Xu, Qianchuan Zhao, Bo XU

机构 * The Key Laboratory of Cognition and Decision Intelligence for Complex Systems Institute of Automation Chinese Academy of Sciences(认知与决策智能复杂系统重点实验室,自动化研究所,中国科学院) Department of Automation Tsinghua University(清华大学自动化系) Moonshot AI Department of Computer Science and Engineering Washington University(华盛顿大学计算机科学与工程系) Tecent AI Lab(腾讯AI实验室)

Comments Website at: https://github.com/RunpengXie/Distributional-Aligned-Learning

详情
英文摘要

Comprehending natural language and following human instructions are critical capabilities for intelligent agents. However, the flexibility of linguistic instructions induces substantial ambiguity across language-conditioned tasks, severely degrading algorithmic performance. To address these limitations, we present a novel method named DAIL (Distributional Aligned Learning), featuring two key components: distributional policy and semantic alignment. Specifically, we provide theoretical results that the value distribution estimation mechanism enhances task differentiability. Meanwhile, the semantic alignment module captures the correspondence between trajectories and linguistic instructions. Extensive experimental results on both structured and visual observation benchmarks demonstrate that DAIL effectively resolves instruction ambiguities, achieving superior performance to baseline methods. Our implementation is available at https://github.com/RunpengXie/Distributional-Aligned-Learning.

URL PDF HTML 收藏
2508.09123 2025-10-07 cs.AI cs.CV

OpenCUA: Open Foundations for Computer-Use Agents

Xinyuan Wang, Bowen Wang, Dunjie Lu, Junlin Yang, Tianbao Xie, Junli Wang, Jiaqi Deng, Xiaole Guo, Yiheng Xu, Chen Henry Wu, Zhennan Shen, Zhuokai Li, Ryan Li, Xiaochuan Li, Junda Chen, Boyuan Zheng, Peihang Li, Fangyu Lei, Ruisheng Cao, Yeqiao Fu, Dongchan Shin, Martin Shin, Jiarui Hu, Yuyan Wang, Jixuan Chen, Yuxiao Ye, Danyang Zhang, Dikang Du, Hao Hu, Huarong Chen, Zaida Zhou, Haotian Yao, Ziwei Chen, Qizheng Gu, Yipu Wang, Heng Wang, Diyi Yang, Victor Zhong, Flood Sung, Y. Charles, Zhilin Yang, Tao Yu

机构 * XLANG Lab, The University of Hong Kong(香港大学XLANG实验室) Moonshot AI Stanford University(斯坦福大学) University of Waterloo(滑铁卢大学) Carnegie Mellon University(卡内基梅隆大学)

Comments Updata author list, modify first page format, correct typos

详情
英文摘要

Vision-language models have demonstrated impressive capabilities as computer-use agents (CUAs) capable of automating diverse computer tasks. As their commercial potential grows, critical details of the most capable CUA systems remain closed. As these agents will increasingly mediate digital interactions and execute consequential decisions on our behalf, the research community needs access to open CUA frameworks to study their capabilities, limitations, and risks. To bridge this gap, we propose OpenCUA, a comprehensive open-source framework for scaling CUA data and foundation models. Our framework consists of: (1) an annotation infrastructure that seamlessly captures human computer-use demonstrations; (2) AgentNet, the first large-scale computer-use task dataset spanning 3 operating systems and 200+ applications and websites; (3) a scalable pipeline that transforms demonstrations into state-action pairs with reflective long Chain-of-Thought reasoning that sustain robust performance gains as data scales. Our end-to-end agent models demonstrate strong performance across CUA benchmarks. In particular, OpenCUA-72B achieves an average success rate of 45.0% on OSWorld-Verified, establishing a new state-of-the-art (SOTA) among open-source models. Further analysis confirms that our approach generalizes well across domains and benefits significantly from increased test-time computation. We release our annotation tool, datasets, code, and models to build open foundations for further CUA research.

URL PDF HTML 收藏
2407.00079 2025-09-04 cs.DC cs.AI cs.AR

Mooncake: A KVCache-centric Disaggregated Architecture for LLM Serving

Ruoyu Qin, Zheming Li, Weiran He, Mingxing Zhang, Yongwei Wu, Weimin Zheng, Xinran Xu

机构 * Moonshot AI Tsinghua University(清华大学)

Comments 23 pages, 13 figures

详情
英文摘要

Mooncake is the serving platform for Kimi, a leading LLM service provided by Moonshot AI. It features a KVCache-centric disaggregated architecture that separates the prefill and decoding clusters. It also leverages the underutilized CPU, DRAM, and SSD resources of the GPU cluster to implement a disaggregated cache of KVCache. The core of Mooncake is its KVCache-centric scheduler, which balances maximizing overall effective throughput while meeting latency-related Service Level Objectives (SLOs). Unlike traditional studies that assume all requests will be processed, Mooncake faces challenges due to highly overloaded scenarios. To mitigate these, we developed a prediction-based early rejection policy. Experiments show that Mooncake excels in long-context scenarios. Compared to the baseline method, Mooncake can achieve up to a 525% increase in throughput in certain simulated scenarios while adhering to SLOs. Under real workloads, Mooncake's innovative architecture enables Kimi to handle 75% more requests.

URL PDF HTML 收藏
2508.14644 2025-08-21 cs.AI

LeanGeo: Formalizing Competitional Geometry problems in Lean

Chendong Song, Zihan Wang, Frederick Pu, Haiming Wang, Xiaohan Lin, Junqi Liu, Jia Li, Zhengying Liu

机构 * Moonshot AI Numina Peking University(北京大学) University of Toronto(多伦多大学)

Comments 28 pages

详情
英文摘要

Geometry problems are a crucial testbed for AI reasoning capabilities. Most existing geometry solving systems cannot express problems within a unified framework, thus are difficult to integrate with other mathematical fields. Besides, since most geometric proofs rely on intuitive diagrams, verifying geometry problems is particularly challenging. To address these gaps, we introduce LeanGeo, a unified formal system for formalizing and solving competition-level geometry problems within the Lean 4 theorem prover. LeanGeo features a comprehensive library of high-level geometric theorems with Lean's foundational logic, enabling rigorous proof verification and seamless integration with Mathlib. We also present LeanGeo-Bench, a formal geometry benchmark in LeanGeo, comprising problems from the International Mathematical Olympiad (IMO) and other advanced sources. Our evaluation demonstrates the capabilities and limitations of state-of-the-art Large Language Models on this benchmark, highlighting the need for further advancements in automated geometric reasoning. We open source the theorem library and the benchmark of LeanGeo at https://github.com/project-numina/LeanGeo/tree/master.

URL PDF HTML 收藏
2508.11739 2025-08-19 cs.LG cs.CV

Scalable Geospatial Data Generation Using AlphaEarth Foundations Model

Luc Houriez, Sebastian Pilarski, Behzad Vahedi, Ali Ahmadalipour, Teo Honda Scully, Nicholas Aflitto, David Andre, Caroline Jaffe, Martha Wedner, Rich Mazzola, Josh Jeffery, Ben Messinger, Sage McGinley-Smith, Sarah Russell

机构 * X, the Moonshot Factory, Bellwether Stanford University(Moonshot Factory,Bellwether 斯坦福大学)

Comments 15 pages, 10 figures, 5 tables

详情
英文摘要

High-quality labeled geospatial datasets are essential for extracting insights and understanding our planet. Unfortunately, these datasets often do not span the entire globe and are limited to certain geographic regions where data was collected. Google DeepMind's recently released AlphaEarth Foundations (AEF) provides an information-dense global geospatial representation designed to serve as a useful input across a wide gamut of tasks. In this article we propose and evaluate a methodology which leverages AEF to extend geospatial labeled datasets beyond their initial geographic regions. We show that even basic models like random forests or logistic regression can be used to accomplish this task. We investigate a case study of extending LANDFIRE's Existing Vegetation Type (EVT) dataset beyond the USA into Canada at two levels of granularity: EvtPhys (13 classes) and EvtGp (80 classes). Qualitatively, for EvtPhys, model predictions align with ground truth. Trained models achieve 81% and 73% classification accuracy on EvtPhys validation sets in the USA and Canada, despite discussed limitations.

URL PDF HTML 收藏