arXivDaily arXiv每日学术速递 周一至周五更新

大厂专区

ByteDance(字节跳动)

至 收录 629
2607.18082 2026-07-21 cs.LG cs.AI 新提交

Enhancing Rubric-based RL via Self-Distillation

通过自蒸馏增强基于评分标准的强化学习

Mingxuan Xia, Yuhang Yang, Chao Ye, Shuai Zhu, Shenzhi Yang, Guangcheng Zhu, Yuhang Zhang, Cheng Peng, Haobo Wang, Siqing Wang

机构 * Zhejiang University(浙江大学) ByteDance(字节跳动)

AI总结 研究基于评分标准的强化学习中探索有限问题,提出标准蒸馏策略优化(CriPO),通过策略内自蒸馏同时解决未探索标准和被抑制标准问题,在医学和科学基准测试中表现出色,减少优化步骤并提升性能。

详情
AI中文摘要

基于评分标准的强化学习在改进大语言模型的开放式任务中展现出潜力。其公认的局限是探索有限,未探索标准(UC)未获优化信号。近期方法通过在展开过程中纳入评分标准信息解决此问题,但引入了训练与推理不匹配。此外,这些方法忽视了被抑制标准(SC)这一失败模式。我们的分析表明SC很普遍。为同时解决UC和SC且不引入训练推理不匹配,我们提出标准蒸馏策略优化(CriPO),通过策略内自蒸馏增强基于评分标准的强化学习。在医学和科学基准测试中,CriPO持续优于基于评分标准的强化学习,以约少两倍的优化步骤实现更强的最终性能。

英文摘要

Rubric-based RL has recently shown promise in improving LLMs on open-ended tasks. A widely recognized limitation of rubric-based RL is limited exploration: criteria that no rollout manages to satisfy (Unexplored Criteria, UC) receive no optimization signal. Recent methods address this by incorporating rubric information as external guidance during rollout, yet they introduce a train-inference mismatch: the policy is optimized on rollouts produced under external guidance while this guidance is absent at inference time, causing error accumulation through autoregressive decoding. Moreover, these exploration-focused approaches overlook a fundamentally different failure mode that we term Suppressed Criteria (SC) -- criteria that are satisfied by some rollouts yet whose learning signals are lost during optimization because scalar reward aggregation assigns them non-positive aggregate advantages. Our analysis reveals that SC are remarkably prevalent: over 57% of samples exhibit this failure mode throughout training, with an average of 1.8 SC per sample. To simultaneously address both UC and SC without introducing training-inference mismatch, we propose Criterion-Distilled Policy Optimization (CriPO), which enhances rubric-based RL via on-policy self-distillation. For UC, CriPO constructs a criterion-injection self-teacher and computes a localized forward-KL loss to inject missing behaviors into the policy. For SC, CriPO employs a counterfactual self-teacher to locate criterion-relevant tokens in negative-advantage rollouts and flips their token-level advantages to positive values, preserving useful patterns that would otherwise be suppressed. Experiments on medicine and science benchmarks demonstrate that CriPO consistently outperforms rubric-based RL, achieving stronger final performance with approximately $2\times$ fewer optimization steps.

URL PDF HTML 收藏
2607.16692 2026-07-21 cs.SE cs.CL 新提交

Dependency-Guided Code Generation: Structured Matrix Decomposition and Consistency-Guided Refinement

依赖引导的代码生成:结构化矩阵分解与一致性引导的细化

Mingqiao Mo, Yangchen Zeng, Zikai Xiao, Xin Xiao, Wenhua Nie, Zhaolu Kang, Guangyuan Dong, Kai Shu, Hao Zhang, Xiaodong Fan

机构 * University of the Chinese Academy of Sciences(中国科学院大学) ByteDance Inc.(字节跳动公司) Zhejiang University(浙江大学) National Taiwan University(台湾国立大学) Peking University(北京大学) Alibaba Group(阿里巴巴集团) Tsinghua University(清华大学) Liaoning Technical University(辽宁技术大学)

AI总结 针对现有代码生成方法无法充分捕捉代码实体依赖关系的问题,提出依赖感知代码生成框架,通过结构化矩阵分解和一致性引导细化生成代码,经实验验证该方法能生成语义对齐和结构保真度更高的代码。

Comments 12 pages

详情
AI中文摘要

现代软件系统日益复杂,使自动代码生成成为软件工程中的一项基本任务。然而,现有方法往往无法充分捕捉代码实体间复杂的多层次依赖关系,导致生成的代码逻辑不完整或难以集成到实际系统中。为解决此局限,我们提出一个依赖感知代码生成框架,通过基于图的表示明确建模代码实体间的交互。我们将依赖分解为两个互补组件:一个捕获强显式关系的量化矩阵和一个对弱隐式交互建模的稀疏低秩分解。通过交替优化过程有效学习分解。在代码生成期间,将学习到的依赖结构作为约束纳入,确保生成代码的语义连贯和结构一致。此外,我们为强依赖引入稀疏三元组表示,显著提高存储效率和计算可扩展性。大量实验表明,与现有方法相比,我们的方法始终能生成具有更高语义对齐和结构保真度的代码。

英文摘要

The increasing complexity of modern software systems has made automated code generation a fundamental task in software engineering. However, existing approaches often fail to adequately capture the intricate, multi-level dependencies among code entities, leading to generated code that is logically incomplete or difficult to integrate into real-world systems. To address this limitation, we propose a dependency-aware code generation framework that explicitly models interactions among code entities through a graph-based representation. We decompose dependencies into two complementary components: a quantized matrix that captures strong, explicit relations, and a sparse low-rank factorization that models weaker, implicit interactions. The decomposition is efficiently learned via an alternating optimization procedure. During code generation, the learned dependency structure is incorporated as a constraint, ensuring both semantic coherence and structural consistency of the generated code. Furthermore, we introduce a sparse triplet representation for strong dependencies, significantly improving storage efficiency and computational scalability. Extensive experiments demonstrate that our approach consistently produces code with superior semantic alignment and structural fidelity compared to existing methods.

URL PDF HTML 收藏
2606.11042 2026-07-20 cs.AI 版本更新

Workflow-GYM: Towards Long-Horizon Evaluation of Computer-use Agentic tasks in Real-World Professional Fields

Workflow-GYM:面向真实世界专业领域的长周期计算机使用代理任务评估

Liya Zhu, Jingzhe Ding, Jian Zhang, Jianbo Xue, Shihao Liang, Ge Zhang, Yi Zhu, Duju Zeng, Xiang Gao, Qingshui Gu, Mailun Gao, Huimin Che, Yan Zhao, Peiheng Zhou, Haojun Wang, Chaobo Xian, Lili Le, Chi Wu, Yiwei Liu, Shengda Long, Jiale Yang, Fangzhi Xu, Sijin Wu, Haodong Duan, Chao He, Zhaojian Li, Minchao Wang, Huan Zhou, Jiani Hou, Chuqian Yu, Weiran Shi, Hongwan Gao, Jiamin Chen, Guanhong Chen, Tingqin Luo, Kaiyuan Zhang, Zhixin Yao, Qing Hua, Yuhao Jiang, Jin Chen, Pu Chen, Zhenyu Hu, Xingyu Li, Zhengxuan Jiang, Meng Cao, Tianfeng Long, Haozhe Wang, Mingzhang Wang, Yichen Zhang, Yiming Dai, Chenchen Zhang, Jiaying Wang, Xinying Liu, Xingzu Liu, Lingling Zhang, Xinjie Chen, Yujia Qin, Wangchunshu Zhou, Zhiyong Wu, Yang Liu, Jiaheng Liu, Lei Zhang, Shen Yan, Wenhao Huang, Zaiyuan Wang, Xiaolong Chang

机构 * ByteDance Seed(字节跳动Seed) M-A-P Humanlaya

AI总结 提出Workflow-GYM基准,评估AI代理在专业软件中执行长周期、高价值工作流的能力,发现最强模型成功率仅略超30%,揭示当前代理在长周期工作流一致性方面的严重不足。

详情
AI中文摘要

近年来,AI代理在处理日益复杂、真实世界任务方面取得了快速发展。然而,现有基准很少评估代理能否操作图形用户界面以完成跨领域的长周期、高价值专业工作流。当前的GUI基准仍主要关注通用软件、相对简单的应用和短周期任务,使得现代代理能否遵循用户指令自主操作领域特定专业软件并以端到端方式完成经济价值工作尚不清楚。为填补这一空白,我们引入Workflow-GYM,一个以专业领域和专门软件环境为中心的长周期GUI任务基准。通过对最先进模型的广泛实验,我们发现即使最强的模型也仅达到略高于30%的成功率,突显出专业长周期GUI工作流对当前GUI代理仍极具挑战性。进一步分析表明,当前代理难以维持长周期工作流的一致性,频繁出现工作流阶段遗漏、错误传播、目标漂移以及对专业软件环境理解不足等问题。我们的发现为当前代理系统的局限性提供了重要见解,并为下一代GUI代理研究指明了关键方向。

英文摘要

Recent years have witnessed the rapid evolution of AI agents toward handling increasingly complex, real-world tasks. However, existing benchmarks rarely evaluate whether agents can operate graphical user interfaces to complete long-horizon, high-value professional workflows across diverse domains. Current GUI benchmarks still predominantly focus on general-purpose software, relatively simple applications, and short-horizon tasks, leaving it largely unknown whether modern agents can follow user instructions to autonomously operate domain-specific professional software and accomplish economically valuable work in an end-to-end manner. To bridge this gap, we introduce Workflow-GYM, a benchmark for long-horizon GUI tasks centered on professional domains and specialized software environments. Through extensive experiments on state-of-the-art models, we find that even the strongest models achieve only slightly above 30% success rates, highlighting that professional long-horizon GUI workflows remain highly challenging for current GUI agents. Further analysis reveals that current agents struggle to maintain long-horizon workflow consistency, frequently exhibiting workflow stage omission, error propagation, objective drift, and insufficient understanding of professional software environments. Our findings provide important insights into the limitations of current agent systems and suggest key directions for the next generation of GUI-agent research.

URL PDF HTML 收藏
2607.00724 2026-07-17 cs.CL 版本更新

MSQA: A Natively Sourced Multilingual and Multicultural SimpleQA Benchmark

MSQA:一个原生来源的多语言多文化SimpleQA基准

Xianru Chen, Yukai Huang, Mingxiang Chen, Xinping Lei, Fangbing Deng, Jin Chen, Ge Zhang, Wenhao Huang, Jiaheng Liu

机构 * M-A-P ByteDance Seed(字节跳动Seed) Beijing University of Posts and Telecommunications(北京邮电大学) Nanjing University(南京大学)

AI总结 提出MSQA基准,包含1064个原生问题覆盖11种语言和5个文化维度,评估18个LLM发现文化能力随预训练暴露程度下降,且推理时补救措施无效。

Comments Due to the company's data approval issue, we need to withdraw the article

详情
AI中文摘要

多语言流利性常常引发一个更强的假设:一个能说用户语言的模型也必须理解该语言编码的文化。我们称之为文化对齐的幻觉。为了直接检验这一假设,我们引入了MSQA,一个包含1064个原生来源问题的基准,涵盖11个语言组、五个文化维度和三个难度层级。与翻译基准不同,MSQA针对本地化知识,并减少了来自以英语为中心的跨语言迁移的捷径。评估18个LLM,我们发现显著的文化退化以及明显的局部性效应:文化能力更紧密地追踪预训练暴露程度,而非一般推理能力。我们进一步表明,常见的推理时补救措施并不能消除这种幻觉。模型对不熟悉的文化问题仍然过度自信,重复采样产生不稳定而非可靠的正确性,检索增强对长尾事实的帮助不均匀。这些发现表明,文化对齐不能仅从多语言能力推断,并且需要比推理时的校准、采样或检索更深入的干预。

英文摘要

Multilingual fluency often invites a stronger assumption: a model that can speak a user's language must also understand the culture encoded by that language. We call this the Illusion of Cultural Alignment. To test this assumption directly, we introduce MSQA, a benchmark of 1,064 natively sourced questions across 11 language groups, five cultural dimensions, and three difficulty tiers. Unlike translated benchmarks, MSQA targets locally grounded knowledge and reduces shortcuts from English-centric cross-lingual transfer. Evaluating 18 LLMs, we find substantial cultural degradation and a pronounced Locality Effect: cultural competence tracks pre-training exposure more closely than general reasoning ability. We further show that common inference-time remedies do not dissolve the illusion. Models remain overconfident on unfamiliar cultural questions, repeated sampling yields unstable rather than reliable correctness, and retrieval augmentation helps unevenly on long-tail facts. These findings indicate that cultural alignment cannot be inferred from multilingual ability alone and requires deeper intervention than calibration, sampling, or retrieval at inference time

URL PDF HTML 收藏
2603.20785 2026-07-17 cs.CV 版本更新

ME-IQA: Memory-Enhanced Image Quality Assessment via Re-Ranking

ME-IQA: 通过重新排序增强的记忆图像质量评估

Kanglong Fan, Tianhe Wu, Wen Wen, Jianzhao Liu, Le Yang, Yabin Zhang, Yiting Liao, Junlin Li, Li Zhang

机构 * City University of Hong Kong(香港城市大学) ByteDance Inc.(字节跳动公司)

AI总结 ME-IQA通过构建记忆库和利用推理摘要检索语义和感知对齐的邻居,将VLM重构为概率比较器,并通过门控反思和记忆巩固提升未来决策,从而实现更密集的失真敏感预测,缓解离散崩溃问题。

Comments Published as a conference paper at ECCV 2026

详情
AI中文摘要

通过推理诱导的视觉-语言模型(VLMs)在图像质量评估(IQA)中引入了文本推理,但它们的标量分数往往缺乏敏感性和崩溃到少数值,即离散崩溃。我们引入了ME-IQA,一种即插即用、测试时记忆增强的重新排序框架。它(i)通过推理摘要构建记忆库并检索语义和感知对齐的邻居;(ii)将VLM重新构造成概率比较器以获得成对偏好概率,并在Thurstone的Case V模型下将此顺序证据与初始分数融合;(iii)通过门控反思和记忆巩固来改进未来决策。这产生了更密集、对失真敏感的预测,并缓解了离散崩溃。在多个IQA基准测试中,实验结果显示出在强大的推理诱导VLM基线、现有的非推理IQA方法以及测试时扩展替代方法上的一致提升。

英文摘要

Reasoning-induced vision-language models (VLMs) advance image quality assessment (IQA) with textual reasoning, yet their scalar scores often lack sensitivity and collapse to a few values, so-called discrete collapse. We introduce ME-IQA, a plug-and-play, test-time memory-enhanced re-ranking framework. It (i) builds a memory bank and retrieves semantically and perceptually aligned neighbors using reasoning summaries, (ii) reframes the VLM as a probabilistic comparator to obtain pairwise preference probabilities and fuse this ordinal evidence with the initial score under Thurstone's Case V model, and (iii) performs gated reflection and consolidates memory to improve future decisions. This yields denser, distortion-sensitive predictions and mitigates discrete collapse. Experiments across multiple IQA benchmarks show consistent gains over strong reasoning-induced VLM baselines, existing non-reasoning IQA methods, and test-time scaling alternatives.

URL PDF HTML 收藏
2607.13124 2026-07-16 cs.LG cs.AI cs.CL 新提交

ShortOPD: Recovering Pruned LLMs with Short-to-Long On-Policy Distillation

ShortOPD:通过短到长的策略蒸馏恢复剪枝后的语言模型

Qingyu Zhang, Qianhao Yuan, Hongyu Lin, Yaojie Lu, Xianpei Han, Le Sun, Xiang Li, Ming Xu, Jiarui Li, Xiuyin Zhao

机构 * ByteDance(字节跳动) Institute of Software, Chinese Academy of Sciences(中国科学院软件研究所) University of Chinese Academy of Sciences(中国科学院大学)

AI总结 研究结构化剪枝在语言模型自由形式生成任务中存在的问题,提出ShortOPD方法,通过短到长的策略蒸馏,检测重复后缀,合理分配展开预算,有效提升压缩模型分数,减少训练时间和展开令牌,推动结构化剪枝接近可部署的生成质量。

详情
AI中文摘要

结构化剪枝是一种对硬件友好的语言模型压缩方式,但大多在多项选择识别任务中得到验证,相同的压缩检查点在实际部署所需的自由形式生成任务中可能会崩溃。本文通过两项观察发现了这种差距。首先,贪心的\textsc{pass}@$1$在压缩后几乎消失,但\textsc{pass}@$k$在重复采样下能大幅恢复。其次,可恢复机制主要因后缀重复而失败。因此,恢复应在压缩模型自身的策略状态上进行密集的令牌级监督训练,策略蒸馏(OPD)通过将预压缩模型用作冻结教师来提供这种监督。然而,长时间的策略展开会将早期恢复预算花费在低信息重复后缀上,延迟损失下降。为缓解这种浪费,本文提出了\textbf{\shortopd},一种短到长的OPD调度,它能检测教师确认的重复后缀,将幸存的前缀视为每次展开的有效长度,并将未来的展开预算分配给策略当前可使用的有效长度。在数学、代码和开放式生成任务中,\shortopd\将压缩模型的分数提高到未恢复值的约$9$倍,以及标准恢复方法(无知识蒸馏的监督微调、知识蒸馏和序列知识蒸馏)的$1.6$ - $4.4$倍,并且在两点内匹配固定的$8192$令牌展开范围,使用四分之一的训练时间($8.5$小时对$35.9$小时)和减少$71\%$的展开令牌。希望该方法有助于使结构化剪枝超越在困惑度和多项选择基准上的微小收益,更接近可部署的生成质量。

英文摘要

Structured pruning is a hardware-friendly way to compress LLMs, but it is mostly validated on multiple-choice recognition tasks, while the same compressed checkpoints can collapse on the free-form generation that deployment actually requires. Two observations trace this gap. First, greedy \textsc{pass}@$1$ nearly vanishes after compression, yet \textsc{pass}@$k$ recovers substantially under repeated sampling: useful generations are demoted, not erased. Second, the recoverable regime fails mainly through suffix repetition. Recovery should therefore train on the compressed model's own on-policy states with dense token-level supervision, which On-Policy Distillation (OPD) provides by reusing the pre-compression model as a frozen teacher. However, long on-policy rollouts spend early recovery budget on low-information repetitive suffixes, delaying loss descent. To mitigate this waste, we propose \textbf{\shortopd}, a short-to-long OPD schedule that detects teacher-confirmed repetitive suffixes, treats the surviving prefix as each rollout's effective length, and allocates future rollout budgets to the effective lengths the policy can currently use. Across math, code, and open-ended generation, \shortopd\ raises the compressed model's score to about $9\times$ its unrecovered value and $1.6$--$4.4\times$ standard recovery recipes (SFT w/o KD, KD, and SeqKD), and it matches a fixed $8192$-token rollout horizon within two points using a quarter of the training time ($8.5$ vs.\ $35.9$ hours) and $71\%$ fewer rollout tokens. We hope this recipe helps move structured pruning beyond marginal gains on perplexity and multiple-choice benchmarks, a step closer to deployment-ready generation quality.

URL PDF HTML 收藏
2603.00546 2026-07-16 cs.AI cs.CV 版本更新

Advancing Multimodal Judge Models through a Capability-Oriented Benchmark and MCTS-Driven Data Generation

通过能力导向的基准和MCTS驱动的数据生成推进多模态评判模型

Zeyu Chen, Huanjin Yao, Ziwang Zhao, Min Yang

机构 * Tsinghua University(清华大学) ByteDance(字节跳动)

AI总结 本文提出M-JudgeBench和Judge-MCTS框架,通过能力导向的基准和MCTS驱动的数据生成提升多模态评判模型的评估能力。

详情
AI中文摘要

利用多模态大语言模型(MLLMs)作为评判者以实现精确且一致的评估,已在各种领域逐渐成为一种新兴范式。评估MLLM作为评判系统的能力和可靠性对于确保可信的评估至关重要。现有的评判基准按任务类型对样本进行分类,但未能捕捉到可靠评估所需的基本判断能力。在本工作中,我们引入M-JudgeBench,一个十维的能力导向基准,旨在全面评估MLLM的判断能力。我们的基准将评估分解为成对的链式思维(CoT)比较、长度偏差避免和过程错误检测任务,共同涵盖十个细粒度子任务。这种设计使能够诊断模型在推理风格、响应长度和跨模型变化方面的可靠性。系统性评估揭示了现有MLLM作为评判系统中的系统性弱点。为了解决这个问题,我们进一步提出Judge-MCTS,一个数据构建框架,生成具有各种正确性和长度的成对推理轨迹。使用Judge-MCTS,我们构建了一个MCTS增强的数据集并训练了M-Judger,一系列强大的评判模型。广泛的实验表明,M-Judger在现有评判基准以及M-JudgeBench上都表现出优越性。总体而言,我们的工作通过M-JudgeBench和Judge-MCTS框架建立了更系统的基础来评估MLLM作为评判者,为未来评判模型评估和能力驱动的评判训练铺平了道路。

英文摘要

Using Multimodal Large Language Models (MLLMs) as judges to achieve precise and consistent evaluations has gradually become an emerging paradigm across various domains. Evaluating the capability and reliability of MLLM-as-a-judge systems is therefore essential for ensuring trustworthy assessment. Existing judge benchmarks categorize samples by task types but fail to capture the fundamental judgment capabilities required for reliable evaluation. In this work, we introduce M-JudgeBench, a ten-dimensional capability-oriented benchmark designed to comprehensively assess the judgment abilities of MLLMs. Our benchmark decomposes evaluation into pairwise Chain-of-Thought (CoT) comparison, length bias avoidance, and process error detection tasks, jointly covering ten fine-grained subtasks. This design enables diagnosis of model reliability across reasoning styles, response lengths, and cross-model variations. Systematic evaluation uncovers the systematic weaknesses in existing MLLM-as-a-judge systems. To address this issue, we further propose Judge-MCTS, a data construction framework generating pairwise reasoning trajectories with various correctness and length. Using Judge-MCTS, we construct an MCTS-augmented dataset and train M-Judger, a series of strong judge models. Extensive experiments demonstrate the superiority of M-Judger on existing judge benchmarks as well as M-JudgeBench. Overall, our work establishes a more principled foundation for evaluating MLLM-as-a-judge through M-JudgeBench and Judge-MCTS framework, paving the way for future research on judge model evaluation and capability-driven judge training.

URL PDF HTML 收藏
2607.12800 2026-07-15 cs.CV 新提交

UniVR: Thinking in Visual Space for Unified Visual Reasoning

UniVR:用于统一视觉推理的视觉空间思考

Zhongwei Ren, Yunchao Wei, Yao Zhao, Weibo Gong, Xiao Liu, Anran Wang, Xiangtai Li, Xiaojie Jin

机构 * Beijing Jiaotong University(北京交通大学) ByteDance(字节跳动)

AI总结 UniVR旨在从纯视觉演示中同时学习复杂推理、物理动力学和长期规划,核心是VR-GRPO强化学习范式。构建VR-X基准进行训练和评估,在VR-X上有25%的提升,还提高了多模态理解基准性能,相关资源已开源。

Comments Code and models are released at: https://maverickren.github.io/UniVR.github.io/

详情
AI中文摘要

从原始视觉数据中直接学习广泛的世界知识是智能的一项基本能力。我们引入了UniVR,首次探索从纯视觉演示中同时学习复杂推理、细粒度物理动力学和长期规划。其核心是VR-GRPO,一种具有互补全局和步级奖励的强化学习范式。该方法在整个推理过程中强制逻辑连贯和物理一致性,无需特定任务启发式或图像-文本对。为训练和评估UniVR,我们构建了VR-X,这是一个从16个不同来源策划的大规模基准,涵盖长期操纵、空间谜题和物理推理。它是首个在纯视觉协议下评估这些异构能力的综合套件。值得注意的是,UniVR在VR-X上实现了高达25%的提升,其卓越的视觉推理还提高了各种多模态理解基准的性能。这些发现强调了视觉空间内推理的巨大潜力,所有代码、数据和模型都已开源以供进一步研究。

英文摘要

Learning broad world knowledge directly from raw visual data is a fundamental capability of intelligence. We introduce UniVR, the first investigation into simultaneously learning complex reasoning, fine-grained physical dynamics, and long-term planning from pure visual demonstrations. At its core, UniVR features VR-GRPO, a reinforcement learning paradigm with complementary global and step-level rewards. This approach enforces logical coherence and physical consistency throughout the reasoning process without requiring task-specific heuristics or image-text pairs. To train and evaluate UniVR, we construct VR-X, a large-scale benchmark curated from 16 diverse sources spanning long-horizon manipulation, spatial puzzles, and physical reasoning. It is the first comprehensive suite to assess these heterogeneous capabilities under a purely visual protocol. Remarkably, UniVR achieves up to a 25% improvement on VR-X, and its superior visual reasoning also boosts performance on various multimodal understanding benchmarks. These findings underscore the vast potential of reasoning within visual spaces, with all code, data, and models are open-sourced for further research.

URL PDF HTML 收藏
2607.11327 2026-07-15 cs.LG cs.AI 版本更新

PRISM Edit: One Vector for All Temporal Answers

PRISM Edit:适用于所有时间答案的单一向量

Chen Huang, Qi Zheng, Ruiqin Zheng, Long Zeng, Yuantong Xu

机构 * Tsinghua University(清华大学) ByteDance(字节跳动)

AI总结 研究针对大语言模型时间事实更新问题,基于因果追踪发现其内部计算支持新旧答案区分,进而引入PRISM Edit,通过优化单一多义词表示及利用固有调制路径,在新基准上评估,相比基线提升了时间一致性等指标且速度更快。

Comments Chen Huang and Qi Zheng contributed equally. Corresponding authors: Long Zeng, Yuantong Xu

详情
AI中文摘要

模型编辑可让大语言模型(LLMs)无需重新训练就能保持更新,但时间事实揭示了当前定位与编辑范式的局限性:更新并不总是替换。当事实发生变化时,新答案应成为当前答案,而旧答案在历史时间背景下可能仍然正确。基于此,我们用因果追踪表明LLMs已通过两阶段内部计算支持这种区分:早期MLP层检索与时间无关的主题表示,后期层用时间上下文对其进行调制以产生时间正确的答案。受此发现启发,我们引入PRISM Edit,它在不修改架构的情况下,跨时间上下文优化单个多义词表示,并利用模型固有的调制路径将其引导至时间正确的预测。我们在新引入的时间编辑基准TimeConflict和时间增强的CounterFact上进行评估。PRISM Edit平均比最佳基线提高了23.3的时间一致性(TC)和33.7的当前相对时间得分(CRS),同时速度快2倍以上。代码和数据可在指定网址公开获取。

英文摘要

Model editing keeps large language models (LLMs) up to date without retraining, but temporal facts expose a limitation of the prevailing locate-and-edit paradigm: an update is not always a replacement. When a fact changes, the new answer should become current while the old answer may remain correct in historical time contexts. Building on this insight, we use causal tracing to show that LLMs already support this distinction via a two-stage internal computation: early MLP layers retrieve a time-agnostic subject representation, and later layers modulate it with temporal context to yield the time-correct answer. Motivated by this finding, we introduce PRISM Edit, which optimizes a single polysemous representation across temporal contexts and leverages the model's inherent modulation pathway to route it to temporally correct predictions, without any architectural modification. We evaluate on TimeConflict, a new temporal editing benchmark we introduce, and on temporally augmented CounterFact. PRISM Edit improves over the best baseline by +23.3 Temporal Consistency (TC) and +33.7 Current Relative-time Score (CRS) on average while being more than 2x faster. Code and data are publicly available at https://github.com/AnonymousStudy972/PRISM-Edit.

URL PDF HTML 收藏
2607.10526 2026-07-15 cs.AI 版本更新

Agents Don't Just Agree, They Remember: Benchmarking Persistent Sycophancy in Stateful Personal Agents

智能体不仅会认同,还会记忆:对有状态个人智能体中持续谄媚行为的基准测试

Xutao Mao, Liangjie Zhao, Leyao Wang, Rui Qian, Qiang Huang, Wentao Wang, Bo Han, Xiang Zheng, Cong Wang

机构 * City University of Hong Kong(香港城市大学) ByteDance(字节跳动) Yale University(耶鲁大学) Fudan University(复旦大学) Hong Kong Baptist University(香港浸会大学) Dalian University of Technology(大连理工大学)

AI总结 研究有状态个人智能体中的持续谄媚行为,引入PASB基准测试,评估真实智能体。通过特定方法隔离写入过程,发现提交边界是关键转折点,提交声明有三种模式,表明智能体谄媚行为是状态写入治理问题,PASB确定相关写入时控制。

详情
AI中文摘要

有状态的个人智能体越来越多地维护长期用户档案、情景记忆和可重复使用的技能。这种持续性将对话中的谄媚行为转变为状态写入失败:被接受的以用户为中心的声明可能会被记录为持久偏好、背景事实或工作流程,并在原始对话结束后被再次使用。我们将此称为持续谄媚行为,并引入了个人智能体谄媚行为基准测试(PASB),这是一个包含1600个任务的基准测试,用于追踪对话声明是否被接受、写入持久智能体状态并在后续中立查询中被再次使用。与之前提供预编写记忆的基准测试不同,PASB评估决定存储内容的真实智能体(Hermes-Agent和OpenClaw)。它通过将四种场景框架与四种时间交付模式相结合,并将五轮持久阶段与清除后的三轮查询阶段分开,来隔离写入过程,确保下游影响仅来自持久状态。在十二个模型中,提交边界是关键转折点:下游失败率从仅会话情节中的45.0%增加到提交后的71.9%,一致增加了27.0个百分点。提交的声明呈现出三种写入时模式:状态提升、归因消除和范围扩大。这些模式在类似记忆或程序框架、重复强化甚至跨领域边界的情况下会变得更强。这些结果表明,智能体谄媚行为从根本上说是一个状态写入治理问题。一旦用户内容被提交到持久记忆中,安全性必须控制智能体写入的内容,而不仅仅是它们所说的内容。PASB确定了在保留存储内容的来源、角色和范围的同时,控制有风险提交所需的写入时控制,而不仅仅是在响应级别进行缓解。

英文摘要

Stateful personal agents increasingly maintain long-term user profiles, episodic memories, and reusable skills. This persistence turns conversational sycophancy into a state-writing failure: accepted user-centric claims can be committed as lasting preferences, background facts, or workflows and later reused after the original conversation is gone. We call this persistent sycophancy and introduce the Personal Agent Sycophancy Benchmark (PASB), a 1,600-task benchmark that traces whether a conversational claim is accepted, written into durable agent state, and reused in a later neutral query. Unlike prior benchmarks that provide pre-written memories, PASB evaluates real agents (Hermes-Agent and OpenClaw) that decide what to store. It isolates the write process by combining four scenario framings with four temporal delivery patterns and separating a five-turn persist stage from a cleared three-turn query stage, ensuring downstream effects arise only from durable state. Across twelve models, the commit boundary is the key inflection point: downstream failure increases from 45.0% in session-only episodes to 71.9% after commitment, a consistent increase of 27.0 percentage points. Committed claims exhibit three write-time patterns: status promotion, attribution removal, and scope broadening. These patterns become stronger under memory-like or procedural framing, repeated reinforcement, and even across domain boundaries. These results show that agent sycophancy is fundamentally a state-writing governance problem. Once user content is committed to durable memory, safety must govern what agents write, not only what they say. PASB identifies the write-time controls needed to gate risky commits while preserving the source, role, and scope of stored content beyond response-level mitigations.

URL PDF HTML 收藏
2604.19632 2026-07-15 cs.CV 版本更新

CreatiParser: Generative Image Parsing of Raster Graphic Designs into Editable Layers

CreatiParser: 从位图图形设计生成可编辑的图层

Weidong Chen, Dexiang Hong, Zhendong Mao, Yutao Cheng, Xinyan Liu, Lei Zhang, Yongdong Zhang

机构 * School of Information Science and Technology, University of Science and Technology of China(科学技术大学信息科学与技术学院) ByteDance Intelligent Creation(字节跳动智能创作) School of Computer Science and Technology, Harbin Institute of Technology (Weihai)(哈尔滨工业大学(威海)计算机科学与技术学院) Institute of Artificial Intelligence, Hefei Comprehensive National Science Center(合肥综合国家科学中心人工智能研究院)

AI总结 本文提出CreatiParser框架,将位图图形设计分解为可编辑的文本、背景和贴纸图层,结合视觉语言模型和多分支扩散架构,提升生成质量与编辑灵活性,实验显示在Parser-40K和Crello数据集上性能优于现有方法。

详情
AI中文摘要

图形设计图像由多个可编辑图层组成,如文本、背景和装饰元素,而大多数生成模型生成的是无显式图层结构的位图输出,限制了后续编辑。现有图形设计解析方法通常依赖多阶段流水线结合布局预测、遮罩和修补,存在误差累积和可控性有限的问题。我们提出一种混合生成框架,将位图转换为可编辑的图层,通过视觉语言模型解析文本区域为文本渲染协议,实现忠实重建和灵活编辑,背景和贴纸图层通过多分支扩散架构生成,支持RGBA。进一步引入ParserReward并结合Group Relative Policy Optimization,使生成质量与人类设计偏好对齐。在Parser-40K和Crello数据集上的大量实验表明,该方法在所有指标上均优于现有方法,例如平均提升23.7%。

英文摘要

Graphic design images consist of multiple editable layers, such as text, background, and decorative elements, while most generative models produce rasterized outputs without explicit layer structures, limiting downstream editing. Existing graphic design parsing methods typically rely on multi-stage pipelines combining layout prediction, matting, and inpainting, which suffer from error accumulation and limited controllability. We propose a hybrid generative framework for raster-to-layer graphic design parsing that decomposes a design image into editable text, background, and sticker layers. Text regions are parsed using a vision-language model into a text rendering protocol, enabling faithful reconstruction and flexible re-editing, while background and sticker layers are generated using a multi-branch diffusion architecture with RGBA support. We further introduce ParserReward and integrate it with Group Relative Policy Optimization to align generation quality with human design preferences. Extensive experiments on two challenging datasets, \emph{i.e.,} the Parser-40K and Crello datasets, demonstrate superior performance over existing methods, \emph{eg.,} achieving an overall average improvement of 23.7\% across all metrics.

URL PDF HTML 收藏
2607.11886 2026-07-14 cs.CV 新提交

Read It Back: Pretrained MLLMs Are Zero-Shot Reward Models for Text-to-Image Generation

读回:预训练的多模态语言模型是文本到图像生成的零样本奖励模型

Runhui Huang, Qihui Zhang, Zhe Liu, Yu Gao, Jie Wu, Hengshuang Zhao

机构 * The University of Hong Kong(香港大学) ByteDance Seed(字节跳动Seed) Peking University(北京大学)

AI总结 研究提出SpectraReward将预训练多模态语言模型转为图像生成强化学习奖励模型,用图像条件提示对数似然作奖励,还引入Self-SpectraReward形成闭环框架。经广泛实验验证,二者能提升生成性能,表明奖励-策略对齐是关键。

详情
AI中文摘要

在本文中,我们提出了SpectraReward,一种无需训练的奖励函数,它将预训练的多模态语言模型转变为用于图像生成强化学习的现成奖励模型。SpectraReward不是要求多模态语言模型判断生成的图像或回答分解的验证问题,而是通过单次图像条件下的教师强制前向传递来衡量从生成的图像中恢复原始提示的程度。我们使用平均图像条件提示对数似然作为奖励,直接重用多模态语言模型的预训练图像-文本对齐能力,无需偏好标签和奖励模型微调。我们进一步引入了Self-SpectraReward,这是统一多模态模型的一种特殊情况,其中策略自身的理解分支作为其生成分支的奖励模型,形成了一个无需外部奖励模型或外部知识的闭环自我改进框架。广泛的实验通过涵盖两个扩散模型、三种强化学习算法、来自四个多模态语言模型家族的九个奖励多模态语言模型主干(参数跨度从4B到235B)以及五个分布外文本到图像基准的广泛图像生成强化学习研究验证了SpectraReward。结果表明,SpectraReward和Self-SpectraReward都显著且持续地提高了生成性能,并且优于先前基于多模态语言模型的奖励训练方法。进一步的分析表明,更大的奖励多模态语言模型并不总是更好,而Self-SpectraReward可以匹配或超过大得多的外部奖励模型,这表明奖励-策略对齐是有效图像生成强化学习的关键因素。

英文摘要

In this paper, we propose SpectraReward, a training-free reward function that turns pretrained MLLMs into off-the-shelf reward models for image-generation reinforcement learning. Instead of asking the MLLM to judge a generated image or answer decomposed verification questions, SpectraReward measures how well the original prompt can be recovered from the generated image through a single image-conditioned, teacher-forced forward pass. We use the average image-conditioned prompt log-likelihood as the reward, directly reusing the MLLM's pretrained image-text alignment ability without preference labels, reward-model fine-tuning. We further introduce Self-SpectraReward, a special case for unified multimodal models where the policy's own understanding branch serves as the reward model for its generation branch, forming a closed-loop self-improving framework without external reward models or external knowledge. Extensive experiments validate SpectraReward through a broad image-generation RL study covering two diffusion models, three RL algorithms, nine reward MLLM backbones from four MLLM families spanning 4B to 235B parameters, and five out-of-distribution text-to-image benchmarks. Results show that both SpectraReward and Self-SpectraReward significantly and consistently improve generation performance and outperform prior MLLM-derived reward training methods. Further analysis reveals that larger reward MLLMs are not always better, while Self-SpectraReward can match or surpass much larger external reward models, suggesting that reward-policy alignment is a key factor for effective image-generation RL. Project Page: https://huangrh99.github.io/SpectraReward/

URL PDF HTML 收藏
2607.11175 2026-07-14 cs.AI 新提交

The Path to Self-Evolving Clinical Systems: Scaling Medical Agents from Assistance to Autonomy

自我进化临床系统之路:将医疗智能体从辅助扩展到自主

Chunzheng Zhu, Lei Tian, Bohan Tan, Ziqi Zhou, Yuxuan Sun, Yijun Wang, Chengchao Lv, Yilin Wen, Yijun He, Jinghao Lin, Yihang Chen, Cheewei Tan, Qianshan Wei, Lei Zhao, Bin Pu, Kenli Li, Yuan Xue, Jianxin Lin

机构 * Hunan University(湖南大学) ByteDance(字节跳动) Duke University(杜克大学) Westlake University(西湖大学) The University of Hong Kong(香港大学) Nanyang Technological University(南洋理工大学) Institute of Automation, Chinese Academy of Sciences(中国科学院自动化研究所) University of Macau(澳门大学) The Ohio State University(俄亥俄州立大学)

AI总结 研究探讨大语言模型等对医疗智能体的重塑,从临床部署出发,将其形式化为决策系统并给出自主性分类。沿统一框架扩展,强调临床环境扩展为关键方向,定位临床自我进化为前沿,还研究了多领域应用及挑战,提供医学成像系统路线图。

Comments Project page: https://github.com/zhcz328/Awesome-Medical-Agents

详情
AI中文摘要

大语言模型和视觉语言模型联合解释和推理图像与文本的能力不断增强,正在重塑医疗智能体,使其从特定任务预测器向能在临床环境中感知、推理、规划、记忆和行动的自主系统转变。本研究从临床部署出发,探讨医疗智能体在实际应用中所需的任务、抗污染基准和交互式训练环境。将医疗智能体形式化为部分可观测下的序列决策系统,并给出了辅助、合作和完全自主操作的三级自主性分类法。沿着由框架扩展、能力扩展和环境扩展组成的统一扩展框架,临床环境扩展被视为在PACS、EHR和FHIR生态系统中运行的智能体最具可行性但未充分探索的方向。临床自我进化被定位为关键研究前沿,借鉴自我改进智能体、智能体训练环境和测试时计算扩展的见解。研究了放射学、病理学、眼科和医院工作流程中的应用以及包括幻觉、级联故障和公平性在内的部署挑战。通过整合300多篇参考文献,特别是2025年至2026年的进展,为实际临床实践中可信的、自我改进的医学成像系统提供了路线图。

英文摘要

The growing ability of large language models and vision language models to jointly interpret and reason over images and text is reshaping medical agents, moving them from task specific predictors toward autonomous systems that perceive, reason, plan, remember, and act in clinical environments. This work departs from the capability first perspective of existing literature and instead begins from clinical deployment, asking what tasks, contamination resistant benchmarks, and interactive training environments are required before medical agents can be trusted in practice. Medical agents are formalized as sequential decision making systems under partial observability, together with a three level autonomy taxonomy spanning assisted, cooperative, and fully autonomous operation. The field is organized along a unified scaling spine consisting of framework scaling, capability scaling, and environment scaling. Within this framework, clinical environment scaling, the integration of tools, data, and clinical gyms, is identified as the most actionable yet underexplored direction for agents operating in PACS, EHR, and FHIR ecosystems. Clinical self evolution, where agents improve through interaction with their environments rather than parameter scaling alone, is further positioned as a key research frontier, drawing insights from self improving agents, agent gyms, and test time compute scaling. Applications across radiology, pathology, ophthalmology, and hospital workflows are examined together with deployment challenges including hallucination, cascading failures, and fairness. By consolidating more than 300 references, with particular emphasis on advances from 2025 to 2026, this work provides a roadmap toward trustworthy, self improving medical imaging systems for real clinical practice.

URL PDF HTML 收藏
2606.31651 2026-07-14 cs.AI 版本更新

FARS: A Fully Automated Research System Deployed at Scale

FARS:一个大规模部署的全自动研究系统

Qiong Tang, Tianxiang Sun, Xiangkun Hu, Xiangyang Liu, Yiran Chen, Yunfan Shao, Bobo Li, Changze Lv, Cheng Xu, Chengsong Huang, Chunyang Li, Dizhan Xue, Hao Bai, Haodong Duan, Hengquan Guo, Hongyang He, Hongyi Chen, Hui Shen, Jiahao Yuan, Jiankai Sun, Jikang Cheng, Jinfeng Xu, Jingqi Tong, Jingye Chen, Jinxiu Liu, Jixuan Leng, Junchi Yu, Kaixun Jiang, Kun Xiang, Kunpeng Yao, Lang Feng, Liangqi Yuan, Longsen Gao, Meng Li, Qi Jia, Qiushi Sun, Shengyuan Ding, Shizhan Gong, Siru Zhong, Terry Jingchen Zhang, Tianle Gu, Tianyi Liang, Weijie Liu, Weikai Yang, Weizhi Fei, Xin Wang, Xinpeng Liu, Xuanwen Ding, Yihong Tang, Yuanli Wang, Yukun Jiang, Yuming Yang, Zhengbao He, Zhikai Chen, Zhikun Xu, Zhuang Li, Zihao Huang

机构 * Analemma National University of Singapore(新加坡国立大学) Fudan University(复旦大学) University College Dublin(都柏林大学) Washington University in St. Louis(圣路易斯华盛顿大学) The Hong Kong University of Science and Technology(香港科技大学) Institute of Automation, Chinese Academy of Sciences(中国科学院自动化研究所) University of Illinois at Urbana-Champaign(伊利诺伊大学厄巴纳-香槟分校) ByteDance(字节跳动) ShanghaiTech University(上海科技大学) University of Warwick(华威大学) Carnegie Mellon University(卡内基梅隆大学) University of Michigan, Ann Arbor(密歇根大学安娜堡分校) East China Normal University(华东师范大学) Stanford University(斯坦福大学) Tencent(腾讯) The University of Hong Kong(香港大学) Shanghai Innovation Institute(上海创新研究院) Nex-AGI Team(Nex-AGI团队)

AI总结 提出FARS系统,通过分阶段智能体协作自动生成研究项目,在67个AI/ML主题上产出166篇论文,经282份评审验证其可产出有价值成果,同时暴露实验范围窄、方法局限和诚信问题。

详情
AI中文摘要

近期的自动化研究系统表明,语言模型智能体可以生成假设、运行实验并撰写完整手稿,但大多数证据仍来自选定的例子、人类设定的主题或少数预定义的研究任务。我们提出了FARS(全自动研究系统),这是一个全自动的AI-for-AI研究系统,旨在跨研究主题大规模运行。FARS通过构思、规划、实验和写作阶段自主生成并推进项目,使用阶段特定的智能体通过共享工作空间进行协调,该工作空间记录提案、代码、日志、结果和手稿。在其首次公开部署中,FARS生成了166篇完整的研究论文,涵盖67个细粒度的AI/ML主题,同时保留中间产物作为可审计的语料库,而非精心挑选的成功案例。我们通过来自志愿评审员的282份结构化评审(涵盖140篇论文)对该语料库进行了评估,包括总体评分、子分数、完整性检查和LLM使用披露。评审表明,FARS可以在大规模公开部署中产生值得评审且偶尔强大的AI/ML研究产物,同时也暴露了在狭窄实验范围、方法局限性和诚信问题方面的重复失败模式。

英文摘要

Recent automated research systems show that language-model agents can generate hypotheses, run experiments, and write complete manuscripts, but most evidence still comes from selected examples, human-framed topics, or a few pre-defined research tasks. We present FARS (Fully Automated Research System), a fully automated AI-for-AI research system designed to operate across research topics at scale. FARS autonomously generates and advances projects through ideation, planning, experimentation, and writing, using stage-specific agents coordinated through a shared workspace that records proposals, code, logs, results, and manuscripts. In its first public deployment, FARS produced 166 complete research papers spanning 67 fine-grained AI/ML topics while preserving intermediate artifacts as an auditable corpus rather than a curated set of successes. We evaluate this corpus with 282 structured reviews from volunteer reviewers covering 140 papers, including overall ratings, sub-scores, integrity checks, and LLM-use disclosure. The reviews indicate that FARS can produce review-worthy and occasionally strong AI/ML research artifacts in a large-scale public deployment, while also exposing recurring failure modes in narrow experimental scope, methodological limitations, and integrity issues.

URL PDF HTML 收藏
2506.16112 2026-07-14 cs.CV 版本更新

AutoV: Loss-Oriented Ranking for Visual Prompt Retrieval in LVLMs

AutoV:面向视觉提示检索的损失导向排名用于大视觉-语言模型

Yuan Zhang, Chun-Kai Fan, Sicheng Yu, Junwen Pan, Tao Huang, Ming Lu, Kuan Cheng, Qi She, Shanghang Zhang

机构 * School of Computer Science, Peking University(北京大学计算机学院) ByteDance Inc.(字节跳动公司) Shanghai Jiao Tong University(上海交通大学)

AI总结 AutoV通过损失导向的提示检索提升大视觉-语言模型在图像理解等任务中的性能。

Comments Accepted by ECCV 2026

详情
AI中文摘要

受大型语言模型中的文本提示启发,视觉提示被探索以增强大型视觉-语言模型(LVLMs)的感知能力。然而,在单一视觉提示设计下,性能趋于饱和,进一步的提示工程变得越来越无效。为了解决这一限制,我们从提示工程转向提示检索,并提出AutoV,一个轻量级的实例自适应视觉提示识别框架。给定一个输入图像和一个文本查询,AutoV会自动从多样化的候选池中定位最合适的视觉提示。训练这样的检索框架需要提示级监督,但提示质量本质上是模糊的,甚至对人类来说也难以可靠评估。为了实现自动监督,我们使用预训练的LVLM评估视觉提示,并根据其预测损失对它们进行标注。使用损失导向的排名作为稳健的训练信号,AutoV学会为每个实例检索查询感知的最优提示,而无需手动标注。实验表明,AutoV在图像理解、描述生成、定位和分类任务中提升了各种LVLMs的性能。例如,AutoV在VizWiz上将LLaVA-OV的性能提高了10.2%,在MMMU上将Qwen2.5-VL的性能提高了3.8%。

英文摘要

Inspired by text prompts in large language models, visual prompts have been explored to enhance the perceptual capabilities of large vision-language models (LVLMs). However, performance tends to saturate under single visual prompt designs, making further prompt engineering increasingly ineffective. To address this limitation, we shift from prompt engineering to prompt retrieval and propose AutoV, a lightweight framework for instance-adaptive visual prompt identification. Given an input image and a textual query, AutoV automatically locates the most suitable visual prompt from a diverse candidate pool. Training such a retrieval framework requires prompt-level supervision, yet prompt quality is inherently ambiguous and difficult to assess reliably, even for humans. To enable automatic supervision, we evaluate visual prompts using a pre-trained LVLM and label them according to their prediction losses. Using the loss-oriented ranking as a robust training signal, AutoV learns to retrieve the query-aware optimal prompt for each instance without manual annotation. Experiments indicate that AutoV enhances the performance of various LVLMs on image understanding, captioning, grounding, and classification tasks. For example, AutoV improves LLaVA-OV by $\textbf{10.2}\%$ on VizWiz and boosts Qwen2.5-VL by $\textbf{3.8}\%$ on MMMU, respectively.

URL PDF HTML 收藏
2511.23191 2026-07-14 cs.CV 版本更新

GeoWorld: Providing Full-frame Geometry Features to Facilitate 3D Scene Generation

GeoWorld:提供全帧几何特征以促进3D场景生成

Yuhao Wan, Lijuan Liu, Jingzhi Zhou, Zihan Zhou, Xuying Zhang, Dongbo Zhang, Shaohui Jiao, Qibin Hou, Ming-Ming Cheng

机构 * VCIP & AAIS, Nankai University(VCIP与AAIS,南开大学) ByteDance Inc.(字节跳动公司) Renmin University of China(中国人民大学) NKIARI, Shenzhen Futian(NKIARI,深圳福田)

AI总结 针对视频模型用于图像到3D场景生成时的几何失真等问题,提出两阶段GeoWorld方法,先由视频模型和多视图几何模型生成全帧几何特征辅助第二阶段,还提几何损失和适配模块,能生成更高保真3D场景且更快。

详情
AI中文摘要

以前利用视频模型进行图像到3D场景生成的工作常常遭受几何失真和内容模糊的困扰。使用视频生成模型根据单帧输入隐式地保持几何一致性是无效的。在本文中,我们提出了一种名为GeoWorld的两阶段方法,通过提供全帧几何特征来革新图像到3D场景生成管道。第一阶段的视频生成模型,随后是多视图几何模型,产生全帧几何特征,然后将其用作几何条件的草图以辅助第二阶段的视频生成模型。提出了一种几何损失来施加现实世界的几何约束,并引入了一个几何适配模块以确保几何特征的有效利用。由于全帧几何建模,我们两阶段方法中的两个较小的视频模型可以生成比SOTA方法更高保真的3D场景,同时甚至更快,例如比混元行者快7.5倍。项目页面:此https URL。

英文摘要

Previous works that leverage video models for image-to-3D scene generation often suffer from geometric distortions and blurry content. Using video generation models to implicitly maintain geometric consistency according to a single-frame input is ineffective. In this paper, we present a two-stage method, named $\textbf{GeoWorld}$, that renovates the image-to-3D scene generation pipeline by providing full-frame geometry features. The first-stage video generation model, followed by a multi-view geometry model, produces $\textbf{full-frame}$ geometry features, which are then used as a mental draft of geometric conditions to aid the second-stage video-generation model. A geometric loss is proposed to impose real-world geometric constraints, and a geometry adaptation module is introduced to ensure the effective utilization of geometry features. Thanks to full-frame geometric modeling, the two smaller video models in our two-stage method can generate higher-fidelity 3D scenes than SOTA methods, while being even faster, e.g. 7.5$\times$ faster than Hunyuan-Voyager. Project page: https://peaes.github.io/GeoWorld.

URL PDF HTML 收藏
2607.07817 2026-07-10 cs.CV cs.AI 新提交

DreamCharacter-1: From 3D Generative Foundation Models to Product-Ready Character Generation

DreamCharacter-1:从3D生成基础模型到产品就绪的角色生成

Weizhe Liu, Yunjie Wu, Xiangqian Shu, Guangwei Wang, Xiangyu Xu, Peng Li, Yujie Li, Hengkai Guo

机构 * ByteDance(字节跳动)

AI总结 研究提出DreamCharacter-1框架,基于3D基础主干,通过几何后期训练、纹理后期训练和推理加速三个组件,将预训练模型校准用于3D角色生成,实验证明该框架能生成高质量角色资产,超越现有方法。

Comments Official Page: https://dreamcharacter-x.github.io/

详情
AI中文摘要

我们展示了DreamCharacter-1,这是一个轻量级的后期适配框架,可将预训练的3D基础模型校准为高保真、产品就绪的3D角色生成。基于3D基础主干,我们的流程包含三个面向任务的组件:几何后期训练、纹理后期训练和推理加速。大量实验表明,DreamCharacter-1能生成视觉上吸引人且结构稳健的3D角色资产,超越现有方法。

英文摘要

We present DreamCharacter-1, a lightweight post-adaptation framework that calibrates pretrained 3D foundation models toward high-fidelity, production-ready 3D character generation. Building upon a 3D foundation backbone, our pipeline incorporates three task-oriented components: (1) geometry post-training, which enhances fine-grained surface details through geometric preference optimization; (2) texture post-training, which synthesizes high-resolution textures and refines the appearance of occluded regions; and (3) inference acceleration, which enables scalable deployment. Extensive quantitative and qualitative experiments demonstrate that DreamCharacter-1 produces visually compelling and structurally robust 3D character assets, consistently surpassing state-of-the-art character generation methods.

URL PDF HTML 收藏
2606.24477 2026-07-10 cs.CV cs.AI cs.SD 新提交

video-SALMONN-R$^3$: Learning to ReWatch, ReAsk, and ReAnswer for Efficient Video Understanding

video-SALMONN-R$^3$: 学习重看、重问和重答以实现高效视频理解

Yixuan Li, Guangzhi Sun, Yudong Yang, Chao Zhang

机构 * Tsinghua University(清华大学) ByteDance(字节跳动) University of Cambridge(剑桥大学)

AI总结 提出video-SALMONN-R$^3$,首个通过强化学习实现重看机制的视频大语言模型,无需链式思维冷启动,采用重答和重问策略提升视频问答效率与准确性。

详情
AI中文摘要

视频大语言模型通常受限于计算和内存预算,导致使用降低的帧率和空间分辨率,可能错过问答所需的关键信息。一种实用且高效的解决方案是两阶段范式:首先进行粗粒度视频理解以定位相关片段,然后以更高的时间或空间保真度重看这些片段。本文提出video-SALMONN-R$^3$,这是首个通过强化学习实现重看机制且不依赖链式思维冷启动的端到端视频大语言模型。该设计消除了昂贵的链式思维数据标注需求,并避免了基于链式思维的有监督微调,后者可能损害预训练的视频理解能力。为解决重看引发的推理优先行为与预训练视频大语言模型回答优先倾向之间的不匹配,我们提出重答策略:模型在首次观看时直接给出答案,重看后对其进行修正。最后,为提升重看过程中的问题遵循度,我们提出重问机制,在重新访问定位片段时重新注入查询。实验结果表明,video-SALMONN-R$^3$在显著降低计算成本的同时,持续优于基础模型和问答有监督微调基线,并超越先前基于重看的方法。代码、模型和数据将在论文被接收后公开。

英文摘要

Video large language models (LLMs) are often constrained by computation and memory budgets, leading them to use reduced frame rates and spatial resolutions, which may cause them to miss critical information for question answering (QA). A practical and efficient solution is a two-stage paradigm: first perform coarse video understanding to localize relevant segments, and then re-watch these segments at higher temporal or spatial fidelity. In this paper, we present video-SALMONN-R$^3$, the first end-to-end video-LLM that enables re-watch through reinforcement learning without relying on chain-of-thought (CoT) cold-start. This design removes the need for costly CoT data annotations and avoids CoT-based supervised fine-tuning (SFT), which can otherwise degrade the pretrained video understanding abilities. To address the mismatch between the reasoning-first behavior induced by re-watch and the answer-first tendency of pretrained video-LLMs, we propose a re-answer strategy, in which the model first produces a direct answer in the first watch and then refines it after re-watching. Finally, to improve question adherence during re-watching, we propose a re-ask mechanism that re-injects the query when revisiting localized segments. Experimental results show that video-SALMONN-R$^3$ consistently outperforms both the base model and the QA-SFT baseline, while surpassing prior re-watch-based approaches with significantly lower computational cost. Code, models, and data will be publicly released upon acceptance.

URL PDF HTML 收藏
2606.06379 2026-07-10 cs.CV cs.AI 版本更新

EasyLens: A Training-Free Plug-and-Play Subtle-Lesion Representation Amplifier for Medical Vision-Language Models

EasyLens: 一种无需训练的即插即用型微病变表示放大器,用于医学视觉语言模型

Qiwei Zeng, Hao Wang, Jinghao Lin, Shuchang Ye, Yuezhe Yang, Yige Peng, Haoyuan Che, Jinman Kim, Lei Bi

机构 * Jilin University(吉林大学) School of Computer Science, The University of Sydney(悉尼大学计算机科学学院) ByteDance(字节跳动) Institute of Translational Medicine, Shanghai Jiao Tong University(上海交通大学转化医学研究院)

AI总结 提出EasyLens,一种无需训练的即插即用模块,通过构建病理-解剖原型空间、反事实推理选择病变相关补丁以及形态引导残差增强,放大医学视觉语言模型对微病变的表示能力。

详情
AI中文摘要

医学视觉语言模型(VLM)在临床图像解读(包括病变检测和报告生成)方面显示出越来越大的潜力。然而,其对微病变的敏感性不足限制了其实用性,因为微病变的视觉证据通常稀疏、低对比度且嵌入复杂的解剖背景中。随着局部视觉标记的聚合,这些微弱的病变线索在全局图像表示中可能变得代表性不足,使得医学VLM难以识别。现有的提高病变敏感性的工作主要依赖于医学领域的视觉编码器预训练、临床术语引导的对齐或可训练的病理表示增强。尽管有效,但这些方法通常需要额外训练或模型特定适配,并可能过度适应特定疾病形态,限制了其在冻结的医学VLM上的适用性。为解决这些限制,我们提出EasyLens,一种无需训练的即插即用型微病变表示放大器,用于医学VLM。EasyLens首先构建EasyBank,一个病理-解剖原型空间,提供病变相关原型和解剖感知的正常参考,用于将可疑补丁与病理和正常解剖模式进行比较。为避免盲目放大正常组织,EasyTag通过反事实原型推理选择病变相关补丁。为抵消全局图像表示中微病变线索的稀释,EasyAmplifier通过形态引导的残差增强强化所选病变相关补丁的表示,从而增加其对全局图像嵌入的贡献。在多个医学图像数据集和冻结的医学VLM骨干上的实验表明,EasyLens改进了微病变检测,并优于现有的编码器增强基线。

英文摘要

Medical vision-language models (VLMs) have shown increasing potential for clinical image interpretation, including lesion detection and report generation. However, their practical utility remains limited by insufficient sensitivity to subtle lesions, whose visual evidence is often sparse, low-contrast, and embedded within complex anatomical context. As local visual tokens are aggregated, these weak lesion cues can become underrepresented in global image representations, making them difficult for medical VLMs to recognize. Existing efforts to improve lesion sensitivity mainly rely on medical-domain vision-encoder pre-training, clinical-term-guided alignment, or trainable pathological representation enhancement. Although effective, these approaches usually require additional training or model-specific adaptation and may overfit to particular disease morphologies, limiting their applicability to frozen medical VLMs. To address these limitations, we propose EasyLens, a training-free plug-and-play subtle-lesion representation amplifier for medical VLMs. EasyLens first constructs EasyBank, a pathology-anatomy prototype space that provides lesion-related prototypes and anatomy-aware normal references for comparing suspicious patches against both pathological and normal anatomical patterns. To avoid blindly amplifying normal tissues, EasyTag selects lesion-relevant patches through counterfactual prototype reasoning. To counteract the dilution of subtle lesion cues in global image representations, EasyAmplifier strengthens the selected lesion-relevant patch representations through morphology-guided residual enhancement, thereby increasing their contribution to the global image embedding. Experiments on multiple medical image datasets and frozen medical VLM backbones show that EasyLens improves subtle-lesion detection and outperforms existing encoder-enhancement baselines.

URL PDF HTML 收藏
2605.19665 2026-07-10 cs.SE cs.AI 版本更新

CriterAlign: Criterion-Centric Rationale Alignment for Code Preference Judging

CriterAlign: 以标准为中心的推理对齐用于代码偏好判断

Zhenyu Li, Aleksandar Cvejic, Zehui Chen, Peter Wonka

机构 * KAUST(卡塔尔人工智能研究 institute) ByteDance(字节跳动)

AI总结 本文提出CriterAlign,一种以标准为中心的推理对齐框架,通过直接的标准级 pairwise 判断、tie-driven 标准细化、swap-consistency 过滤和最终 pairwise 合成,改进了代码偏好判断的准确性,同时引入Human-Preference-Aligned Guidance (HPAG)来提升性能。

详情
AI中文摘要

成对的人类偏好预测是评估代码生成系统的核心,其中质量往往依赖于任务特定的权衡,而不仅仅是功能正确性。虽然基于评分表的LLM判断通过将评估分解为显式标准来提高可解释性,但大多数现有流程仍然是逐点的:它们独立评分每个响应,并通过比较聚合分数来推导偏好。我们证明这种设计与成对的代码偏好预测不匹配,并且可能在强单体判断下表现不佳。我们提出了CriterAlign,一种以标准为中心的框架,通过直接的标准级成对判断、tie驱动的标准细化、swap一致性过滤和最终成对合成,将基于评分表的判断适应于成对偏好评估。我们进一步引入Human-Preference-Aligned Guidance (HPAG),通过从训练示例中提取人类偏好与单体判断预测之间的反复推理缺口进行离线合成,并注入到标准生成器、标准判断器和最终判断器中。在BigCodeReward上,CriterAlign将Qwen2.5-VL-32B单体判断的准确率从60.4%提升到66.3%,消融实验确认了成对标准设计和HPAG的贡献。

英文摘要

Pairwise human preference prediction is central to evaluating code-generation systems, where quality often depends on task-specific trade-offs beyond functional correctness. While rubric-based LLM judges improve interpretability by decomposing evaluation into explicit criteria, most existing pipelines remain pointwise: they score each response independently and derive preferences by comparing aggregated scores. We show that this design is poorly matched to pairwise code preference prediction and can underperform a strong monolithic judge. We propose CriterAlign, a criterion-centric framework that adapts rubric-based judging to pairwise preference evaluation through direct criterion-level pairwise judgments, tie-driven criterion refinement, swap-consistency filtering, and final pairwise synthesis. We further introduce Human-Preference-Aligned Guidance (HPAG), synthesized offline from training examples by extracting recurring rationale gaps between human preferences and monolithic judge predictions, and injected into the criterion generator, criterion judge, and final judge. On BigCodeReward, CriterAlign improves a Qwen2.5-VL-32B monolithic judge from 60.4% to 66.3% accuracy, with ablations confirming the contributions of pairwise criterion design and HPAG.

URL PDF HTML 收藏
2604.10180 2026-07-10 cs.DC cs.LG 版本更新

Tessera: Unlocking Heterogeneous GPUs through Kernel-Granularity Disaggregation

Tessera:通过内核粒度解耦解锁异构GPU

Tiancheng Hu, Jin Qin, Zheng Wang, Junhao Hu, Yuzheng Wang, Lei Chen, Yizhou Shan, Mingxing Zhang, Ting Cao, Chunwei Xia, Huimin Cui, Tao Xie, Chenxi Wang

机构 * Peking University(北京大学) Key Lab of HCST (PKU), MOE(高可信软件技术教育部重点实验室(北京大学)) University of Chinese Academy of Sciences(中国科学院大学) Institute for AI Industry Research, Tsinghua University(清华大学人工智能研究院) Tsinghua University(清华大学) University of Leeds(利兹大学) ByteDance(字节跳动)

AI总结 Tessera通过内核粒度解耦提升异构GPU上的大模型推理性能和成本效率,利用离线分析与在线适应结合,实现计算与硬件能力的匹配。

详情
AI中文摘要

Tessera通过内核粒度解耦提升异构GPU上的大模型推理性能和成本效率,利用离线分析与在线适应结合,实现计算与硬件能力的匹配。

英文摘要

Disaggregation maps parts of an AI workload to different types of GPUs, offering a path to utilize modern heterogeneous GPU clusters. However, existing solutions operate at a coarse granularity and are tightly coupled to specific model architectures, leaving much room for performance improvement. This paper presents Tessera, the first kernel disaggregation system to improve performance and cost efficiency on heterogeneous GPUs for large model inference. Our key insight is that kernels within a single application exhibit diverse resource demands, making them the most suitable granularity for aligning computation with hardware capabilities. Tessera integrates offline analysis with online adaptation by extracting precise inter-kernel dependencies from PTX to ensure correctness, overlapping communication with computation through a pipelined execution model, and employing workload-aware scheduling with lightweight runtime adaptation. Extensive evaluations across five heterogeneous GPUs and four model architectures, scaling up to 16 GPUs, show that Tessera improves serving throughput and cost efficiency by up to 2.3x and 1.6x, respectively, compared to existing disaggregation methods, while generalizing to model architectures where prior approaches do not apply. Surprisingly, a heterogeneous GPU pair under Tessera can even exceed the throughput of two homogeneous high-end GPUs at a lower cost.

URL PDF HTML 收藏
2607.06987 2026-07-09 cs.LG 新提交

UP: Unbounded Positive Asymmetric Optimization for Breaking the Exploration-Stability Dilemma

UP:用于打破探索-稳定性困境的无界正不对称优化

Chongyu Fan, Pengfei Liu, Jingjia Huang, Sijia Liu, Yi Lin

机构 * ByteDance Seed(字节跳动种子) Michigan State University(密歇根州立大学)

AI总结 研究针对强化学习探索-稳定性困境,提出无界正不对称优化(UP)方法,通过特殊设计重组优化过程,在多种算法、模型架构及训练模式下增强探索能力,实现卓越推理精度,是通用即插即用的强化学习训练增强方法。

详情
AI中文摘要

强化学习已成为增强大语言模型复杂推理能力的标准范式。现代强化学习框架依靠重要性采样来实现样本效率,但存在探索-稳定性困境。纯重要性采样常导致训练不稳定,标准裁剪机制虽缓解不稳定但严格限制策略更新预算。通过形式化概率容量概念,发现保守裁剪过早截断正确但低置信度推理路径的更新预算,抑制探索。提出无界正不对称优化(UP),通过停止梯度算子将策略锚定到当前状态,重组优化过程。不对称设计释放未裁剪稳定梯度以最大化探索,同时对负优势保持标准裁剪防止不稳定。实验表明UP增强探索能力,在多种算法、模型架构和训练模式下实现卓越推理精度。

英文摘要

Reinforcement learning (RL) has become the standard paradigm for enhancing the complex reasoning capabilities of large language models (LLMs). To achieve sample efficiency, modern RL frameworks rely on importance sampling (IS). However, these algorithms suffer from an exploration-stability dilemma. Pure IS often leads to catastrophic training instability, while standard clipping mechanisms used to mitigate this instability strictly constrain the policy update budget. By formalizing the concept of Probability Capacity (Cap), we reveal that conservative clipping structurally stifles exploration by prematurely truncating the update budget for correct but low-confidence reasoning paths. To break free from these constraints, we propose Unbounded Positive Asymmetric Optimization (UP), a universal and plug-and-play objective. UP theoretically restructures the optimization process by anchoring the policy to its current state via the stop-gradient operator. This asymmetric design unleashes unclipped, stable gradients for positive advantages to maximize exploration, while maintaining standard clipping safeguards for negative advantages to prevent training instability. Furthermore, our formulation readily extends across different optimization granularities, including token-level (GRPO, DAPO) and sequence-level (GSPO) frameworks. Extensive experiments demonstrate that UP enhances exploration capacity and achieves superior reasoning accuracy across diverse RL algorithms (DAPO, GSPO, and GRPO), model architectures (Dense, MoE, and vision-language), and training modalities (language and multimodal), validating UP as a truly universal plug-and-play enhancement for RL-based training.

URL PDF HTML 收藏
2607.05394 2026-07-09 cs.LG cs.AI cs.CL 新提交

Weak-to-Strong Generalization via Direct On-Policy Distillation

通过直接在线策略蒸馏实现弱到强泛化

Shiyuan Feng, Huan-ang Gao, Haohan Chi, Hanlin Wu, Zhilong Zhang, Zheng Jiang, Bingxiang He, Wei-Ying Ma, Ya-Qin Zhang, Hao Zhou

机构 * SIA-Lab of Tsinghua AIR and ByteDance Seed(清华-字节跳动联合研究中心SIA实验室) Institute for AI Industry Research (AIR), Tsinghua University(清华大学人工智能产业研究院) Department of Computer Science and Technology, Tsinghua University(清华大学计算机科学与技术系) Peking University(北京大学)

AI总结 针对可验证奖励RL在大模型上成本高的问题,提出Direct-OPD方法迁移小模型RL的策略偏移,无需在大模型上跑RL即可实现性能提升。

Comments Project Page: https://bytedtsinghua-sia.github.io/Direct-OPD/

详情
AI中文摘要

带可验证奖励的强化学习(RLVR)是提升语言模型推理能力的有效方案,但由于训练过程中目标模型需要生成大量轨迹,在每款新的强模型上重复执行该流程成本极高。随着模型规模扩大,后训练本身就成为了性能瓶颈。本文研究一种弱到强的替代方案:先在轨迹生成成本更低的小模型上运行RL,再复用该RL流程学到的知识来优化更强的目标模型。直接蒸馏经过RL训练的弱教师模型效果不佳,因为教师的最终策略同时混合了RL带来的有效增益和小模型自身的固有局限。本文提出直接在线策略蒸馏(Direct-OPD)方法,转而迁移教师模型由RL诱导的策略偏移。Direct-OPD将经过RL训练的教师模型与其自身未经过RL训练的参考模型做对比,把二者的对数似然比作为面向学生模型的稠密隐式奖励。通俗来说,这一对检查点可以标识出RL让弱模型更倾向或更不倾向采取的动作,而Direct-OPD会将该信号应用于更强学生模型自身的在线策略状态上。该方法可以直接复用弱模型的RL监督信号,无需训练显式奖励模型,也无需在目标模型上运行稀疏奖励RL。实验表明,Direct-OPD可以稳定地利用更弱的教师模型来提升更强的目标模型性能:值得注意的是,仅需在8张A100 GPU上运行4小时,它就能将Qwen3-1.7B在2024年AIME数据集上的准确率从48.3%提升至62.4%。该方法性能优于步数匹配的直接RL,还支持多轮策略偏移的顺序组合。本文结果表明,RL的产出可以作为隐式奖励信号跨模型规模复用,而不只是作为待模仿的最终模型。

英文摘要

Reinforcement learning with verifiable rewards (RLVR) is a powerful recipe for improving language-model reasoning, but it is expensive to repeat on every new strong model because the target model must generate many rollouts during training. As models scale, post-training itself becomes a bottleneck. We study a weak-to-strong alternative: run RL on a smaller model where rollouts are cheaper, then reuse what that RL run learned to improve a stronger target model. Directly distilling the post-RL weak teacher is not enough, because the teacher's final policy mixes useful RL gains with the limitations of the smaller model. We propose Direct On-Policy Distillation (Direct-OPD), which transfers the teacher's RL-induced policy shift instead. Direct-OPD compares the post-RL teacher with its own pre-RL reference and treats their log-ratio as a dense implicit reward for the student. In plain terms, the checkpoint pair tells us which actions RL made the weak model more or less likely to take, and Direct-OPD applies that signal on the stronger student's own on-policy states. This directly reuses the weak model's RL supervision signal without running sparse-reward RL on the target model. Empirically, Direct-OPD consistently leverages weaker teachers to improve stronger target models; notably, it boosts Qwen3-1.7B from 48.3% to 58.3% on AIME 2024 in just 4 hours on 8 A100 GPUs. It outperforms step-matched direct RL and enables the sequential composition of multiple policy shifts. Our results show that RL outcomes can be reused across model scales as implicit reward signals, not merely as final models to imitate.

URL PDF HTML 收藏
2606.27377 2026-07-08 cs.CV cs.CL cs.LG 新提交

DanceOPD: On-Policy Generative Field Distillation

DanceOPD:在策略生成场蒸馏

Wei Zhou, Xiongwei Zhu, Zelin Xu, Bo Dong, Lixue Gong, Yongyuan Liang, Meng Chu, Leigang Qu, Lingdong Kong, Wei Liu, Tat-Seng Chua

机构 * ByteDance Seed(字节跳动Seed) NUS(新加坡国立大学) UMD(马里兰大学) HKUST(香港科技大学)

AI总结 提出DanceOPD框架,通过将每个样本路由到能力场并查询低噪声学生诱导状态,用速度MSE损失训练流匹配模型,实现文本到图像、局部编辑和全局编辑的多能力组合,提升目标能力同时保持生成质量。

Comments Technical Report; 40 pages, 13 figures, 9 tables; Project Page at https://danceopd.github.io/ GitHub Repo at https://github.com/worldbench/DanceOPD

详情
AI中文摘要

现代图像生成需要一个统一多种能力的单一模型,包括文本到图像(T2I)、局部编辑和全局编辑。然而,这些能力很少自然对齐且经常冲突。例如,编辑往往会降低T2I性能,而全局和局部编辑相互干扰。因此,有效组合这些能力已成为图像生成模型训练的核心挑战。为了解决这个问题,我们引入了DanceOPD,一种用于流匹配模型的在策略生成场蒸馏框架,它将每个样本路由到一个能力场,查询一个低噪声学生诱导状态,并使用简单的速度MSE目标进行训练。每个能力源被定义为共享流状态空间上的速度场,学生从其自身展开状态上查询的场中学习以组合专家能力。该公式还吸收了算子定义的场,如无分类器引导。在T2I、编辑、真实感场吸收和CFG吸收上的全面实验表明,我们的方法改进了多能力组合,在保持锚定生成质量的同时增强了目标能力。我们相信这项工作为流匹配模型中的生成场蒸馏建立了一条实用途径。

英文摘要

Modern image generation demands a single model that unifies diverse capabilities, including text-to-image (T2I), local editing, and global editing. However, these capabilities are rarely naturally aligned and often conflict. For instance, editing tends to degrade T2I performance, while global and local editing interfere with each other. Consequently, effectively composing these capabilities has become a central challenge for image generation model training. To tackle this, we introduce DanceOPD, an on-policy generative field distillation framework for flow-matching models that routes each sample to one capability field, queries one low-noise student-induced state, and trains with a simple velocity MSE objective. With each capability source defined as a velocity field over the shared flow state space, the student learns from fields queried on its own rollout states to compose expert capabilities. This formulation also absorbs operator-defined fields such as classifier-free guidance. Comprehensive experiments on T2I, editing, realism-field absorption, and CFG absorption show that our approach improves multi-capability composition, strengthening target capabilities while preserving anchor generation quality. We believe this work establishes a practical route for generative field distillation in flow-matching models.

URL PDF HTML 收藏
2606.11032 2026-07-08 cs.CV 新提交

U-TTT: Towards Generalizable PET Image Denoising via Test-Time Training

U-TTT:通过测试时训练实现可泛化的PET图像去噪

Zhiwen Yang, Jiayin Li, Hao Lu, Hui Zhang, Zihua Wang, Yan Xu

机构 * School of Biological Science and Medical Engineering, Beihang University(北京航空航天大学生物科学与医学工程学院) Department of Biomedical Engineering, Tsinghua University(清华大学生物医学工程系) School of Aerospace Engineering, Tsinghua University(清华大学航天航空学院) ByteDance Inc.(字节跳动有限公司)

AI总结 针对PET图像去噪模型在分布偏移下性能退化的问题,提出U-TTT模型,集成测试时训练(TTT)层,通过自监督动态调整参数,并设计双域自适应机制(空间和频率TTT层),在未见剂量水平和扫描仪下实现最优去噪和泛化。

Comments This paper introduces the first TTT-based model for image restoration

详情
AI中文摘要

现有的用于正电子发射断层扫描(PET)图像去噪的深度学习模型在分布偏移下常常遭受严重的性能退化,从根本上限制了其稳健的临床部署。这种泛化能力的缺乏源于固定参数模型的传统范式,该范式在训练后无法适应测试数据的变化(例如,剂量水平或扫描仪类型)。为了克服这一限制并实现稳健的泛化,我们引入了U-TTT,一种新颖的U形模型,它集成了测试时训练(TTT)层,通过自监督在推理过程中动态调整模型参数,从而适应每个测试实例的特定特征。此外,为了全面捕捉3D PET数据的复杂退化,U-TTT具有双域自适应机制,包括空间测试时训练(S-TTT)层和频率测试时训练(F-TTT)层。S-TTT层捕捉并校正空间结构退化,而F-TTT层抑制全局噪声谱并恢复精细的高频细节。大量实验表明,U-TTT在PET去噪性能上达到了最先进水平,并在具有挑战性的分布偏移下(包括未见剂量水平和未见扫描仪)展现出优越的泛化能力。我们的代码将在此https URL提供。

英文摘要

Existing deep learning models for Positron Emission Tomography (PET) image denoising often suffer from severe performance degradation under distribution shifts, fundamentally restricting their robust clinical deployment. This lack of generalization stems from the conventional paradigm of fixed-parameter models that cannot adapt to variations in test data (e.g., dose levels or scanner types) after training. To overcome this limitation and achieve robust generalization, we introduce U-TTT, a novel U-shaped model that integrates Test-Time Training (TTT) layers to dynamically adjust model parameters during inference through self-supervision, thereby adapting to the specific characteristics of each test instance. Furthermore, to comprehensively capture the complex degradations of 3D PET data, U-TTT features a dual-domain adaptation mechanism comprising a Spatial Test-Time Training (S-TTT) layer and a Frequency Test-Time Training (F-TTT) layer. The S-TTT layer captures and corrects spatial structural degradations, while the F-TTT layer suppresses global noise spectra and restores delicate high-frequency details. Extensive experiments demonstrate that U-TTT achieves state-of-the-art PET denoising performance and exhibits superior generalization under challenging distribution shifts, including both unseen dose levels and unseen scanners. Our code will be available at https://github.com/Yaziwel/U-TTT.

URL PDF HTML 收藏
2604.13030 2026-07-08 cs.CV 版本更新

Generative Refinement Networks for Visual Synthesis

生成细化网络用于视觉合成

Jian Han, Jinlai Liu, Jiahuan Wang, Bingyue Peng, Zehuan Yuan

机构 * ByteDance(字节跳动)

AI总结 本文提出生成细化网络(GRN),通过近无损分层二进制量化解决离散化瓶颈,结合全局细化机制和熵引导采样策略,提升图像生成质量与效率。

Comments code: https://github.com/bytedance/GRN

详情
AI中文摘要

尽管扩散模型在视觉生成领域占据主导地位,但它们计算效率低下,无论复杂度如何都采用统一计算量。相比之下,自回归(AR)模型本质上具有复杂度意识,但常受限于损失性离散化和误差累积。本文引入生成细化网络(GRN),一种新一代视觉合成范式,以解决这些问题。其核心是通过理论近无损的分层二进制量化(HBQ)解决离散化瓶颈,实现与连续模型相当的重建质量。基于HBQ的潜在空间,GRN从根本上升级了AR生成,通过全局细化机制逐步完善和修正艺术作品——如同人类艺术家作画。此外,GRN整合了熵引导的采样策略,使生成过程具备复杂度意识和自适应步长,不牺牲视觉质量。在ImageNet基准上,GRN在图像重建(0.56 rFID)和类条件图像生成(1.81 gFID)方面建立了新纪录。我们还将GRN扩展到更具挑战性的文本到图像和文本到视频生成,实现了同等尺度上的优越性能。我们释放所有模型和代码以促进进一步的GRN研究。

英文摘要

While diffusion models dominate the field of visual generation, they are computationally inefficient, applying a uniform computational effort regardless of different complexity. In contrast, autoregressive (AR) models are inherently complexity-aware, as evidenced by their variable likelihoods, but are often hindered by lossy discrete tokenization and error accumulation. In this work, we introduce Generative Refinement Networks (GRN), a next-generation visual synthesis paradigm that addresses these issues. At its core, GRN addresses the discrete tokenization bottleneck through a theoretically near-lossless Hierarchical Binary Quantization (HBQ), achieving a reconstruction quality comparable to continuous counterparts. Built upon HBQ's latent space, GRN fundamentally upgrades AR generation with a global refinement mechanism that progressively perfects and corrects artworks -- like a human artist painting. Besides, GRN integrates an entropy-guided sampling strategy, enabling complexity-aware, adaptive-step generation without compromising visual quality. On the ImageNet benchmark, GRN establishes new records in image reconstruction (0.56 rFID) and class-conditional image generation (1.81 gFID). We also scale GRN to more challenging text-to-image and text-to-video generation, delivering superior performance on an equivalent scale. We release all models and code to foster further research on GRN.

URL PDF HTML 收藏
2112.02353 2026-07-08 cs.CV cs.LG 版本更新

Label Hierarchy Transition: Delving into Class Hierarchies to Enhance Deep Classifiers

标签层次转换:深入研究类层次结构以增强深度分类器

Renzhen Wang, De cai, Kaiwen Xiao, Xixi Jia, Xiao Han, Deyu Meng

机构 * School of Mathematics and Statistics and Ministry of Education Key Lab of Intelligent Networks and Network Security, Xi’an Jiaotong University(数学与统计学学院和教育部智能网络与网络安全重点实验室,西安交通大学) Tencent AI Lab(腾讯AI实验室) ByteDance(字节跳动) SINGPATH AI Lab, SingPath Medical Technology Pte. Ltd.(SingPath AI实验室,SingPath医疗科技私人有限公司) School of Mathematics and Statistics, Xidian University(数学与统计学学院,西安电子科技大学) College of Biomedical Engineering, Sichuan University(生物医学工程学院,四川大学) School of Mathematics and Statistics, Ministry of Education Key Lab of Intelligent Networks and Network Security, Xi’an Jiaotong University(数学与统计学学院和教育部智能网络与网络安全重点实验室,西安交通大学) Macao Institute of Systems Engineering, Macau University of Science and Technology(澳门系统工程研究院,澳门科技大学)

AI总结 研究针对层次分类中现有方法未充分利用类别相关性的问题,提出基于深度学习的LHT统一概率框架,由转换网络和混淆损失构成,实验证明其优于现有方法,还扩展到皮肤病变诊断任务展现潜力。

详情
AI中文摘要

层次分类旨在将对象分类到层次化的类别结构中,如鸟类可按目、科、种的三级层次分类。现有方法常将其解耦为一系列多类分类任务,但这种多任务学习策略未能充分利用层次结构中不同级别各类别间的相关性。本文提出基于深度学习的统一概率框架标签层次转换(LHT)来应对层次分类挑战。LHT框架由转换网络和混淆损失组成,转换网络专注于显式学习标签层次转换矩阵,可有效编码类层次结构中的潜在相关性,混淆损失促使分类网络在训练中学习不同标签层次间的相关性。该框架只需少量修改就能适应任何现有深度网络。通过一系列公共基准数据集进行层次分类问题实验,结果表明该方法优于当前最先进方法。此外,还将LHT框架扩展到皮肤病变诊断任务,验证了其在计算机辅助诊断中的巨大潜力。方法代码可在指定链接获取。

英文摘要

Hierarchical classification aims to sort the object into a hierarchical structure of categories. For example, a bird can be categorized according to a three-level hierarchy of order, family, and species. Existing methods commonly address hierarchical classification by decoupling it into a series of multi-class classification tasks. However, such a multi-task learning strategy fails to fully exploit the correlation among various categories across different levels of the hierarchy. In this paper, we propose Label Hierarchy Transition (LHT), a unified probabilistic framework based on deep learning, to address the challenges of hierarchical classification. The LHT framework consists of a transition network and a confusion loss. The transition network focuses on explicitly learning the label hierarchy transition matrices, which has the potential to effectively encode the underlying correlations embedded within class hierarchies. The confusion loss encourages the classification network to learn correlations across different label hierarchies during training. The proposed framework can be readily adapted to any existing deep network with only minor modifications. We experiment with a series of public benchmark datasets for hierarchical classification problems, and the results demonstrate the superiority of our approach beyond current state-of-the-art methods. Furthermore, we extend our proposed LHT framework to the skin lesion diagnosis task and validate its great potential in computer-aided diagnosis. The code of our method is available at \href{https://github.com/renzhenwang/label-hierarchy-transition}{https://github.com/renzhenwang/label-hierarchy-transition}.

URL PDF HTML 收藏
2607.05155 2026-07-07 cs.CL cs.LG 新提交

EdgeBench: Unveiling Scaling Laws of Learning from Real-World Environments

EdgeBench:揭示从真实环境中学习的缩放定律

Deyao Zhu, Xin Zhou, Shengling Qin, Xuekai Zhu, Hangliang Ding, Shu Zhong, Zixin Wen, Zhonglin Xie, Chenhui Gou, Linxuan Ren, Yueyang Wang, Junfeng Zhong, Rui Liu, Tian Gao, Yangguang Lin, Jingyuan Zhang, Maojia Song, Xuan Qi, Jinhong Wu, Chenyang Zhang, Yinzhu Piao, Ziru Niu, Hongbin Lin, Lingxiang Meng, Peng Tang, Chengyao Tang, Shanyu Wu, Huanyu Zheng, Yu Liu, Liya Zhu, He Wang, Ming Ding, Ziyu Wan, Hao Liu, Sibo Wang, Haotian Zhu, Xintian Zhang, Nan Chai, Yipeng Liu, Panhao Lai, Sihang Yuan, Zixin Su, Ge Zhang, Wangchunshu Zhou, Yantao Du, Wenhao Huang, Guang Shi

机构 * ByteDance(字节跳动)

AI总结 本文提出含134项长周期真实任务的EdgeBench平台,通过分析大量智能体交互数据,发现环境学习性能符合对数S型缩放定律,为智能体现实经验学习研究提供支撑。

详情
AI中文摘要

预训练缩放定律表明,模型能力会随数据和计算量增长呈现可预测的提升,但智能体部署后从真实环境中学习的过程仍远未被充分理解。通过分析智能体在134项真实任务中累计约38000小时的环境交互数据,我们发现了据我们所知的首个可证明环境学习的整体性能遵循对数S型缩放定律的证据,该规律的拟合精度极高,R²可达0.998。跨多代模型的观测还显示,智能体的学习速度大致每三个月提升一倍。上述发现依托EdgeBench实现,这是一套包含134项超长期限真实任务的评测套件,覆盖科学发现、软件工程、组合优化、专业知识工作、形式化数学和交互式游戏领域。每项任务都可在丰富的多层级反馈下支撑智能体至少12小时的连续运行,由大量专家投入构建完成。我们公开发布其中51项任务及完整评测框架,以加速智能体从真实经验中学习这一方向的研究。

英文摘要

Pretraining scaling laws reveal that model capability improves predictably with data and compute. But learning from real world environments after deployment remains far less understood. Analyzing roughly 38,000 hours of agent interaction with the environment across 134 real world tasks, we find, to the best of our knowledge, the first evidence that overall performance during environment learning follows a log-sigmoid scaling law with remarkably high precision, reaching R^2 = 0.998. Across model generations, we also find that agent learning speed roughly doubles every three months. This discovery stems from EdgeBench, a suite of 134 real world tasks with ultra-long horizons, spanning scientific discovery, software engineering, combinatorial optimization, professional knowledge work, formal mathematics, and interactive games. Each task sustains at least 12 hours of continuous agent operation under rich, multilevel feedback, and is built through substantial expert effort. We publicly release 51 tasks and our full evaluation framework to accelerate the study of how agents learn from real world experience.

URL PDF HTML 收藏
2607.04144 2026-07-07 cs.RO 新提交

Semantic-Guided Progressive Object Removal with Gaussian Splatting

基于高斯点云的语义引导渐进式物体移除

Xianliang Huang, Chen Xiao, Yuanxiang Ni, Guanming Liu, Mingkai Liu, Dikai Fan, Xiao Liu, Hao Zhang

机构 * PICO, ByteDance Inc.(字节跳动公司PICO) Shenzhen Institute of Advanced Technology, Chinese Academy of Sciences(中国科学院深圳先进技术研究院) Fudan University(复旦大学)

AI总结 提出结合语义引导块匹配与区域渐进式细化的框架用于3D物体移除。利用DINOv2编码语义指导,通过RPR策略分区域优化,基于高斯点云实现高效重建,在物体移除上性能优于现有高斯方法。

Comments 8 pages, 4 figures

详情
AI中文摘要

从重建的3D场景中移除不需要的物体是计算机视觉中的一项重要任务,支持AR/VR、机器人技术和数字内容创作等应用。现有方法通常在一步中完成整个掩码区域,且未有效利用其他视图的语义信息,导致处理复杂几何细节和纹理困难。在这项工作中,我们提出了一个新颖的框架,该框架集成了语义引导块匹配(SBM)和区域渐进式细化(RPR),用于高质量的3D物体移除。首先,我们利用DINOv2对多视图观察的语义引导进行编码,并对最佳匹配令牌进行解码,以完成目标视图中的缺失区域,同时保持跨视图一致性。其次,我们引入了一种RPR策略,将目标掩码分割成多个子区域,并选择性地细化那些视觉质量较差的区域。我们的方法基于高斯点云构建,确保了高效计算的高保真场景重建。实验结果表明,我们的方法在3D物体移除的感知质量和连贯性方面优于现有的基于高斯的方法。

英文摘要

Removing unwanted objects from reconstructed 3D scenes is an important task in computer vision, supporting applications in AR/VR, robotics, and digital content creation. Existing methods typically complete the entire masked region in a single step and without effectively utilizing semantic information from other views, leading to difficulties in handling complex geometric details and textures. In this work, we propose a novel framework that integrates Semantic-guided Block Matching (SBM) and Region-Wise Progressive Refinement (RPR) for high-quality 3D object removal. First, we leverage DINOv2 to encode semantic guidance from multi-view observations, and the best match tokens are decoded to complete missing regions in the target view while maintaining cross-view consistency. Second, we introduce a RPR strategy that segments the target mask into multiple subregions and selectively refines those with poor visual quality. Our method is built upon Gaussian Splatting, ensuring high-fidelity scene reconstruction with efficient computation. Experimental results demonstrate that our approach outperforms existing Gaussian-based methods in terms of perceptual quality and coherence in 3D object removal.

URL PDF HTML 收藏
2607.03162 2026-07-07 cs.AI cs.HC 新提交

APeB: Benchmarking Personalization Ability of Large Language Model Agents

APeB:大型语言模型智能体个性化能力基准测试

Garry Yang, Zizhe Chen, Xinru Chen, Yongqiang Chen, Jianxiang Wang, Deyu Zou, Linyi Ding, Jialiang Wu, Yunzhong He, Yu Gong, James Cheng, Huaixiao Tou

机构 * The Chinese University of Hong Kong(香港中文大学) ByteDance(字节跳动)

AI总结 研究大型语言模型智能体在处理原始、未明确查询时的个性化问题,通过引入个性化产品搜索测试平台构建基准测试,评估发现模型处理明确查询较好,但早期查询有困难,简单方法能提升性能。

Comments NA

详情
AI中文摘要

由大型语言模型驱动的智能体在用户发出原始、未明确的查询时,在个性化方面存在困难。在这种情况下,智能体必须推断潜在意图,从嘈杂的交互历史中提取偏好,并在相互竞争的选项中进行选择。现有基准很少测试这种能力……

英文摘要

LLM-powered agents struggle with personalization when users issue raw, underspecified queries. In this setting, agents must infer latent intent, extract preferences from noisy interaction histories, and select among competing alternatives. Existing benchmarks rarely test this capability, as they often rely on user-refined queries or simplified histories. We introduce personalized product search (PPS), a testbed for agentic personalization under raw queries and diverse histories. We construct Agent Personalized Benchmark (APeB) from action logs, pairing underspecified intents with rich histories and user-viewed candidate items. Evaluating state-of-the-art LLMs with multi-step agent workflows, we find that models handle explicit queries well but struggle with early-stage queries requiring intent and preference discovery. Rubric analysis attributes this gap mainly to ineffective history use. A simple history-aware query-refinement pipeline, VQRA, yields consistent gains, highlighting the need for dedicated history-utilization modules in personalized agents.

URL PDF HTML 收藏