arXivDaily arXiv每日学术速递 周一至周五更新

大厂专区

Salesforce

至 收录 174
2505.16839 2026-07-17 cs.CV 版本更新

LaViDa: A Large Diffusion Language Model for Multimodal Understanding

LaViDa:用于多模态理解的大型扩散语言模型

Shufan Li, Konstantinos Kallidromitis, Hritik Bansal, Akash Gokul, Yusuke Kato, Kazuki Kozuka, Jason Kuen, Zhe Lin, Kai-Wei Chang, Aditya Grover

机构 * UCLA(加州大学洛杉矶分校) Panasonic AI Research(松下人工智能研究) Adobe Research(Adobe研究) Salesforce Research(Salesforce研究)

AI总结 研究针对现有视觉语言模型在快速推理和可控生成方面的不足,提出基于离散扩散模型构建 LaViDa 系列视觉语言模型,采用多种新技术,在多模态基准测试中性能出色,成为自回归视觉语言模型的有力替代。

Comments 26 pages, 8 figures

详情
AI中文摘要

现代视觉语言模型(VLM)可解决各种需要视觉推理的任务。在现实场景中,VLM 理想特性包括快速推理和可控生成。现有自回归 VLM 在这些方面存在困难。离散扩散模型(DM)提供了有前景的替代方案,在多模态任务中的潜力未被充分探索。我们引入基于 DM 的 LaViDa 系列 VLM,通过为 DM 配备视觉编码器并联合微调以遵循多模态指令。LaViDa 采用了如互补掩码、前缀 KV 缓存和时间步长移位等新技术。实验表明,LaViDa 在多模态基准测试中性能优于或与自回归 VLM 竞争,具有速度 - 质量权衡灵活、可控性和双向推理等优势。

英文摘要

Modern Vision-Language Models (VLMs) can solve a wide range of tasks requiring visual reasoning. In real-world scenarios, desirable properties for VLMs include fast inference and controllable generation (e.g., constraining outputs to adhere to a desired format). However, existing autoregressive (AR) VLMs like LLaVA struggle in these aspects. Discrete diffusion models (DMs) offer a promising alternative, enabling parallel decoding for faster inference and bidirectional context for controllable generation through text-infilling. While effective in language-only settings, DMs' potential for multimodal tasks is underexplored. We introduce LaViDa, a family of VLMs built on DMs. We build LaViDa by equipping DMs with a vision encoder and jointly fine-tune the combined parts for multimodal instruction following. To address challenges encountered, LaViDa incorporates novel techniques such as complementary masking for effective training, prefix KV cache for efficient inference, and timestep shifting for high-quality sampling. Experiments show that LaViDa achieves competitive or superior performance to AR VLMs on multi-modal benchmarks such as MMMU, while offering unique advantages of DMs, including flexible speed-quality tradeoff, controllability, and bidirectional reasoning. On COCO captioning, LaViDa surpasses Open-LLaVa-Next-8B by +4.1 CIDEr with 1.92x speedup. On bidirectional tasks, it achieves +59% improvement on Constrained Poem Completion. These results demonstrate LaViDa as a strong alternative to AR VLMs. Code and models will be released in the camera-ready version.

URL PDF HTML 收藏
2607.13431 2026-07-16 cs.LG cs.AI cs.CL 新提交

Discrete Diffusion Models: A Unified Framework from Tokenization to Generation

离散扩散模型:从词元化到生成的统一框架

Ye Yuan, Weien Li, Rui Song, Zeyu Li, Haochen Liu, Xiangyu Kong, Zixuan Dong, Linfeng Du, Zipeng Sun, Weixu Zhang, Jiaxin Huang, Changjiang Han, Yonghan Yang, Zichen Zhao, Xiuyuan Hu, Haolun Wu, Yankai Chen, Fengran Mo, Jikun Kang, Bowei He, Philip S. Yu, Xue Liu

机构 * McGill University(麦吉尔大学) Mila - Quebec AI Institute(米拉-魁北克人工智能研究所) University of Cambridge(剑桥大学) University of Toronto(多伦多大学) MBZUAI - Mohamed bin Zayed University of Artificial Intelligence(穆罕默德·本·扎耶德人工智能大学) Tsinghua University(清华大学) Rochester Institute of Technology(罗彻斯特理工学院) Salesforce(Salesforce公司) University of Illinois Chicago(伊利诺伊大学芝加哥分校)

AI总结 研究离散扩散模型,引入统一框架从离散状态空间构建审视该模型,让现有公式成为共同设计空间实例,揭示训练、推理等方面权衡,为未来研究提供方向。

详情
AI中文摘要

离散去噪扩散模型(DDMs)最近成为离散数据自回归建模的有力替代方案,具有并行生成和迭代全局细化能力。与连续扩散不同,离散扩散模型的状态空间由离散状态空间的构建方式决定。本文引入统一概念框架,通过构建底层离散状态空间来审视离散扩散模型。在此框架下,现有公式成为共同设计空间的不同实例,还揭示了训练目标、推理算法等方面的常见权衡,为未来研究指明方向。

英文摘要

Discrete denoising diffusion models (DDMs) have recently emerged as a compelling alternative to autoregressive (AR) modeling for discrete data, offering parallel generation and iterative global refinement capabilities. Unlike continuous diffusion, where the state space is fixed, DDMs are fundamentally shaped by how the discrete state space is constructed: the tokenization scheme, the vocabulary topology, and domain-specific structural alphabets. This work introduces a unified conceptual framework that views discrete diffusion models through the construction of the underlying discrete state space. Within this framework, existing formulations, including transition-matrix, masking/absorbing-state, and score/ratio-based approaches, emerge as different instantiations of a common design space. The framework further exposes common design trade-offs across training objectives, inference algorithms, scaling behavior, systems optimization, and evaluation protocols, suggesting several promising directions for future research.

URL PDF HTML 收藏
2607.11862 2026-07-14 cs.CV cs.AI 新提交

Evidence-Backed Video Question Answering

有证据支持的视频问答

Shijie Wang, Honglu Zhou, Ziyang Wang, Ran Xu, Caiming Xiong, Silvio Savarese, Chen Sun, Juan Carlos Niebles

机构 * Salesforce, Palo Alto, CA, USA(Salesforce公司) Brown University, Providence, RI, USA(布朗大学)

AI总结 研究视频问答中模型缺乏可视化依据的问题,提出E-VQA任务及ST-Evidence基准,开发数据集ST-Evidence-Instruct,通过微调提高模型表现,为可解释的视频理解建立基线。

Journal ref ECCV 2026

详情
AI中文摘要

当前的视频大语言模型在问答方面表现出色,但大多像黑匣子一样运作,提供无可视化依据的文本答案。现有的可解释性方法依赖文本理由或稀疏边界框,难以捕捉复杂的视频动态。我们提出有证据支持的视频问答(E-VQA),要求模型联合输出语义答案和精确的时空证据。为此引入ST-Evidence基准,评估显示问答准确性和真正视觉感知之间存在关键解耦。我们开发数据集ST-Evidence-Instruct,在该数据上微调可提高模型表现,为可解释的视频理解建立了强大基线。

英文摘要

Current Video Large Language Models (Video LLMs) excel in question answering (QA) but largely operate as black boxes, providing textual answers without verifiable visual grounding. Existing explainability efforts rely on textual rationales or sparse bounding boxes, which struggle to capture complex video dynamics such as occlusions and non-rigid deformations. We propose Evidence-Backed Video Question Answering (E-VQA), a novel task requiring models to jointly output a semantic answer and precise spatio-temporal evidence: temporal segments and dense, tracked object segmentation masklets. To support this, we introduce ST-Evidence, the first human-verified benchmark for both discriminative and generative pixel-level grounding. Evaluations of state-of-the-art models reveal a critical decoupling between QA accuracy and true visual perception that scaling alone fails to bridge. To address this, we develop scalable, automated generation pipelines to create ST-Evidence-Instruct, a 160k-scale dataset bridging high-level reasoning with fine-grained grounding. Fine-tuning grounded Video LLMs on this data yields substantial gains over the corresponding size-matched UniPixel baselines (e.g., +27.2 t-mean and +13.8 J&F on a 7B model), establishing a robust baseline for explainable, evidence-backed video understanding. Code and data are available at https://github.com/SalesforceAIResearch/EVQA.

URL PDF HTML 收藏
2607.10806 2026-07-14 cs.CL cs.AI 新提交

Abstractiveness Metrics for Evaluating Text Summarization: A Refined Formulation with Empirical Validation

用于评估文本摘要的抽象性度量:一种经过实证验证的精确公式

Praveenkumar Katwe, Rakesh Chandra Balabantaray, Kali Prasad Vittala

机构 * International Institute of Information Technology, Bhubaneswar(布巴内斯瓦尔国际信息技术学院) Salesforce India Pvt Ltd(Salesforce印度私人有限公司)

AI总结 研究旨在量化文本摘要抽象性,引入RA、SA和AR度量,利用文档长度调和平均及非重叠因子公式,在四个模型上评估100个XSUM文档,成功区分提取式与抽象式模型,抽象率可识别需人工评估的摘要。

Comments 13 pages, 8 figures, code at https://github.com/katweNLP/AbstractionStudy. Extended and revised version of: Katwe et al., IEEE OCIT 2022 (doi:10.1109/OCIT56763.2022.00022)

详情
AI中文摘要

量化生成摘要中的抽象性对于评估超越像ROUGE这样的表面级度量的摘要模型至关重要。我们引入了参考抽象(RA)、摘要抽象(SA)和抽象率(AR)——一组有原则的启发式度量,用于衡量摘要与源文本的提取式复制的差异程度。该公式使用由三次非重叠因子调制的文档长度的调和平均值,产生维度一致、有界的输出,对提取式 - 抽象式边界具有非线性敏感性。在四个摘要模型(BART - large - cnn、Pegasus - xsum、DistilBart、MT5 - small)上对100个XSUM文档进行的评估表明,这些度量成功地区分了提取式模型(SA约为0.12 - 0.26)和抽象式模型(SA约为0.96 - 1.77),并且抽象率识别出需要人工评估是否存在潜在幻觉的摘要。代码和结果可在这个https网址获取。

英文摘要

Quantifying abstractiveness in generated summaries is essential for evaluating summarization models beyond surface-level metrics like ROUGE. We introduce Reference Abstraction (RA), Summary Abstraction (SA), and Abstraction Ratio (AR) -- a set of principled heuristic metrics that measure how much a summary diverges from extractive copying of the source text. The formulation uses the harmonic mean of document lengths modulated by a cubic non-overlap factor, yielding dimensionally consistent, bounded output with non-linear sensitivity to the extractive-abstractive boundary. Evaluation on 100 XSUM documents across four summarization models (BART-large-cnn, Pegasus-xsum, DistilBart, MT5-small) demonstrates that the metrics successfully discriminate between extractive models (SA ~ 0.12-0.26) and abstractive models (SA ~ 0.96-1.77), and that the Abstraction Ratio identifies summaries requiring manual evaluation for potential hallucination. Code and results are available at https://github.com/katweNLP/AbstractionStudy.

URL PDF HTML 收藏
2607.07985 2026-07-10 cs.CL cs.AI cs.SD eess.AS 新提交

A Reliability Assessment of LALM Audio Judges for Full-Duplex Voice Agents

全双工语音智能体的LALM音频评判的可靠性评估

A. Sayyad, J. Emmons, S. Jones, T. Lin, H. Krishnan

机构 * Salesforce Applied AI Research(Salesforce应用人工智能研究院)

AI总结 研究评估Gemini模型作为全双工语音智能体音频评判的可靠性,以Gemini 2.5 Flash为基准与人类评分者对比,测试多个维度,发现其在一些维度表现良好,不同模型有差异,确定部署需谨慎领域,为LALM部署提供实证依据。

Comments 28 pages total (12 main body, 1 reference, 15 appendix). In main body: 2 diagrams, 3 table, 2 charts

详情
AI中文摘要

我们报告了Gemini模型作为音频评判的实证可靠性,该模型能直接从原始立体声波形中对全双工智能体对话进行评分,测试了Gemini家族中的三个模型:2.5 Flash、3.5 Flash和3.1 Pro。主要证据基础使用Gemini 2.5 Flash作为基准模型,在209个立体声会话中与三位经过校准的人类评分者进行验证,在8个生产维度上评分:跨越13个口音和条件层次的152个全双工对话,以及57个对抗性缺陷注入片段。Gemini 2.5 Flash的证据在三项测试中是一致的。在8个维度中的5个维度上,LALM与人类的斯皮尔曼等级相关系数与人类之间的相关系数最多相差0.07,在8个维度中的7个维度上,这两个数量的95%自举置信区间重叠。在8个维度中的6个维度上,60%至92%的会话中LALM与三位评分者的人类平均值相差在1分以内。在48个(缺陷,维度)单元格中的45个单元格中,在纽科姆 - 威尔逊95%置信区间下,LALM与人类一样敏感或更敏感。排序能力在Gemini家族中具有转移性:3.5 Flash将简单一致性提高到8个维度中的8个,而3.1 Pro尽管等级相关性相当,但在几个维度上的评分明显低于人类。模型交换应专门在校准上重新验证,而不能仅从等级相关性假设。我们确定了四个部署需要谨慎的领域,并且估计仅人工评分我们当前的评估节奏成本比等效的LALM工作量大约高出两个数量级。这里呈现的数据为在证据支持的维度上部署LALM作为替代或第四评分者提供了可靠的实证基础。

英文摘要

We report the empirical reliability of Gemini models as audio judges that score full-duplex agent conversations directly from the raw stereo waveform, tested across three models in the Gemini family: 2.5 Flash, 3.5 Flash, and 3.1 Pro. Our primary evidence base uses Gemini 2.5 Flash as the ground-truth model, validated against three calibrated human raters on 209 stereo sessions, scored on 8 production dimensions: 152 full-duplex conversations across 13 accent-and-condition strata, together with 57 adversarial defect-injected clips. The evidence for Gemini 2.5 Flash is consistent across three tests. (i) On 5 of 8 dimensions the LALM-human Spearman rho departs from the pairwise human-human rho by at most 0.07, and on 7 of 8 dimensions the two quantities 95 percent bootstrap confidence intervals overlap. (ii) The LALM agrees with the three-rater human mean within 1 point on 60 to 92 percent of sessions on 6 of 8 dimensions. (iii) On 45 of 48 (defect, dimension) cells the LALM is as sensitive as humans or better under Newcombe-Wilson 95 percent confidence intervals, though most of these are underpowered nulls rather than demonstrated parity. Rank-ordering ability transfers across the Gemini family: 3.5 Flash improves simple agreement to 8 of 8 dimensions, while 3.1 Pro rates several dimensions markedly lower than humans despite comparable rank correlation. A model swap should be re-validated on calibration specifically, not assumed from rank-correlation alone. We identify four areas where deployment requires care, and we estimate that human rating alone for our current evaluation cadence costs roughly two orders of magnitude more than the equivalent LALM workload. The data presented here provides a defensible empirical basis for deploying the LALM as a substitute or fourth rater on the dimensions where the evidence supports it.

URL PDF HTML 收藏
2607.02032 2026-07-07 cs.AI cs.CL 新提交

PACE: A Proxy for Agentic Capability Evaluation

PACE:智能体能力评估的代理框架

Yueqi Song, Lintang Sutawika, Jiarui Liu, Lindia Tjuatja, Jiayi Geng, Yunze Xiao, Daniel Lee, Aditya Bharat Soni, Vincent Lo, Xiang Yue, Graham Neubig

机构 * Carnegie Mellon University(卡内基梅隆大学) Salesforce AI Research(Salesforce人工智能研究)

AI总结 提出PACE框架,通过从非智能体评测中选取原子实例构建代理基准,以低成本高精度预测智能体基准性能,在14个模型上实现MAE<4%、Spearman>0.80。

详情
AI中文摘要

在SWE-Bench和GAIA等基准上评估LLM智能体可能昂贵、耗时且需要复杂基础设施。单次评估可能花费数千美元并需要数天完成。相比之下,测试个体能力(如推理、代码生成)的非智能体LLM基准运行快速且成本低廉。本文研究是否可以通过在少量精心选择的原子评估实例上的表现来准确预测昂贵智能体基准的性能。我们引入PACE,一个通过从现有非智能体评估中选择实例来构建代理基准的框架,这些实例的聚合得分最可靠地预测模型在智能体基准上的表现。给定涵盖原子能力的候选实例池,PACE拟合一个回归模型,将模型在紧凑源实例子集上的得分映射到目标智能体基准得分。该子集本身通过结合两种互补的实例选择策略(目标相关局部选择和全局信息全局选择)来策划。我们将PACE应用于本文的4个目标智能体基准,得到PACE-Bench,即我们在论文中评估的具体代理基准。跨14个模型、4个智能体基准和19个非智能体基准的实验表明,PACE-Bench以远低于完整智能体评估成本1%的成本,预测智能体得分的留一交叉验证平均绝对误差低于4%,Spearman相关系数高于0.80,成对模型排名准确率约85%。我们进一步分析所选代理实例,揭示每个智能体基准独特要求的技能。PACE使从业者能够在模型开发、选择和路由过程中获得智能体性能的可靠估计,而无需进行完整的智能体评估。

英文摘要

Evaluating LLM agents on benchmarks like SWE-Bench and GAIA can be expensive, time-consuming, and requires complex infrastructure. A single evaluation can cost thousands of dollars and take days to complete. In contrast, non-agentic LLM benchmarks that test individual capabilities (e.g., reasoning, code generation) are fast and cheap to run. In this paper, we investigate whether performance on expensive agentic benchmarks can be accurately predicted by the performance on a small, carefully selected subset of atomic evaluation instances. We introduce PACE, a framework that constructs proxy benchmarks by selecting instances from existing non-agentic evaluations whose aggregate scores most reliably predict model performances on agentic benchmarks. Given a pool of candidate instances spanning atomic capabilities, PACE fits a regression that maps a model's scores on a compact subset of source instances to its score on the target agentic benchmark. The subset itself is curated by combining two complementary instance-selection strategies, target-relevance local selection and globally informative global selection. We apply PACE to the 4 target agentic benchmarks in this paper, which yields PACE-Bench, the concrete proxy benchmark that we evaluate in the paper. Experiments across 14 models, 4 agentic benchmarks, and 19 non-agentic benchmarks show that PACE-Bench predicts agentic scores with leave-one-out cross-validation (LOOCV) mean absolute error (MAE) under 4%, Spearman correlation above 0.80, and pairwise model-ranking accuracy around 85%, all at much less than 1% of the full agentic evaluation cost. We further analyze the selected proxy instances, revealing which skills each agentic benchmark uniquely demands. PACE enables practitioners to obtain reliable estimates of agentic performance during model development, selection, and routing, without the overhead of full agent evaluation.

URL PDF HTML 收藏
2607.01844 2026-07-03 cs.DC cs.AI 新提交

Mixture-of-Parallelisms: Towards Memory-Efficient Training Stack for Mixture-of-Experts Models

并行化混合:面向混合专家模型的内存高效训练栈

Xuan-Phi Nguyen, Shrey Pandit, Yiran Zhao, Semih Yavuz, Silvio Savarese, Shafiq Joty

机构 * Salesforce AI Research(Salesforce AI研究院)

AI总结 提出一种内存高效的MoE模型训练范式,结合多种并行技术,在有限硬件资源下实现万亿参数、百万上下文长度的无损训练,吞吐量比FSDP2基线提升4.7-8.2倍。

Comments Work in progress

详情
AI中文摘要

本文展示了一种针对混合专家(MoE)模型的内存高效训练栈。该训练范式在MoE模型训练管道的不同层和阶段组合并专门化各种现有和新型的并行技术。它利用这些技术,在CPU、CPU内存、GPU HBM内存以及GPU集群的CPU-GPU、GPU-GPU和节点间通信带宽的物理约束下,实现最大效率。它还包含一种新颖的优化器步骤策略,以实现高吞吐量和内存效率,使从业者能够在不到12个8x H200 GPU节点上,以百万上下文长度,对万亿参数规模的模型进行无损预训练/微调,并具有最先进的吞吐量和内存效率。在我们的实验中,MoP的每GPU吞吐量比强调优的FSDP2基线高出4.7倍至8.2倍(且差距随规模扩大而增大),并在长达100万token的上下文长度下维持训练,而基线在超过64-128K时内存耗尽。

英文摘要

This paper showcases a memory-efficient training stack for Mixture-of-Experts (MoE) models. It is a training paradigm that combines and specializes various existing and novel parallelism techniques at different layers and stages of the Mixture-of-Experts (MoE) model training pipeline. It leverages these techniques to achieve maximal efficiency given the physical constraints of CPU, CPU memory, GPU HBM memory, and the CPU-GPU, GPU-GPU, and node-node communication bandwidth of the GPU cluster. It also contains a novel strategy for the optimizer step to achieve high throughput and memory efficiency, enabling practitioners to conduct lossless pre-training/fine-tuning of trillion-parameter scale models, at a million context length, with just under 12 8x H200 GPU nodes, with state-of-the-art throughput and memory efficiency. In our experiments, MoP delivers 4.7x--8.2x higher per-GPU throughput than a strongly-tuned FSDP2 baseline (with the gap widening at larger scale) and sustains training at context lengths up to 1M tokens, where the baseline runs out of memory beyond 64--128K.

URL PDF HTML 收藏
2606.20676 2026-06-23 cs.CV cs.AI cs.CL 新提交

Jury Duty: Calibration and Orientation Failures in MLLM-as-a-Judge Under Cultural Ambiguity

陪审团职责:文化模糊性下MLLM作为评判者的校准与定向失败

Daniel Lee, Harsh Sharma, Eunkyu Park, Pranav Narayanan Venkit, Jeonghwan Kim, Kah Mun Chia, Andreas Vlachos, Shafiq Joty

机构 * Salesforce AI Research(Salesforce AI 研究院) University of Cambridge(剑桥大学) University of Colorado Boulder(科罗拉多大学博尔德分校) Carnegie Mellon University(卡内基梅隆大学) UIUC(伊利诺伊大学厄巴纳-香槟分校)

AI总结 针对MLLM作为评判者在跨文化场景中的偏差问题,提出VOIR DIRE基准,通过分析校准失败(压缩尺度)和定向失败(默认一种文化规范),揭示模型偏向于更宽容的文化解读。

Comments Under Review

详情
AI中文摘要

MLLM作为评判者通常通过与人类标注的一致性来验证,但当人类群体在文化上异质时,该指标无法定义。我们引入了VOIR DIRE,一个包含626个跨文化图像-提示对的多模态基准,覆盖美国和中国大陆的食品、时尚和建筑领域,标注者群体内部可靠(a=0.86/0.74),但跨群体评估存在分歧(Q1 r=-0.12)。在六个MLLM中,偏差分解为两种失败:正向性地板校准失败(压缩尺度使用)和定向失败(默认一种文化规范)。在这个语料库中,有争议的项目被采样以分裂两个群体,地板效应机械地验证了更宽容的中文解读;人物提示部分恢复了校准,但定向残余仍然存在,证明倾斜不能简化为尺度压缩。参考群体上下文演示加深了定向残余并抬高了高端,而不是恢复低端的使用。模型来源增加了约0.10 MAE的小附加倾斜,该倾斜在演示下近似不变。我们建议分别报告与每个参考群体的一致性,并将跨群体分歧视为评判者属性。

英文摘要

MLLM-as-a-Judge is conventionally validated by agreement with human annotations, but this metric is undefined when the human pool is culturally heterogeneous. We introduce VOIR DIRE, a multimodal benchmark of 626 culturally paired image--prompt artifacts spanning U.S. and mainland Chinese contexts across food, fashion, and architecture, with annotator pools that are within-pool reliable (a = 0.86/0.74) but cross-pool divergent on evaluation (Q1 r = -0.12). Across six MLLMs, the bias decomposes into two failures: a positivity-floor calibration failure (compressed scale use) and an orientation failure (default to one cultural norm). On this corpus, where contested items are sampled to split the two pools, the floor mechanically validates the more-permissive Chinese reading; persona prompting partially recovers calibration, but the orientation residual survives, evidence the tilt is not reducible to scale compression. Reference-pool in-context demonstrations deepen the orientation residual and inflate the high end rather than restoring use of the low end. Model origin adds a small additive tilt (~0.10 MAE) that is approximately invariant under demonstration. We recommend reporting alignment against each reference pool separately and treating cross-pool divergence as a judge property.

URL PDF HTML 收藏
2506.06952 2026-06-19 cs.CV 版本更新

LaTtE-Flow: Layerwise Timestep-Expert Flow-based Transformer

LaTtE-Flow: 基于层间时间步专家流的Transformer

Ying Shen, Zhiyang Xu, Jiuhai Chen, Shizhe Diao, Jiaxin Zhang, Yuguang Yao, Joy Rimchala, Ismini Lourentzou, Lifu Huang

机构 * University of Illinois Urbana-Champaign(伊利诺伊大学厄巴纳-香槟分校) University of Maryland(马里兰大学) Nvidia(英伟达) Salesforce AI Research(Salesforce AI研究) Intuit AI Research(Intuit AI研究)

AI总结 提出LaTtE-Flow,一种基于预训练视觉语言模型的高效统一架构,通过层间时间步专家流和条件残差注意力机制,实现图像理解与生成,生成速度提升约6倍。

Comments Unified multimodal model, Flow-matching

详情
AI中文摘要

多模态基础模型在统一图像理解与生成方面取得了最新进展,为在单一框架内处理广泛的视觉-语言任务开辟了令人兴奋的途径。尽管取得了进展,现有的统一模型通常需要大量的预训练,并且与专门针对每项任务的模型相比,难以达到相同的性能水平。此外,许多这些模型存在图像生成速度慢的问题,限制了它们在实时或资源受限环境中的实际部署。在这项工作中,我们提出了基于层间时间步专家流的Transformer(LaTtE-Flow),一种新颖且高效的架构,可在单个多模态模型中统一图像理解与生成。LaTtE-Flow建立在强大的预训练视觉语言模型(VLM)之上,以继承强大的多模态理解能力,并通过新颖的层间时间步专家流架构扩展它们,以实现高效的图像生成。LaTtE-Flow将流匹配过程分布到专门的Transformer层组中,每组负责不同的时间步子集。这种设计通过在每个采样时间步仅激活一小部分层,显著提高了采样效率。为了进一步提升性能,我们提出了一种时间步条件残差注意力机制,用于跨层高效的信息重用。实验表明,LaTtE-Flow在多模态理解任务上取得了强劲的性能,同时与最近的统一多模态模型相比,实现了具有竞争力的图像生成质量,推理速度提高了约6倍。

英文摘要

Recent advances in multimodal foundation models unifying image understanding and generation have opened exciting avenues for tackling a wide range of vision-language tasks within a single framework. Despite progress, existing unified models typically require extensive pretraining and struggle to achieve the same level of performance compared to models dedicated to each task. Additionally, many of these models suffer from slow image generation speeds, limiting their practical deployment in real-time or resource-constrained settings. In this work, we propose Layerwise Timestep-Expert Flow-based Transformer (LaTtE-Flow), a novel and efficient architecture that unifies image understanding and generation within a single multimodal model. LaTtE-Flow builds upon powerful pretrained Vision-Language Models (VLMs) to inherit strong multimodal understanding capabilities, and extends them with a novel Layerwise Timestep Experts flow-based architecture for efficient image generation. LaTtE-Flow distributes the flow-matching process across specialized groups of Transformer layers, each responsible for a distinct subset of timesteps. This design significantly improves sampling efficiency by activating only a small subset of layers at each sampling timestep. To further enhance performance, we propose a Timestep-Conditioned Residual Attention mechanism for efficient information reuse across layers. Experiments demonstrate that LaTtE-Flow achieves strong performance on multimodal understanding tasks, while achieving competitive image generation quality with around 6x faster inference speed compared to recent unified multimodal models.

URL PDF HTML 收藏
2602.03846 2026-06-17 cs.LG cs.AI 版本更新

PLATE: Plasticity-Tunable Efficient Adapters for Geometry-Aware Continual Learning

PLATE: 可塑性可调的几何感知持续学习高效适配器

Romain Cosentino

机构 * Salesforce AI Research(Salesforce人工智能研究)

AI总结 提出无需旧任务数据的持续学习方法PLATE,利用预训练网络的几何冗余性,通过结构化低秩更新显式控制可塑性-保留权衡,提升最坏情况保留保证。

详情
AI中文摘要

我们为预训练模型开发了一种持续学习方法,该方法不需要访问旧任务数据,解决了基础模型适应中预训练分布通常不可用的实际障碍。我们的关键观察是,预训练网络表现出大量的几何冗余性,并且这种冗余性可以通过两种互补的方式加以利用。首先,冗余神经元提供了预训练时代主导特征方向的代理,使得可以直接从预训练权重构建近似受保护的更新子空间。其次,冗余性为可塑性的放置位置提供了自然偏差:通过将更新限制在冗余神经元的子集并约束剩余的自由度,我们获得了在旧数据分布上功能漂移减少且最坏情况保留保证改善的更新族。这些见解导致了PLATE(可塑性可调的高效适配器),一种不需要过去任务数据的持续学习方法,它提供了对可塑性-保留权衡的显式控制。PLATE通过结构化低秩更新ΔW = B A Q^T参数化每一层,其中B和Q从预训练权重一次性计算并保持冻结,只有A在新任务上训练。代码可在https://this URL获取。

英文摘要

We develop a continual learning method for pretrained models that \emph{requires no access to old-task data}, addressing a practical barrier in foundation model adaptation where pretraining distributions are often unavailable. Our key observation is that pretrained networks exhibit substantial \emph{geometric redundancy}, and that this redundancy can be exploited in two complementary ways. First, redundant neurons provide a proxy for dominant pretraining-era feature directions, enabling the construction of approximately protected update subspaces directly from pretrained weights. Second, redundancy offers a natural bias for \emph{where} to place plasticity: by restricting updates to a subset of redundant neurons and constraining the remaining degrees of freedom, we obtain update families with reduced functional drift on the old-data distribution and improved worst-case retention guarantees. These insights lead to \textsc{PLATE} (\textbf{Pla}sticity-\textbf{T}unable \textbf{E}fficient Adapters), a continual learning method requiring no past-task data that provides explicit control over the plasticity-retention trade-off. PLATE parameterizes each layer with a structured low-rank update $ΔW = B A Q^\top$, where $B$ and $Q$ are computed once from pretrained weights and kept frozen, and only $A$ is trained on the new task. The code is available at https://github.com/SalesforceAIResearch/PLATE.

URL PDF HTML 收藏
2606.14958 2026-06-16 cs.CV cs.IR cs.LG 新提交

MVEB: Massive Video Embedding Benchmark

MVEB:大规模视频嵌入基准

Adnan El Assadi, Roman Solomatin, Isaac Chung, Chenghao Xiao, Deep Shah, Manan Dey, Shriya Sudhakar, Zacharie Bugaud, Wissam Siblini, Ayush Sunil Munot, Yashwanth Devavarapu, Rakshitha Ireddi, Michelle Yang, Márton Kardos, Niklas Muennighoff, Kenneth Enevoldsen

机构 * Harvard University(哈佛大学) SaluteDevices MIRAI Zendesk Shanghai University of Finance and Economics(上海财经大学) Google LLC Salesforce Cornell University(康奈尔大学) Astera Institute(Astera研究院) Independent Contributor(独立贡献者) Indian Institute of Technology, Kharagpur(印度理工学院,克拉格浦分校) Barclays(巴克莱银行) Aarhus University(奥胡斯大学) Stanford University(斯坦福大学)

AI总结 提出MVEB基准,包含23个任务评估33种视频嵌入模型,发现无单一模型占优,音频贡献取决于标注来源,并集成到MTEB生态。

详情
AI中文摘要

我们介绍了大规模视频嵌入基准(MVEB),这是一个包含23个任务的视频嵌入基准,涵盖分类、零样本分类、聚类、配对分类、检索和以视频为中心的问答。我们评估了33个模型,发现没有单一模型占优:基于MLLM的嵌入在分类、聚类、配对分类和问答上领先;多模态绑定在检索和零样本分类上领先;没有对比适应训练的生成式MLLM在跨模态任务上崩溃。成对的仅视频与音频+视频评估表明,音频的贡献取决于数据集标注来源:当标签来自两种模态时音频有帮助,当仅来自视觉时则有害,这一差距在模型族中一致为6个百分点。MVEB源自MVEB+(一个包含184个任务的任务池),旨在保持任务多样性的同时降低评估成本。它集成到MTEB生态系统中,以实现跨文本、图像、音频和视频的统一评估。我们在https://github.com/embeddings-benchmark/mteb上发布MVEB和所有184个任务,以及代码和排行榜。

英文摘要

We introduce the Massive Video Embedding Benchmark (MVEB), a 23-task benchmark for video embeddings spanning classification, zero-shot classification, clustering, pair classification, retrieval, and video-centric question answering. We evaluate 33 models and find that no single model dominates: MLLM-based embeddings lead on classification, clustering, pair classification, and QA; multimodal binding leads on retrieval and zero-shot classification; generative MLLMs without contrastive adaptation collapse on cross-modal tasks. Paired video-only vs. audio+video evaluations show that audio's contribution depends on dataset annotation provenance: audio helps when labels were produced from both modalities and hurts when they were produced from visuals alone, a six-point gap consistent across model families. MVEB is derived from MVEB+, a 184-task pool, and is designed to maintain task diversity while reducing evaluation cost. It integrates into the MTEB ecosystem for unified evaluation across text, image, audio, and video. We release MVEB and all 184 tasks along with code and a leaderboard at https://github.com/embeddings-benchmark/mteb.

URL PDF HTML 收藏
2606.13003 2026-06-16 cs.AI cs.CL cs.MA 新提交

The Illusion of Multi-Agent Advantage

多智能体优势的错觉

Prathyusha Jwalapuram, Hehai Lin, Chuyuan Li, Fangkai Jiao, Sudong Wang, Yifei Ming, Zixuan Ke, Chengwei Qin, Giuseppe Carenini, Shafiq Joty

机构 * Salesforce Research(Salesforce研究院) HKUST (Guangzhou)(香港科技大学(广州)) University of British Columbia(不列颠哥伦比亚大学) Nanyang Technological University(南洋理工大学)

AI总结 通过系统评估,发现自动生成的多智能体系统在性能和成本效率上均不如单智能体基线(如思维链自一致性),揭示了现有评估框架的缺陷和架构膨胀问题。

详情
AI中文摘要

普遍观点认为多智能体系统优于单智能体系统,其优势包括上下文保护、并行处理和分布式决策。然而,这一主张的经验支持主要依赖于与使用优先考虑孤立推理任务的基准测试的单智能体基线的比较,这些基准测试未能充分评估这些优势。我们专注于自动生成的多智能体系统(旨在比手动设计的系统具有更强的泛化能力),对单智能体系统(特别是思维链自一致性)进行了严格、系统的评估。在传统推理数据集和具有交互式多步骤工作流的任务(例如 BrowseComp-Plus)上,我们证明自动多智能体系统始终不如思维链自一致性,尽管其成本高达10倍。为了将这些失败与任务结构固有的局限性隔离开来,我们引入了一个为多智能体系统量身定制的诊断性合成数据集,该数据集具有显式任务分解、上下文分离和并行化潜力。我们表明,专家设计的多智能体系统在该数据集上的原始性能和成本效率方面始终优于自动生成的架构,这表明现有的评估框架未能考虑增加计算成本的边际效用,从而掩盖了复杂多智能体系统的关键架构缺陷和低效性。关键的是,对生成的多智能体系统架构的系统解构表明,当前的自动化设计范式产生了架构膨胀,优先考虑表面复杂性,但这并未转化为功能效用,暴露了与多智能体原则的根本性错位。

英文摘要

Prevailing wisdom posits that Multi-Agent Systems (MAS) are superior to Single-Agent Systems (SAS), citing advantages like context protection, parallel processing and distributed decision-making. However, empirical support for this claim relies primarily on comparisons with SAS baselines using benchmarks that prioritize isolated reasoning tasks, which do not adequately assess these advantages. Focusing on automatically generated MAS that are designed for enhanced generalizability over manually-designed counterparts, we perform a rigorous, systematic evaluation against SAS, specifically Chain-of-Thought with Self-Consistency (CoT-SC). Across traditional reasoning datasets and tasks with interactive multi-step workflows (e.g., BrowseComp-Plus), we demonstrate that automatic MAS consistently underperform CoT-SC despite being up to 10x more expensive. To isolate these failures from limitations inherent to task structure, we introduce a diagnostic synthetic dataset tailored for MAS featuring explicit task decomposition, context separation and parallelization potential. We show that expert-architected MAS consistently outperforms automatically generated architectures in both raw performance and cost-efficiency on this dataset, demonstrating that existing evaluation frameworks mask critical architectural gaps and inefficiencies of complex MAS by failing to account for the marginal utility of increased computational cost. Critically, systematic deconstruction of the generated MAS architectures reveals that current automated design paradigms produce architectural bloat that prioritizes superficial complexity which does not translate into functional utility, exposing a fundamental misalignment with multi-agent principles.

URL PDF HTML 收藏
2606.12918 2026-06-15 cs.CR cs.AI 新提交

MAStrike: Shapley-Guided Collusive Red-Teaming on Multi-Agent Systems

MAStrike: 基于Shapley值的多智能体系统合谋红队测试

Chejian Xu, Zhaorun Chen, Jingyang Zhang, Freddy Lecue, Avni Kothari, Sarah Tan, Wenbo Guo, Bo Li

机构 * University of Illinois, Urbana-Champaign(伊利诺伊大学厄巴纳-香槟分校) Virtue AI University of Chicago(芝加哥大学) Wells Fargo(摩根大通) Salesforce University of California, Santa Barbara(加州大学圣芭芭拉分校)

AI总结 提出MAStrike框架,通过Shapley值分析识别多智能体系统中脆弱智能体联盟,生成角色感知的对抗攻击,并迭代优化以绕过防御,显著优于启发式基线。

详情
AI中文摘要

分层多智能体系统(MAS)正迅速部署在金融和软件工程等高危工作流中。在这些系统中,安全本质上是分布在不同角色智能体上的,显著扩大了攻击面,特别是在特权提升和跨智能体合谋等协调对抗行为下。现有的MAS红队测试方法仍然有限:它们依赖启发式选择目标智能体并扰动孤立的消息流,留下了关键问题未解答,即哪些智能体对系统安全最负责,以及受损智能体如何协调以绕过防御。我们提出MAStrike,一个用于分层MAS中合谋红队测试的闭环框架。我们首次提出针对MAS的智能体级Shapley值分析,量化每个智能体在任务特定分布下对系统鲁棒性的边际贡献。在此归因指导下,MAStrike识别脆弱智能体联盟并生成协调的、角色感知的对抗操纵。这些攻击通过结构化因果诊断迭代优化,将失败案例归因于阻止对抗尝试的未受损智能体。我们进一步构建了全面的MAS红队测试基准和可控环境,涵盖不同的分层拓扑和领域,包括金融、软件工程和CRM。在多个前沿模型构建的MAS上进行的广泛实验表明,MAStrike显著优于启发式基线。我们的分析进一步揭示了智能体间非平凡的Shapley值分布和高阶交互结构,揭示了先前单智能体或基于模板的方法忽略的关键漏洞和协调模式。

英文摘要

Hierarchical multi-agent systems (MAS) are rapidly being deployed in high-stakes workflows across domains such as finance and software engineering. In these systems, safety and security are inherently distributed across role-specialized agents, significantly expanding the attack surface, particularly under coordinated adversarial behaviors such as privilege escalation and cross-agent collusion. Existing red-teaming approaches for MAS remain limited: they rely on heuristic selection of target agents and perturb isolated message streams, leaving critical questions unanswered as which agents are most responsible for system safety, and how compromised agents can coordinate to bypass defenses. We propose MAStrike, a closed-loop framework for collusive red-teaming in hierarchical MAS. We propose the first agent-level Shapley value analysis for MAS, quantifying each agent's marginal contribution to system robustness under task-specific distributions. GGuided by this attribution, MAStrike identifies vulnerable agent coalitions and generates coordinated, role-aware adversarial manipulations. These attacks are iteratively refined through structured causal diagnosis, attributing failure cases to uncompromised agents that block adversarial attempts. We further build a comprehensive MAS red-teaming benchmark and controllable environments spanning diverse hierarchical topologies and domains, including finance, software engineering, and CRM. Extensive experiments across MAS built on multiple frontier models show that MAStrike substantially outperforms heuristic baselines. Our analysis further uncovers non-trivial Shapley value distributions and higher-order interaction structures among agents, revealing critical vulnerabilities and coordination patterns that are overlooked by prior single-agent or template-based methods.

URL PDF HTML 收藏
2505.12992 2026-06-15 cs.LG cs.AI cs.CL stat.ML 版本更新

Fractured Chain-of-Thought Reasoning

断裂链式思维推理

Baohao Liao, Hanze Dong, Yuhui Xu, Doyen Sahoo, Christof Monz, Junnan Li, Caiming Xiong

机构 * University of Amsterdam(阿姆斯特丹大学) eBay Microsoft(微软) Google Research(谷歌研究) Salesforce

AI总结 提出断裂采样策略,通过截断推理链、调整轨迹数和解数,在推理时实现精度与成本的帕累托最优。

详情
AI中文摘要

推理时扩展技术通过在不重新训练的情况下利用额外的推理计算,显著增强了大型语言模型(LLMs)的推理能力。类似地,链式思维(CoT)提示及其扩展Long CoT通过生成丰富的中间推理轨迹来提高准确性,但这些方法会带来大量的token成本,阻碍了它们在延迟敏感场景中的部署。在这项工作中,我们首先证明截断CoT(即在完成推理前停止并直接生成最终答案)通常在使用显著更少token的情况下与完整CoT采样相匹配。基于这一见解,我们引入了断裂采样,这是一种统一的推理时策略,沿着三个正交轴在完整CoT和仅解决方案采样之间进行插值:(1)推理轨迹的数量,(2)每条轨迹的最终解数量,以及(3)推理轨迹被截断的深度。通过在五个不同的推理基准和多个模型规模上进行大量实验,我们证明断裂采样始终实现优越的精度-成本权衡,在Pass@k与token预算之间产生陡峭的对数线性缩放增益。我们的分析揭示了如何在这些维度上分配计算以最大化性能,为更高效和可扩展的LLM推理铺平了道路。代码可在该https URL获取。

英文摘要

Inference-time scaling techniques have significantly bolstered the reasoning capabilities of large language models (LLMs) by harnessing additional computational effort at inference without retraining. Similarly, Chain-of-Thought (CoT) prompting and its extension, Long CoT, improve accuracy by generating rich intermediate reasoning trajectories, but these approaches incur substantial token costs that impede their deployment in latency-sensitive settings. In this work, we first show that truncated CoT, which stops reasoning before completion and directly generates the final answer, often matches the full CoT sampling while using dramatically fewer tokens. Building on this insight, we introduce Fractured Sampling, a unified inference-time strategy that interpolates between full CoT and solution-only sampling along three orthogonal axes: (1) the number of reasoning trajectories, (2) the number of final solutions per trajectory, and (3) the depth at which reasoning traces are truncated. Through extensive experiments on five diverse reasoning benchmarks and several model scales, we demonstrate that Fractured Sampling consistently achieves superior accuracy-cost trade-offs, yielding steep log-linear scaling gains in Pass@k versus token budget. Our analysis reveals how to allocate computation across these dimensions to maximize performance, paving the way for more efficient and scalable LLM reasoning. Code is available at https://github.com/BaohaoLiao/frac-cot.

URL PDF HTML 收藏
2606.13598 2026-06-12 cs.AI cs.CL cs.LG cs.MA 新提交

Reward Modeling for Multi-Agent Orchestration

多智能体编排的奖励建模

King Yeung Tsang, Zihao Zhao, Vishal Venkataramani, Haizhou Shi, Zixuan Ke, Semih Yavuz, Shafiq Joty, Hao Wang

机构 * Rutgers University(罗杰斯大学) Salesforce AI Research(Salesforce人工智能研究)

AI总结 提出OrchRM框架,通过自监督学习从多智能体执行中间产物构建奖励模型,无需人工标注,实现高效编排器训练和测试时扩展,在多个领域提升性能并降低计算成本。

Comments Preprint; work in progress

详情
AI中文摘要

基于大型语言模型(LLM)的多智能体系统(MAS)需要有效的编排来协调专门化的智能体,然而训练这样的编排器受到有限监督和高计算成本的阻碍。我们提出了编排奖励建模(OrchRM),一种无需人工标注即可评估编排质量的自监督框架。OrchRM利用多智能体执行过程中的中间产物来构建Bradley-Terry奖励模型训练的胜负对。与现有的依赖昂贵子智能体展开的MAS测试时扩展和编排器训练框架不同,OrchRM直接在编排层面操作,实现了高效且高性能的奖励引导编排器训练和MAS测试时扩展。OrchRM在token使用上提高了高达10倍的训练效率,同时将MAS测试时扩展的准确率提升了高达8%。这些增益在多个领域(包括数学推理、基于网络的问答和多跳推理)中一致迁移,证明了编排级奖励建模作为鲁棒多智能体编排的可扩展方向。代码将在此https URL提供。

英文摘要

Multi-Agent Systems (MAS) built on Large Language Models (LLMs) require effective orchestration to coordinate specialized agents, yet training such orchestrators is hindered by limited supervision and high computational cost. We propose Orchestration Reward Modeling (OrchRM), a self-supervised framework for evaluating orchestration quality without human annotations. OrchRM leverages intermediate artifacts from multi-agent executions to construct win-lose pairs for Bradley-Terry reward model training. Unlike existing MAS test-time scaling and orchestrator training frameworks that rely on costly sub-agent rollouts, OrchRM operates directly at the orchestration level, enabling efficient and high-performing reward-guided orchestrator training and MAS test-time scaling. OrchRM improves training efficiency by up to 10x in token usage while improving MAS test-time scaling performance by up to 8% in accuracy. These gains consistently transfer across multiple domains, including mathematical reasoning, web-based question answering, and multi-hop reasoning, demonstrating orchestration-level reward modeling as a scalable direction for robust multi-agent orchestration. Code will be available at https://github.com/Wang-ML-Lab/OrchRM.

URL PDF HTML 收藏
2606.11625 2026-06-11 cs.LG 新提交

TimeRouter: Efficient and Adaptive Routing of Time-Series Foundation Models

TimeRouter: 时间序列基础模型的高效自适应路由

Kanghui Ning, Yushan Jiang, Kashif Rasul, Anderson Schneider, Yuriy Nevmyvaka, Dongjin Song

机构 * University of Connecticut(康涅狄格大学) Salesforce AI Research JP Morgan AI Research(摩根大通人工智能研究院)

AI总结 提出TimeRouter框架,通过轻量判别路由、选择性门控和集成回退实现时间序列基础模型的自适应选择,无需LLM推理,在GIFT-EVAL榜单取得最优性能。

详情
AI中文摘要

时间序列基础模型(TSFMs)作为新兴智能时间序列系统中的预测专家越来越受到探索。然而,TSFMs表现出异质性归纳偏差,且没有单一模型能在所有预测场景中持续占优,使得专家选择成为关键挑战。现有系统通常将此决策委托给基于LLM的控制器,导致大量推理开销。我们提出TimeRouter,一种高效路由框架,通过轻量判别路由、选择性门控和集成回退,利用预训练TSFM池的经验互补性。具体而言,TimeRouter结合了学习路由头、选择性门控和集成回退,在推理时无需调用LLM即可实现自适应专家选择。TimeRouter在GIFT-EVAL榜单上取得了最先进性能,LB MASE为0.6765。除了基准性能,我们的消融研究为TSFM路由设计提供了经验见解,强调了池组成和选择性门控的重要性。综合来看,这些结果使TimeRouter成为未来基于基础模型池的智能时间序列系统的模块化轻量路由层。我们的代码见此链接。

英文摘要

Time-series foundation models (TSFMs) are increasingly explored as predictive experts within emerging agentic time-series systems. However, TSFMs exhibit heterogeneous inductive biases, and no single model consistently dominates across forecasting regimes, making expert selection a critical challenge. Existing systems often delegate this decision to LLM-based controllers, incurring substantial inference overhead. We present TimeRouter, an efficient routing framework that leverages empirical complementarity across a pool of pretrained TSFMs through lightweight discriminative routing, selective gating, and ensemble fallback. Concretely, TimeRouter combines a learned routing head, a selective gate, and an ensemble fallback, enabling adaptive expert selection without invoking an LLM at inference time. TimeRouter achieves state-of-the-art performance on the GIFT-EVAL leaderboard, with an LB MASE of 0.6765. Beyond benchmark performance, our ablation studies provide empirical insights into TSFM routing design, highlighting the importance of pool composition and selective gating. Taken together, these results position TimeRouter as a modular and lightweight routing layer for future agentic time-series systems built upon foundation-model pools. Our code is available at https://github.com/UConn-DSIS/TimeRouter.

URL PDF HTML 收藏
2606.10722 2026-06-10 cs.CL 新提交

Continual LLM Upcycling: A Predictor-Gated Bank-Wise Sparsity Training Recipe for Dense-to-Sparse LLMs

持续LLM升级:一种用于稠密到稀疏LLM的预测器门控银行级稀疏训练方案

Ruixuan Huang, Jinyuan Shi, Hantao Huang, Yifan Huang, Ziyi Guan, Hao Zeng, Ian En-Hsu Yen, Minghui Yu

机构 * Nanyang Technological University(南洋理工大学) Salesforce AI Huawei Noah's Ark Lab(华为诺亚方舟实验室)

AI总结 提出一种从稠密检查点构建通道稀疏大语言模型的持续训练方法,通过预测器门控稀疏SwiGLU FFN和银行级top-k规则实现4倍稀疏性,并修复长上下文失败模式。

详情
AI中文摘要

我们研究稠密到稀疏的持续训练,作为从稠密检查点构建通道稀疏大语言模型的一种方式。从Qwen2.5-8B稠密骨干网络开始,我们在32K上下文中继续训练,并在32K阶段引入预测器门控稀疏SwiGLU FFN。对于每个token和层,我们使用低秩预测器生成FFN通道路由logits。然后应用银行级top-k规则,在每个64通道的银行中保留16个通道,从而在FFN中间激活中实现4倍稀疏性。与事后稀疏推理方法不同,路由模块被放置在主要语言建模路径上,并在持续训练期间进行优化,使稠密模型能够升级为面向硬件的稀疏模型。我们报告了架构、训练方案、基准性能以及训练经验。我们还识别了RULER-CWE上的层局部长上下文失败模式,并提出了一种单层修复算法,显著改善了受影响长度范围内的性能。

英文摘要

We study dense-to-sparse continual training as a way to construct channel-sparse large language models from dense checkpoints. Starting from a Qwen2.5-8B dense backbone, we continue training at 32K context and introduce a predictor-gated sparse SwiGLU FFN in the 32K stage. For each token and layer, we use a low-rank predictor to produce FFN-channel routing logits. We then apply a bank-wise top-k rule to retain 16 channels in every 64-channel bank, yielding 4x sparsity in the FFN intermediate activation. Unlike post-hoc sparse inference methods, the routing module is placed on the main language modeling path and optimized during continual training, enabling the dense model to be upcycled into a hardware-oriented sparse model. We report the architecture, training recipe, benchmark performance, and training lessons. We also identify a layer-local long-context failure mode on RULER-CWE and propose a single-layer repair algorithm that substantially improves the affected length range.

URL PDF HTML 收藏
2606.06741 2026-06-08 cs.AI cs.CL cs.LG 新提交

OpenSkill: Open-World Self-Evolution for LLM Agents

OpenSkill: 面向LLM智能体的开放世界自我进化

Zhiling Yan, Dingjie Song, Hanrong Zhang, Wei Liang, Yuxuan Zhang, Yutong Dai, Lifang He, Philip S. Yu, Ran Xu, Xiang Li, Lichao Sun

机构 * Lehigh University(莱维大学) University of Illinois Chicago(伊利诺伊大学芝加哥分校) University of British Columbia(不列颠哥伦比亚大学) Vector Institute(向量研究所) Salesforce AI Research(Salesforce人工智能研究) Massachusetts General Hospital and Harvard Medical School(麻省总医院和哈佛医学院)

AI总结 提出OpenSkill框架,使智能体在无目标任务监督下,利用开放世界资源自举构建技能和验证信号,实现自我进化,在多个基准上取得最佳自动通过率。

Comments 20 pages, 4 figures and 8 tables. Code is avalable at https://github.com/OpenLAIR/OpenSkill

详情
AI中文摘要

自我进化智能体需要在部署后进行适应,但现有方法假设存在可用的学习循环,例如精心策划的技能、成功的轨迹或验证信号。真实的开放世界部署可能不提供这些,只提供一个任务提示。在这项工作中,我们研究开放世界自我进化,其中智能体必须从零开始构建其技能和自身的验证信号,使用开放世界资源但没有目标任务监督。我们提出OpenSkill,一个启动这个循环的框架:它从文档、代码库和网络中获取基础知识和验证锚点,将它们综合成可迁移的技能,并根据自建的虚拟任务(基于锚点而非目标答案)来优化这些技能。因此,开放世界既提供了要学习的知识,也提供了一个独立于监督的练习环境,目标任务监督保留用于最终评估。在三个基准和两个目标智能体上,OpenSkill在满足无监督约束的同时取得了最佳自动通过率。分析表明,其技能无需特定模型适应即可跨模型迁移,并且其自建验证器与真实结果一致,尽管从未访问过这些结果。

英文摘要

Self-evolving agents requires adaptation after deployment, but existing approaches assume a usable learning loop, such as curated skills, successful trajectories, or verifier signals. Real open-world deployments may provide none of these, offering only a task prompt. In this work, we study open-world self-evolution, where an agent must build both its skills and its own verification signals from scratch, using open-world resources but no target-task supervision. We propose OpenSkill, a framework that bootstraps this loop: it acquires grounded knowledge and verification anchors from documentation, repositories, and the web, synthesizes them into transferable skills, and refines those skills against self-built virtual tasks grounded in the anchors rather than in target answers. The open world thus supplies both the knowledge to be learned and a supervision-independent practice environment, with target-task supervision reserved for final evaluation. Across three benchmarks and two target agents, OpenSkill attains the best automated pass rate while satisfying the no-supervision constraint. Analysis shows its skills transfer across models without model-specific adaptation, and its self-built verifier aligns with ground-truth outcomes despite never accessing them.

URL PDF HTML 收藏
2606.05613 2026-06-05 cs.AI

Multilingual Fine-Tuning via Localized Gradient Conflict Resolution

通过局部梯度冲突解决的多语言微调

Long P. Hoang, Yiran Zhao, Wei Lu, Wenxuan Zhang

机构 * Singapore University of Technology and Design(新加坡科技设计大学) Salesforce AI Research(Salesforce人工智能研究) Nanyang Technological University(南洋理工大学)

AI总结 提出Bucket-Level MOO框架,将多语言微调重构为多目标优化问题,通过局部梯度冲突解决提升多语言性能。

详情
AI中文摘要

大型语言模型(LLMs)的快速发展已将跨语言多功能性确立为现代系统的定义特征。然而,微调这些模型经常引发跨语言的负面干扰。为了解决这个问题,我们将多语言微调重构为多目标优化(MOO)问题。具体来说,我们引入了Bucket-Level MOO,一个可扩展的分布式框架,它在参数桶上局部应用基于梯度的MOO算法。这使得冲突感知更新成为可能,而无需重建完整梯度向量的高昂通信开销。理论上,我们证明了这种局部解决自然地强制执行精炼帕累托平稳性,这是帕累托最优性的一个严格更紧的必要条件。实验上,Bucket-Level MOO通过驱动LLMs构建特定的语言维度来减轻干扰,提高了表示的可分离性。在四个基础LLM上的广泛实验表明,我们的方法在标准微调范式上显著提高了所见和未见的多语言性能。

英文摘要

The rapid evolution of Large Language Models (LLMs) has established cross-lingual versatility as a defining feature of modern systems. However, fine-tuning these models frequently induces negative interference across languages. To address this, we reformulate multilingual fine-tuning as a multi-objective optimization (MOO) problem. Specifically, we introduce Bucket-Level MOO, a scalable distributed framework that applies gradient-based MOO algorithms locally on parameter buckets. This enables conflict-aware updates without the prohibitive communication overhead of reconstructing full gradient vectors. Theoretically, we prove this localized resolution natively enforces Refined Pareto Stationarity, a strictly tighter necessary condition for Pareto optimality. Empirically, Bucket-Level MOO mitigates interference by driving LLMs to construct distinct language-specific dimensions, improving representational separability. Extensive experiments across four base LLMs demonstrate that our method significantly improves both seen and unseen multilingual performance over standard fine-tuning paradigms.

URL PDF HTML 收藏
2512.05774 2026-06-05 cs.CV cs.AI cs.CL

Active Video Perception: Iterative Evidence Seeking for Agentic Long Video Understanding

主动视频感知:用于代理长视频理解的迭代证据寻求

Ziyang Wang, Honglu Zhou, Shijie Wang, Junnan Li, Caiming Xiong, Silvio Savarese, Mohit Bansal, Michael S. Ryoo, Juan Carlos Niebles

机构 * Salesforce AI Research(Salesforce AI研究院) University of North Carolina at Chapel Hill(北卡罗来纳大学教堂山分校)

AI总结 本文提出了一种主动视频感知框架AVP,通过迭代计划-观察-反思过程,主动决定视频内容的观察目标和时间,以提高长视频理解的准确性和效率。

Comments Website: https://activevideoperception.github.io/

详情
AI中文摘要

长视频理解(LVU)具有挑战性,因为回答现实世界查询往往依赖于稀疏、时间分散的线索,这些线索隐藏在数小时的大部分冗余和无关内容中。尽管代理流程提高了视频推理能力,但现有框架依赖于查询无关的描述器来感知视频信息,这浪费了计算资源并模糊了细粒度的时间和空间信息。受主动感知理论的启发,我们主张LVU代理应主动决定观察什么、何时和在哪里观察,并持续评估当前观察是否足够回答查询。我们提出了主动视频感知(AVP),一种证据寻求框架,将视频视为交互环境,并直接从像素中获取紧凑、查询相关的证据。具体而言,AVP运行一个迭代的计划-观察-反思过程,使用MLLM代理。在每个轮次中,计划者提出有针对性的视频交互,观察者执行以提取时间戳证据,反思者评估证据对查询的充分性,要么终止并给出答案,要么触发进一步观察。在五个LVU基准测试中,AVP实现了最高整体准确率,有显著提升。值得注意的是,AVP在平均整体准确率上比最佳代理方法高出5.7%,同时仅需18.4%的推理时间和12.4%的输入令牌。

英文摘要

Long video understanding (LVU) is challenging because answering real-world queries often depends on sparse, temporally dispersed cues buried in hours of mostly redundant and irrelevant content. While agentic pipelines improve video reasoning capabilities, prevailing frameworks rely on a query-agnostic captioner to perceive video information, which wastes computation on irrelevant content and blurs fine-grained temporal and spatial information. Motivated by active perception theory, we argue that LVU agents should actively decide what, when, and where to observe, and continuously assess whether the current observation is sufficient to answer the query. We present Active Video Perception (AVP), an evidence-seeking framework that treats the video as an interactive environment and acquires compact, queryrelevant evidence directly from pixels. Concretely, AVP runs an iterative plan-observe-reflect process with MLLM agents. In each round, a planner proposes targeted video interactions, an observer executes them to extract time-stamped evidence, and a reflector evaluates the sufficiency of the evidence for the query, either halting with an answer or triggering further observation. Across five LVU benchmarks, AVP achieves highest overall accuracy with significant improvements. Notably, AVP outperforms the best agentic method by 5.7% in average overall accuracy while only requires 18.4% inference time and 12.4% input tokens.

URL PDF HTML 收藏
2510.22768 2026-06-05 cs.CL

Seeing is Believing? Evaluating Vision-Language Model Susceptibility in Agent-to-Agent Multimodal Persuasion

见多识广?评估面向Agent-to-Agent多模态说服的视觉语言模型易受性

Haoyi Qiu, Yilun Zhou, Pranav Narayanan Venkit, Kung-Hsiang Huang, Jiaxin Zhang, Nanyun Peng, Chien-Sheng Wu

机构 * University of California, Los Angeles(加州大学洛杉矶分校) Salesforce AI Research(Salesforce AI研究)

AI总结 本文研究了在多智能体多模态说服场景中,视觉语言模型对多模态内容的易受性,提出了MMPersuade框架和数据集,通过实验揭示了多模态输入在说服中的优势,以及说服对象的领域和格式依赖性,以及心理策略在不同上下文和模型架构下的效果差异。

详情
AI中文摘要

随着自主代理越来越多地互动,它们不可避免地试图互相影响。尽管先前在纯文本环境下研究了Agent-to-Agent (A2A) 说服的动力学,但视觉语言模型 (VLMs) 的兴起带来了更复杂的挑战:多模态内容传达了更丰富的信息,同时整合了微妙且难以检测的说服线索。为了研究这种易受性,我们提出了MMPersuade,一个统一的框架和数据集用于A2A多模态说服。我们建模了说服者代理(利用图像和心理策略)与说服对象VLM之间的互动。我们的基准涵盖商业、主观和行为,以及对抗性情境,并通过功能调用评估说服,以捕捉超出口头回应的行为变化。在六个VLM上的实验揭示了三个发现:(1)多模态输入在说服中始终优于纯文本说服,原始视觉信号在对抗性情境中独特地增加易受性,通过绕过文本激活的安全防御;(2)说服对象的易受性高度依赖于领域和格式,现实和社区风格的格式在商业情境中驱动易受性,而不同格式在对抗性情境中占主导地位;(3)心理策略的有效性取决于上下文和模型架构,更强大的模型抵抗良性说服,但在对抗性多模态输入下更易受攻击。我们的框架为构建更稳健和对齐的VLMs提供了基础,以在多代理环境中使用。

英文摘要

As autonomous agents increasingly interact, they inevitably attempt to influence one another. While prior work in text-only settings has explored the dynamics of Agent-to-Agent (A2A) persuasion, the rise of Vision-Language Models (VLMs) introduces a more complex challenge: multimodal content conveys richer information while integrating subtle, hard-to-detect persuasive cues. To study this vulnerability, we present MMPersuade, a unified framework and dataset for A2A multimodal persuasion. We model interactions between a persuader agent, which leverages images and psychological strategies, and a persuadee VLM. Our benchmark spans commercial, subjective and behavioral, and adversarial contexts, and evaluates persuasion via function-calling that capture behavioral shifts beyond verbal responses. Experiments on six VLMs reveal three findings: (1) multimodal inputs consistently outperform text-only persuasion, with raw visual signals uniquely increasing susceptibility in adversarial settings by bypassing text-activated safety defenses; (2) persuadee vulnerability is highly domain- and format-dependent, with realistic and community-style formats driving susceptibility in commercial settings while different formats dominate in adversarial ones; and (3) psychological strategy efficacy varies with context and model architecture, as more capable models resist benign persuasion yet become more susceptible under adversarial multimodal inputs. Our framework provides a foundation for building more robust and aligned VLMs in multi-agent environments.

URL PDF HTML 收藏
1905.08930 2026-06-04 math.NA cs.LG cs.NA math.PR math.ST stat.ML stat.TH

Heavy Hitters and Bernoulli Convolutions

重 hitters与伯努利卷积

Alexander Kushkuley

机构 * Salesforce/Demandware

AI总结 本文提出了一种简单的事件频率近似算法,该算法对事件时效性敏感。算法通过迭代更新类别点击分布,在标准n维单纯形上生成随机游走路径。在某些条件下,这种随机游走具有自相似性,并对应于有偏伯努利卷积。算法评估自然地导致对有偏(有限和无限)伯努利卷积矩的估计。

Comments 1) fixed some typos and a reference 2) expanded section 3

详情
AI中文摘要

提出了一种非常简单的事件频率近似算法,该算法对事件时效性敏感。该算法通过迭代更新类别点击分布,在标准n维单纯形上生成(路径)随机游走。在某些条件下,这种随机游走具有自相似性,并对应于有偏伯努利卷积。算法评估自然地导致对有偏(有限和无限)伯努利卷积矩的估计。

英文摘要

A very simple event frequency approximation algorithm that is sensitive to event timeliness is suggested. The algorithm iteratively updates categorical click-distribution, producing (path of) a random walk on a standard $n$-dimensional simplex. Under certain conditions, this random walk is self-similar and corresponds to a biased Bernoulli convolution. Algorithm evaluation naturally leads to estimation of moments of biased (finite and infinite) Bernoulli convolutions.

URL PDF HTML 收藏
2606.02798 2026-06-03 cs.AI

BehaviorBench: Modeling Real-World User Decisions from Behavioral Traces

BehaviorBench: 从行为轨迹建模真实用户决策

Liangwei Yang, Jielin Qiu, Zixiang Chen, Ming Zhu, Juntao Tan, Zhiwei Liu, Wenting Zhao, Zhujun Lan, Akshara Prabhakar, Silvio Savarese, Huan Wang, Shelby Heinecke

机构 * Salesforce AI Research(Salesforce AI研究院)

AI总结 提出 BehaviorBench 基准,利用真实世界行为轨迹(预测市场与链上记录)评估个性化决策建模,包含信念预测和交易预测两个任务层。

详情
AI中文摘要

许多决策支持场景需要系统适应个体用户,但针对该问题的评估数据仍然有限。现有的用户理解基准通常依赖模拟用户或模型生成的行为,尽管近期研究警告基于模型的模拟可能系统性地偏离人类行为。我们引入了 extsc{BehaviorBench},一个从真实世界行为轨迹评估个性化决策建模的基准。 extsc{BehaviorBench} 从观测到的公开预测市场和链上记录重建钱包级别的决策历史,并将其组织为两个互补的任务层:\emph{信念预测},预测用户在市场中最终的公开立场和置信度;以及\emph{交易预测},预测个体交易的方向和数量。在 2000 个评估钱包中,该基准包含 141,445 个信念实例和 1,485,972 个交易实例,并具有用于基于检索的评估的不相交支持池。我们在四种历史接口下评估前沿和开放权重生成模型:无个性化、直接近期历史、生成用户画像和检索支持钱包证据。个性化在信念预测上比交易预测更一致地提升性能,模型排名在不同任务层和指标间变化,不同的历史接口暴露了不同的失败模式。 extsc{BehaviorBench} 提供了一个评估设置,用于研究个性化方法是否能够利用真实世界行为证据而非仅依赖模拟用户。

英文摘要

Many decision-support settings require systems that adapt to individual users, but evaluation data for this problem remain limited. Existing benchmarks for user understanding often rely on simulated users or model-generated behavior, even though recent work cautions that model-based simulations can diverge systematically from human behavior. We introduce \textsc{BehaviorBench}, a benchmark for evaluating personalized decision modeling from real-world behavioral traces. \textsc{BehaviorBench} reconstructs wallet-level decision histories from observed public prediction-market and on-chain records, and organizes them into two complementary task layers: \emph{Belief prediction}, which predicts a user's final revealed stance and confidence in a market, and \emph{Trade prediction}, which predicts the direction and amount of individual transactions. Across 2,000 evaluation wallets, the benchmark contains 141,445 Belief instances and 1,485,972 Trade instances, with disjoint support pools for retrieval-based evaluation. We evaluate frontier and open-weight generative models under four history interfaces: no personalization, direct recent history, generated user profiles, and retrieved support-wallet evidence. Personalization improves Belief prediction more consistently than Trade prediction, model rankings change across task layers and metrics, and different history interfaces expose different failure modes. \textsc{BehaviorBench} provides an evaluation setting for studying whether personalized methods can use real-world behavioral evidence rather than simulated users alone.

URL PDF HTML 收藏
2605.30341 2026-05-29 cs.CV cs.AI

GPIC: A Giant Permissive Image Corpus for Visual Generation

GPIC:用于视觉生成的大型许可图像数据集

Keshigeyan Chandrasegaran, Kyle Sargent, Suchir Agarwal, Michael Jang, Michael Poli, Juan Carlos Niebles, Justin Johnson, Jiajun Wu, Li Fei-Fei

机构 * Stanford University(斯坦福大学) Radical Numerics University of Michigan(密歇根大学) Salesforce Research(Salesforce研究)

AI总结 提出GPIC,一个约28万亿像素的大型许可图像数据集,包含1亿训练样本,通过最先进的视觉语言模型标注,用于视觉生成建模研究。

Comments 25 pages; Dataset: https://huggingface.co/datasets/stanford-vision-lab/giant-permissive-image-corpus; Project website: https://gpic.stanford.edu

详情
AI中文摘要

研究视觉生成建模的可扩展方法需要大型、可访问且稳定的数据集。我们引入了GPIC,一个约28万亿像素的大型许可图像数据集。GPIC包含由最先进的视觉语言模型标注的多样化互联网图像,包括1亿训练样本、20万验证样本和100万测试样本。此外,所有GPIC图像均获得研究及商业用途的许可。GPIC经过安全过滤、去重,并集中托管在Hugging Face上。我们为GPIC上的生成建模提供了一个基准测试协议。最后,我们提供了GPIC上像素空间流匹配的参考基线。我们的数据集、基准和模型可在https://huggingface.co/datasets/stanford-vision-lab/gpic获取。评估工具包和代码可在https://gpic.stanford.edu获取。

英文摘要

Studying scalable methods for visual generative modeling requires large, accessible, and stable datasets. We introduce GPIC, a Giant Permissive Image Corpus of approximately 28 trillion pixels. GPIC comprises diverse internet images captioned by a state-of-the-art vision-language model, including 100M training, 200K validation, and 1M test examples. Moreover, all GPIC images are permissively licensed for both research and commercial use. GPIC is safety-filtered, deduplicated, and centrally hosted on Hugging Face. We provide a benchmarking protocol for generative modeling on GPIC. Finally, we provide a reference baseline for pixel-space flow matching on GPIC. Our dataset, benchmark, and models are available at https://huggingface.co/datasets/stanford-vision-lab/gpic. Evaluation toolkit and code are available at https://gpic.stanford.edu

URL PDF HTML 收藏
2605.29218 2026-05-29 cs.AI cs.CL

GTA: Generating Long-Horizon Tasks for Web Agents at Scale

GTA:大规模生成面向Web智能体的长程任务

Tenghao Huang, Kung-Hsiang Huang, Prafulla Kumar Choubey, Yilun Zhou, Muhao Chen, Jonathan May, Chien-Sheng Wu

机构 * University of Southern California(南加州大学) Salesforce AI Research(Salesforce人工智能研究) University of California, Davis(加州大学戴维斯分校)

AI总结 提出GTA框架,通过集成爬取、检索式种子生成、上下文内生成和自动质量控制,为Web智能体生成带可执行轨迹的真实长程任务,解决现有基准缺乏过程监督和可扩展性问题。

Comments Published at Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics

详情
AI中文摘要

Web智能体将语言模型与浏览和工具使用能力相结合,有望成为开放的Web助手。然而,进展日益受到缺乏可扩展的过程级监督的限制。现有基准大多为手动构建,仅提供粗略的起始-目标注释,缺乏中间轨迹,而最近的自动生成方法仍然昂贵、有偏且浅显。这些限制阻碍了对必须泛化到现实、多跳、跨页面任务的智能体进行可靠训练和评估。我们引入了一个可扩展的框架GTA,它集成了爬取、基于检索的种子生成、上下文内生成和自动质量控制,以生成与可执行轨迹配对的真实任务。该设计将爬取与生成解耦以提高效率,将任务基于站点图以强制组合性,并通过确定性重放和系统验证确保密集监督。我们在超过50个涵盖电子商务、政府、论坛和新闻的网站上实例化了该流程,并具有多语言和多跳覆盖。由此产生的基准揭示了显著的人机性能差距,并实现了详细的诊断。我们的贡献有三方面:(i)形式化多跳Web智能体任务生成,(ii)提出一个高效且经过验证的自动数据创建流程,以及(iii)发布一个具有可重复评估的动态基准。

英文摘要

Web agents, which couple language models with browsing and tool-use capabilities, show promise as open web assistants. Yet progress is increasingly limited by the lack of scalable, process-level supervision. Existing benchmarks are largely manually constructed, providing only coarse start-goal annotations without intermediate trajectories, while recent automatic generation efforts remain expensive, biased, and shallow. These limitations prevent reliable training and evaluation of agents that must generalize to realistic, multi-hop, cross-page tasks. We introduce a scalable framework, GTA, that integrates crawling, retrieval-based seeding, in-context generation, and automated quality control to produce realistic tasks paired with executable trajectories. This design decouples crawling from generation for greater efficiency, grounds tasks in the site graph to enforce compositionality, and ensures dense supervision through deterministic replays and systematic validation. We instantiate the pipeline on over 50 websites covering e-commerce, government, forums, and news, with multilingual and multi-hop coverage. The resulting benchmark reveals a significant human-agent performance gap and enables detailed diagnostics. Our contributions are three-fold: (i) formalizing multi-hop web-agent task generation, (ii) proposing an efficient and validated pipeline for automatic data creation, and (iii) releasing a dynamic benchmark with reproducible evaluation.

URL PDF HTML 收藏
2603.00309 2026-05-28 cs.AI cs.MA

DIG to Heal: Scaling General-purpose Agent Collaboration via Explainable Dynamic Decision Paths

DIG to Heal: 通过可解释的动态决策路径扩展通用智能体协作

Hanqing Yang, Hyungwoo Lee, Yuhang Yao, Zhiwei Liu, Kay Liu, Jingdi Chen, Carlee Joe-Wong

机构 * Carnegie Mellon University(卡内基梅隆大学) University of Arizona(亚利桑那大学) Zoom Salesforce Amazon(亚马逊)

AI总结 提出动态交互图(DIG)框架,将通用LLM智能体的涌现协作建模为时变因果网络,首次实现协作过程的可观察、可解释与实时纠错。

详情
AI中文摘要

日益流行的智能体AI范式有望利用多个通用大语言模型(LLM)智能体的能力协作完成复杂任务。尽管许多智能体AI系统通过预定义工作流或固定智能体角色来降低复杂性,但理想情况是支持真正自主的智能体,能够在多个交互智能体之间实现涌现协作。然而在实践中,这种非结构化交互常常导致冗余工作和级联故障,难以解释或纠正。在这项工作中,我们研究了由通用LLM智能体组成的多智能体系统,这些智能体通过涌现协作解决问题,而不依赖预定义角色、控制流或通信约束。我们引入了动态交互图(DIG),它将涌现协作捕获为智能体激活和交互的时变因果网络。DIG首次使涌现协作变得可观察和可解释,能够直接从智能体的协作路径中实时识别、解释和纠正协作引发的错误模式。因此,DIG填补了理解通用LLM智能体如何在真正智能体化的多智能体系统中共同解决问题的关键空白。项目网页见:https://happyeureka.github.io/dig。

英文摘要

The increasingly popular agentic AI paradigm promises to harness the power of multiple, general-purpose large language model (LLM) agents to collaboratively complete complex tasks. While many agentic AI systems reduce complexity through predefined workflows or fixed agent roles, the ideal is to support truly autonomous agents capable of emergent collaboration across many interacting agents. Yet in practice, such unstructured interactions often lead to redundant work and cascading failures that are difficult to interpret or correct. In this work, we study multi-agent systems composed of general-purpose LLM agents that solve problems through emergent collaboration, without relying on predefined roles, control flows, or communication constraints. We introduce the Dynamic Interaction Graph (DIG), which captures emergent collaboration as a time-evolving causal network of agent activations and interactions. DIG makes emergent collaboration observable and explainable for the first time, enabling real-time identification, explanation, and correction of collaboration-induced error patterns directly from agents' collaboration paths. Thus, DIG fills a critical gap in understanding how general LLM agents solve problems together in truly agentic multi-agent systems. The project webpage can be found at: https://happyeureka.github.io/dig.

URL PDF HTML 收藏
2605.27068 2026-05-27 cs.CL cs.AI cs.MA

QUACK: Questioning, Understanding, and Auditing Communicated Knowledge in Multimodal Social Deduction Agents

QUACK: 多模态社交推理智能体中的沟通知识质疑、理解与审计

Ye Yuan, Rui Song, Weien Li, Zeyu Li, Haochen Liu, Xiangyu Kong, Changjiang Han, Yonghan Yang, Zichen Zhao, Zixuan Dong, Fuyuan Lyu, Bowei He, Haolun Wu, Jikun Kang, Xue Liu

机构 * McGill University(麦吉尔大学) Mila - Quebec AI Institute(魁北克人工智能研究所) University of Cambridge(剑桥大学) MBZUAI - Mohamed bin Zayed University of Artificial Intelligence(MBZUAI - 摩苏尔·本·扎耶德人工智能大学) University of Toronto(多伦多大学) Salesforce

AI总结 提出QUACK框架,通过游戏结果、行为轨迹和话语一致性三级评估,自动审计多模态社交推理智能体语言与感知行为的一致性,发现最强智能体仍有15.1%的空间幻觉和过半无据指控。

详情
AI中文摘要

社交推理游戏已成为探测大型语言模型智能体推理、欺骗、协调和信念建模的热门测试平台。然而,大多数环境仅通过胜率等游戏结果评分,且主要局限于纯文本交互,难以判断智能体的语言是否真正基于其感知和行动,也难以识别其行为背后的失败模式。为填补这一空白,我们引入了QUACK,一个用于审计多模态社交推理中智能体语言基础的开源环境和评估框架。QUACK在三个层面评估智能体:游戏结果、行为轨迹和话语级一致性。其核心的陈述验证流水线从引擎日志重建每个智能体的真实轨迹,并对照检查每个讨论声明,自动标记空间幻觉、无据指控、欺骗崩溃和语言-行动不一致。在同质和跨模型对抗设置下评估三个前沿视觉语言模型,我们发现即使是最强的智能体,其可验证的空间声明中有15.1%是幻觉,且超过一半的指控缺乏有据证据。我们在https://github.com/AAAAA-Academia-Attractions/QUACK发布完整的引擎、评估框架、工具包和日志。

英文摘要

Social deduction games have become a popular testbed for probing reasoning, deception, coordination, and belief modeling in Large Language Model (LLM) agents. However, most environments are scored only by game outcomes such as win rates and largely remain to text-only interaction, making it difficult to tell whether an agent's language is actually grounded in what it perceived and did, or to identify the failure modes underlying its behavior. To address this gap, we introduce QUACK, an open-source environment and evaluation framework for auditing the grounding of agent language in multimodal social reasoning. QUACK evaluates agents at three levels: game outcomes, behavioral trajectories, and utterance-level consistency. Its core Statement Verification Pipeline reconstructs each agent's ground-truth trajectory from engine logs and checks every discussion claim against it, automatically flagging spatial hallucination, unsupported accusation, deception collapse, and language-action inconsistency. Evaluating three frontier VLMs in both homogeneous and cross-model adversarial settings, we find that even the strongest agent hallucinates 15.1% of its verifiable spatial claims and makes over half of its accusations without grounded evidence. We release the full engine, evaluation framework, toolkit, and logs at https://github.com/AAAAA-Academia-Attractions/QUACK.

URL PDF HTML 收藏
2501.18196 2026-05-26 cs.LG

GDformer: Going Beyond Subsequence Isolation for Multivariate Time Series Anomaly Detection

GDformer:超越子序列隔离的多变量时间序列异常检测

Qingxiang Liu, Xiaoliang Luo, Chenghao Liu, Sheng Sun, Di Yao, Lvchun Wang, Wei Yu, Yuxuan Liang

机构 * The Hong Kong University of Science and Technology (Guangzhou)(香港科学与技术大学(广州)) China Mobile (Jiangxi) Virtual Reality Technology Co., Ltd.(中国移动(江西)虚拟现实技术有限公司) Salesforce AI Research(Salesforce AI研究) Institute of Computing Technology, Chinese Academy of Sciences(中国科学院计算技术研究所)

AI总结 提出全局字典增强Transformer(GDformer),通过基于字典的交叉注意力机制学习整个序列中所有正常点的全局表示,并利用原型捕获正常点-全局相关权重分布,实现基于表示相似性的统一检测准则,在五个基准数据集上达到最先进性能。

详情
AI中文摘要

无监督的多变量时间序列异常检测是一项具有挑战性的任务,因为需要在不访问异常点的情况下推导出紧凑的检测标准。现有方法主要基于重构误差或关联分歧,两者都局限于有限视野的孤立子序列,难以提供统一的序列级标准。在本文中,我们提出了全局字典增强Transformer(GDformer),采用改进的基于字典的交叉注意力机制,以培养整个序列中所有正常点共享的全局表示。相应地,交叉注意力图反映了点与全局表示之间的相关权重,这自然导致了基于表示相似性的检测标准。为了促进更紧凑的检测边界,引入了原型来捕获正常点-全局相关权重的分布。GDformer在五个真实世界基准数据集上一致实现了最先进的无监督异常检测性能。进一步的实验验证了全局字典在不同数据集之间具有良好的可迁移性。

英文摘要

Unsupervised anomaly detection of multivariate time series is a challenging task, given the requirements of deriving a compact detection criterion without accessing the anomaly points. The existing methods are mainly based on reconstruction error or association divergence, which are both confined to isolated subsequences with limited horizons, hardly promising unified series-level criterion. In this paper, we propose the Global Dictionary-enhanced Transformer (GDformer) with a renovated dictionary-based cross attention mechanism to cultivate the global representations shared by all normal points in the entire series. Accordingly, the cross-attention maps reflect the correlation weights between the point and global representations, which naturally leads to the representation-wise similarity-based detection criterion. To foster more compact detection boundary, prototypes are introduced to capture the distribution of normal point-global correlation weights. GDformer consistently achieves state-of-the-art unsupervised anomaly detection performance on five real-world benchmark datasets. Further experiments validate the global dictionary has great transferability among various datasets.

URL PDF HTML 收藏
2605.23204 2026-05-25 cs.AI

AutoResearch AI: Towards AI-Powered Research Automation for Scientific Discovery

AutoResearch AI:迈向人工智能驱动的科研自动化以实现科学发现

Guiyao Tie, Jiawen Shi, Dingjie Song, Yixiao Huang, Ziji Sheng, Xueyang Zhou, Daizong Liu, Pan Zhou, Yongchao Chen, Ran Xu, Lifang He, Qingsong Wen, Manling Li, Cong Lu, Shuai Li, Pengtao Xie, Yixuan Yuan, Rui Meng, Lei Xing, Lichao Sun, Caiming Xiong, Philip S. Yu, Jianfeng Gao

机构 * Huazhong University of Science and Technology(华中科技大学) Lehigh University(莱斯大学) Tsinghua University(清华大学) Wuhan University(武汉大学) Salesforce Research(Salesforce研究) Squirrel AI Learning(Squirrel AI学习) Northwestern University(西北大学) Independent(独立) Shanghai Jiao Tong University(上海交通大学) University of California San Diego(加州大学圣地亚哥分校) Chinese University of Hong Kong(香港中文大学) University of Illinois Chicago(伊利诺伊大学香槟分校) Stanford University(斯坦福大学) Google Cloud AI Research(谷歌云AI研究) Recursive Superintelligence(递归超级智能) Microsoft Research(微软研究院)

AI总结 本文综述了AI驱动的科研工作流自动化(AutoResearch)的发展,分析了从任务级AI到工作流级研究自动化的转变,并提出了五个评估维度(新颖性、有效性、影响力、可靠性和溯源),指出自主性受领域条件限制。

Comments 49 pages, 12 figures, 10 tables

详情
AI中文摘要

科学研究正在被AI系统重塑,这些系统从孤立的辅助转向更长周期的工作流,涵盖文献基础、假设生成、实验、验证、报告和修订。这一转变标志着从面向科学的任务级AI向工作流级研究自动化的过渡。然而,当前系统仍然碎片化,在自主性、领域范围、执行环境、验证机制和人类监督方面存在差异,同时在证据保存、可重复性、弱方向拒绝、溯源追踪、跨领域鲁棒性和负责任的科学闭环方面仍面临挑战。本综述通过AutoResearch(定义为AI驱动的科学工作流自动化的演进谱系)审视这些发展。其中,Vibe Research表示人类引导的基于提示的辅助和人工验证执行区域,而新兴的AI主导系统协调发现循环的更大部分,但尚未实现稳健的自主性。我们分析了研究系统如何在工作流中重新分配控制、证据、执行、验证和问责,并围绕五个工作流条件组织该领域:文献与研究基础;假设形成与规划;实验与工具使用;反馈、验证与评审;报告与知识传播。我们进一步综合了AI科学家系统、混合主动协同研究框架、基准测试、领域部署和开源基础设施。最后,我们提出五个评估维度——新颖性、有效性、影响力、可靠性和溯源——并表明AutoResearch的自主性是领域条件化的,在结构化、可执行且快速可验证的环境中更为可信,但在具身、延迟、异构、伦理或机构问责的背景下则受限。

英文摘要

Scientific research is being reshaped by AI systems that move beyond isolated assistance toward longer-horizon workflows spanning literature grounding, hypothesis generation, experimentation, validation, reporting, and revision. This shift marks a transition from task-level AI for science to workflow-level research automation. Yet current systems remain fragmented, differing in autonomy, domain scope, execution environment, validation mechanism, and human oversight, while still struggling with evidence preservation, reproducibility, weak-direction rejection, provenance tracking, cross-domain robustness, and accountable scientific closure. This survey examines these developments through AutoResearch, defined as the developmental spectrum of AI-powered scientific workflow automation. Within it, Vibe Research denotes the human-steered region of prompt-based assistance and human-verified execution, whereas emerging AI-led systems coordinate larger portions of the discovery loop without achieving robust autonomy. We analyze how research systems redistribute control, evidence, execution, validation, and accountability across workflows and organize the field around five workflow conditions: literature and research grounding; hypothesis formation and planning; experimentation and tool use; feedback, validation, and review; and reporting and knowledge communication. We further synthesize AI scientist systems, mixed-initiative co-research frameworks, benchmarks, domain deployments, and open-source infrastructures. Finally, we propose five evaluation dimensions--novelty, validity, impact, reliability, and provenance--and show that AutoResearch autonomy is domain-conditioned, being more credible in structured, executable, and rapidly verifiable settings but limited in embodied, delayed, heterogeneous, ethical, or institutionally accountable contexts.

URL PDF HTML 收藏
2605.23108 2026-05-25 cs.SE cs.AI

Philosophical Dispositions as Behavioral Constraints for AI-Assisted Code Review: An Empirical Study

哲学倾向作为AI辅助代码评审的行为约束:一项实证研究

Kaushal Bansal

机构 * Salesforce, Inc.(Salesforce公司)

AI总结 提出一种通过哲学倾向约束AI评审者行为的系统,在7个代码库上评估,发现其与人类评审者46%一致,75%发现独特,且无假阳性。

详情
AI中文摘要

AI辅助代码评审工具通常作为通用的“专家评审者”代理运行,无论需要何种分析类型,都会产生同质化的发现。我们提出一个系统,通过哲学倾向——基于特定认识论传统(皮浪怀疑论、新正理逻辑、第欧根尼犬儒主义、儒家关系伦理)的连贯人格视角,将注意力引导到结构上不同类型的问题上——来约束AI评审者行为。每种倾向通过否定方式定义(即拒绝做什么),配备自我监控的失败模式(hamartia),并通过角色协议按顺序编排。我们在跨越5种编程语言(Python、Go、C++、Java、Terraform)、5个组织(2个企业、3个开源)和2个时间时代(AI前2020年、AI后2024-2026年)的7个代码库的50个合并拉取请求上评估该系统。该倾向系统与人类评审者达到46%的一致性(验证信号质量),以75%的比率识别出独特发现,并且在总共601个发现中,没有发现被作者判定为假阳性(未评估评分者间一致性,这仍是一个局限)。受控基线比较表明,51%的倾向发现是同一模型使用通用“专家评审者”提示不会产生的,这些独特发现针对结构、操作和逻辑问题,而非标准代码级别问题。初步跨模型验证(Claude Opus vs. GPT Codex 5.3-xhigh)在3个PR上显示100%的框架结构遵循度和39%的发现级别一致性,表明该框架在保持模型特定分析视角的同时提供了真正的行为约束。

英文摘要

AI-assisted code review tools typically operate as generic "expert reviewer" agents, producing homogeneous findings regardless of the analysis type needed. We present a system that constrains AI reviewer behavior through philosophical dispositions -- coherent personality lenses grounded in specific epistemological traditions (Pyrrhonist Skepticism, Navya-Ny=aya logic, Diogenes' Cynicism, Confucian relational ethics) that direct attention to structurally different types of issues. Each disposition is defined apophatically (by what it refuses to do), equipped with a self-monitoring failure mode (hamartia), and orchestrated in sequence by role protocols. We evaluate this system on 50 merged pull requests across 7 repositories spanning 5 programming languages (Python, Go, C++, Java, Terraform), 5 organizations (2 enterprise, 3 open-source), and 2 temporal eras (pre-AI 2020, post-AI 2024--2026). The disposition system achieves 46% convergence with human reviewers (validating signal quality), identifies unique findings at a 75% rate, and produces no findings judged false-positive by the author across 601 total findings (inter-rater agreement was not assessed and remains a limitation). A controlled baseline comparison demonstrates that 51% of disposition findings are not produced by the same model using generic "expert reviewer" prompting, and these unique findings target structural, operational, and logical concerns rather than standard code-level issues. Preliminary cross-model validation (Claude Opus vs.\ GPT Codex 5.3-xhigh) on 3 PRs shows 100% framework-structure adherence with 39% finding-level agreement, suggesting the framework provides real behavioral constraint while preserving model-specific analytical perspective.

URL PDF HTML 收藏