arXivDaily arXiv每日学术速递 周一至周五更新

高校专区

University of Washington(华盛顿大学)

至 收录 1178
2607.18218 2026-07-21 cs.CV cs.AI 新提交

GigaPath-Flash and GigaTIME-Flash: Efficient Pathology Foundation Models for Whole-Slide and Tumor Microenvironment Analysis

GigaPath-Flash和GigaTIME-Flash:用于全切片和肿瘤微环境分析的高效病理学基础模型

Naoto Usuyama, Jeya Maria Jose Valanarasu, Sicong Yao, Hanwen Xu, Jaspreet Bagga, Guanghui Qin, Robert E. Kramer, Cliff Wong, Soohee Lee, Hao Qiu, Theodore Zhengde Zhao, Racheli Ben Shimol, Angela Crabtree, Kevin Matlock, Eduardo Alejandro Lozano Garcia, Naiteek Sangani, Alberto Santamaria-Pang, Jason Entenmann, Alexandra Q. Bartlett, Bill J. Wright, Bernard A. Fox, Brian Piening, Sheng Zhang, Sheng Wang, Tristan Naumann, Carlo Bifulco, Hoifung Poon

机构 * Microsoft Research(微软研究院) Paul G. Allen School of Computer Science and Engineering, University of Washington(华盛顿大学保罗·G·艾伦计算机科学与工程学院) Providence Genomics(普罗维登斯基因组学公司) Earle A. Chiles Research Institute, Providence Cancer Institute(普罗维登斯癌症研究所厄尔·A·奇尔斯研究所) Providence Research Network(普罗维登斯研究网络)

AI总结 研究针对计算病理学中模型局限,提出GigaPath-Flash和GigaTIME-Flash模型用于全切片和肿瘤微环境分析。前者结合特定编码器,计算量少性能优;后者扩展架构预测肿瘤免疫微环境,速度快内存省,共同为相关领域提供开放许可模型及权重。

Comments Models: https://aka.ms/gigapath-flash (GigaPath-Flash) and https://aka.ms/gigatime-flash (GigaTIME-Flash)

详情
AI中文摘要

基础模型已成为计算病理学的驱动力,有潜力通过从大规模组织病理学数据中学习可转移表示来改变癌症诊断、预后和治疗选择。然而,大多数预训练模型仅在图像块级别运行,使用受限许可证且计算成本高,限制了大规模切片级临床和研究应用。本文介绍了GigaPath-Flash和GigaTIME-Flash,用于全切片病理学AI和空间蛋白质组学预测的高效模型。GigaPath-Flash结合了在大规模真实世界组织病理学数据上预训练的22M参数ViT-S块编码器和21M参数LongNet切片编码器,其紧凑块编码器从十亿参数GigaPath(ViT-g)教师模型中提炼而来。GigaPath-Flash以少50倍的计算量保留了GigaPath 97%的平均切片级性能。GigaTIME-Flash扩展此架构以直接从常规H&E图像预测肿瘤免疫微环境,在预测质量上超越了基于CNN的原始GigaTIME,速度快6倍且GPU内存使用少8倍。这些模型与GigaPath和GigaTIME一起形成了一个基于大规模真实世界临床数据预训练的、开放权重且遵循Apache-2.0许可的模型家族。通过发布所有模型和权重,为计算病理学、免疫肿瘤学和精准健康提供了可访问的构建模块。

英文摘要

Foundation models have emerged as a driving force in computational pathology, with the potential to transform cancer diagnosis, prognosis, and treatment selection by learning transferable representations from large-scale histopathology data. A growing landscape of pathology foundation models now spans diverse data sources, architectures, and downstream applications. However, most pretrained models operate only at the image-tile level, use restrictive licenses, and remain computationally expensive, limiting large-scale slide-level clinical and research use. Here, we introduce GigaPath-Flash and GigaTIME-Flash, efficient models for whole-slide pathology AI and spatial proteomics prediction. GigaPath-Flash combines a 22M-parameter ViT-S tile encoder with a 21M-parameter LongNet slide encoder, both pretrained on large-scale real-world histopathology data. Its compact tile encoder is distilled from the billion-parameter GigaPath (ViT-g) teacher and shared by both models. GigaPath-Flash retains 97% of GigaPath's average slide-level performance with 50x less compute. GigaTIME-Flash extends this backbone to predict the tumor immune microenvironment directly from routine H&E images. It surpasses the original CNN-based GigaTIME in prediction quality while running 6x faster and using 8x less GPU memory. Together with GigaPath and GigaTIME, these models form an open-weight, Apache-2.0-licensed family pretrained on large-scale real-world clinical data. By releasing all models and weights, we provide accessible building blocks for computational pathology, immuno-oncology, and precision health.

URL PDF HTML 收藏
2607.18088 2026-07-21 cs.LG cs.CV 新提交

The Label Complexity of Class-Conditional Coverage under Distribution Shift

分布偏移下类条件覆盖的标签复杂性

Weijia Han, Lisha Qu

机构 * University of Washington(华盛顿大学)

AI总结 研究分布偏移下类条件覆盖的标签复杂性,指出无标签方法难以兼顾有效性和效率,精确了恢复每类有效性的成本,通过骨架动作识别等案例研究表明相关现象在多模态基准上存在。

Comments 16 pages main text, 22-page supplement

详情
AI中文摘要

许多识别系统的标准评估中存在分布偏移,因为基准在训练和测试分割中设置了不相交的条件。在这种偏移下,分割共形预测使边际覆盖率接近标称水平,而每类覆盖率却悄然失效。我们刻画了恢复每类有效性的成本。首先,存在一个不可能的情况:一旦偏移同时作用于协变量和标签,目标类条件得分律就无法从源标签和未标记的目标样本中识别出来,所以没有无标签方法能同时实现有效且高效的每类覆盖率。其次,我们精确了成本:仅每类有效性只需每类少量目标标签,而要同时实现有效性和每类效率所需的标签数量随着效率容差的平方反比和类数的对数增长,且有匹配的上下界。第三,在评估的预测驱动推理家族中,即使在无界未标记目标池上最有利地使用分类器自己的伪标签,在覆盖率崩溃时效率最多提高一个小常数因子。骨架动作识别是我们的真实数据案例研究。仅使用源标签进行每类校准可在偏移保持边际覆盖率时恢复大部分每类差距,且在边际覆盖率本身崩溃时停止起作用。三种严重程度不断增加的真实偏移追踪了这个边界,并且在自然图像损坏基准上也出现了同样的崩溃和恢复,超越了任何单一模态。

英文摘要

Standard evaluation of many recognition systems contains distribution shift by construction, since benchmarks place disjoint conditions in the training and test splits. Under such a shift, split conformal prediction keeps marginal coverage near the nominal level while per-class coverage fails silently: on a real cross-subject skeleton benchmark, marginal coverage stays near ninety percent, the worst action class is covered about seventy percent of the time, and ten of the sixty classes fall below eighty percent coverage. We characterize the cost of restoring per-class validity. First, an impossibility: once the shift acts jointly on the covariates and the labels, the target class-conditional score law is unidentified from source labels and an unlabeled target sample, so no label-free method attains per-class coverage that is at once valid and efficient. Second, we make the cost precise: per-class validity alone needs only a handful of target labels per class, while the label count necessary and sufficient for validity together with per-class efficiency grows as the inverse square of the efficiency tolerance and the logarithm of the number of classes, with matching upper and lower bounds. Third, within the evaluated prediction-powered inference family, even the most favorable use of the classifier's own pseudo-labels on an unbounded unlabeled target pool improves efficiency by at most a small constant factor where coverage collapses. Skeleton action recognition is our real-data case study. A per-class calibration using source labels alone recovers a substantial share of the per-class gap while the shift preserves marginal coverage, and stops helping exactly when marginal coverage itself breaks. Three real shifts of increasing severity trace this boundary, and the same collapse and recovery appears on a natural-image corruption benchmark, beyond any single modality.

URL PDF HTML 收藏
2607.17479 2026-07-21 cs.CV 新提交

TraversRL: Traversable Pedestrian Pathway Generation With Reinforcement Learning

TraversRL:基于强化学习的可通行行人路径生成

Bin Han, Robert Wolfe, Bill Howe

机构 * University of Washington(华盛顿大学) Rutgers University(罗格斯大学)

AI总结 研究旨在从航拍图像生成行人路径,核心方法是用TraversRL视觉条件模型,通过特定动作空间和奖励机制迭代生长路径网络,主要贡献是相比基线提升交并比与连通性指标,结合奖励生成更优网络,证明该建模方法的有效性。

Comments Accepted to ECCV 2026, main conference

详情
AI中文摘要

从航拍图像自动生成行人路径需要构建适用于路线规划的连通网络,而非仅检测人行道位置。现有基于分割的方法常生成不可靠的导航图。我们引入TraversRL,一个视觉条件模型,从航拍图像迭代生长路径网络。它使用长短方向距离段的动作空间,结合图级和逐步奖励。在三个视觉骨干网络和三个交叉数据集上,TraversRL相对于分割基线大幅提高了缓冲交并比,连通性指标翻倍。结合全局和局部奖励能产生更优网络。结果表明将路径提取建模为旅行者视角的序列决策过程并用强化学习优化最终图质量,能生成更可靠的行人网络。

英文摘要

Automatically generating pedestrian pathways from aerial images requires producing a connected network suitable for routing, not just detecting where sidewalks appear. Sidewalks and crossings, in contrast to roads, may be partially occluded, implicitly defined, and exhibit complex connectivity patterns. Existing segmentation-based approaches focus on labeling pixels to infer segments, but often produce disconnected or fragmentary graphs that are unreliable for navigation. We introduce TraversRL, a vision-conditioned model that iteratively grows a pathway network from an aerial image, simulating a traveler navigating the built environment. TraversRL uses an action space of short and long direction-distance segments designed to adapt to complex patterns and span occlusions, and uses a combination of graph-level and step-wise rewards to balance overall connectivity with precise edge placement. Across three visual backbones and three intersection datasets, TraversRL substantially improves buffered IoU with the ground-truth graph relative to a state-of-the-art segmentation baseline, and more than doubles metrics of connectivity. Moreover, combining global and local rewards produces cleaner graphs with fewer spurious branches while further improving overall performance. These results demonstrate that modeling pathway extraction as a sequential decision process from the perspective of a traveler, while optimizing for final graph quality with reinforcement learning, produces significantly more reliable pedestrian networks.

URL PDF HTML 收藏
2607.16922 2026-07-21 cs.CV 新提交

Pedestrian Archetypes Extension -- More Pedestrian Models for Autonomous Vehicle Safety Testing

行人原型扩展——用于自动驾驶车辆安全测试的更多行人模型

Taorui Huang, Namita Gaidhani, Ritvik Bansal, S M Jubaer, Regina Lim, Rhett Zhao, Gavin Rafael Selin, Sunnie Deng Gao, Hasnain N Syed

机构 * Stanford University(斯坦福大学) University of California San Diego(加利福尼亚大学圣地亚哥分校) University of Washington(华盛顿大学)

AI总结 研究在行人原型基础上,通过注释YouTube行车记录仪视频,识别出7种新的行人原型,介绍其行为,阐述与旧原型差异并提供视频证据,为自动驾驶车辆安全测试提供更多行人模型。

Comments Extended version of Pedestrian Archetypes paper (published in IEEE IV 2025)

详情
AI中文摘要

在我们之前的工作《行人原型》中,我们将行人原型定义为唯一标识特定类型行人的行为集合。第一篇论文提出了12种行人原型,包括漫步者、醉酒者、分心者、闪行者、优柔寡断者、盲人、群体、乱穿马路者、老年人、儿童、多变者和停车行人。引入这些原型是为了超越单一行为标签,提供一种更自然的方式来描述危险行人在现实交通场景中实际逐步的行为方式。然而,在对YouTube行车记录仪视频进行进一步注释后,我们识别出7种额外的行人原型,它们与先前提出的原型存在明显的行为差异。这些新原型捕捉到了原始分类法无法完全解释的行人行为模式。在本预印本中,我们介绍每种新原型,定义其基本和可选行为,解释其与先前提出的原型的不同之处,并提供显示该原型实际应用的视频帧证据。

英文摘要

In our prior work, Pedestrian Archetypes, we defined pedestrian archetypes as collections of behaviors that uniquely identify a specific type of pedestrian. The first paper proposed 12 pedestrian archetypes, including the Wanderer, Drunk, Distracted, Flash, Indecisive, Blind, Flock, Jaywalker, Elderly, Kid, Eventful, and Parked Pedestrian. These archetypes were introduced to move beyond single behavior labels and provide a more natural way to describe how dangerous pedestrians actually behave progressively in real-world traffic scenarios. However, upon further annotation of YouTube dash-cam videos, we identified 7 additional pedestrian archetypes with observable and significant behavioral differences from the previously proposed ones. These new archetypes capture pedestrian behavior patterns that could not be fully explained by the original taxonomy. In this pre-print, we introduce each new archetype, define its essential and optional behaviors, explain how it differs from previously proposed archetypes, and provide video-frame evidence showing the archetype in action.

URL PDF HTML 收藏
2607.16324 2026-07-21 cs.CV cs.LG 新提交

SGMCE: Segment-Grounded Morphological Concept Explanation for Malaria Parasite Species Identification in Thick Blood Smears

SGMCE:厚血涂片疟原虫物种识别的基于片段的形态学概念解释

Ahmed Tahiru Issah, Charles B. Delahunt, Carine Mukamakuza

机构 * Carnegie Mellon University(卡内基梅隆大学) University of Washington(华盛顿大学)

AI总结 研究旨在解决厚血涂片疟原虫物种识别中深度学习缺乏形态学证据的问题,提出SGMCE框架,无需额外训练等,通过提取特征、查询GPT-4o生成自然语言解释,经多指标验证,在多种检测中取得较好结果。

Comments Accepted for publication in the 30th Conference on Medical Image Understanding and Analysis (MIUA 2026), Dublin. To appear in Springer Lecture Notes in Computer Science (LNCS)

详情
AI中文摘要

疟疾流行地区的疟疾诊断依赖于厚血涂片中疟原虫的物种水平识别,但深度学习检测器在分类检测时未提供预测的形态学证据,限制了显微镜检查人员在病例层面审核这些预测的能力。我们提出了SGMCE(基于片段的形态学概念解释),这是一个无需额外训练、形态学注释和标记解释数据的事后解释框架,它能产生基于厚涂片形态学的每个检测的自然语言解释。对于每个检测,SGMCE提取掩码引导的裁剪缩略图,使用自适应掩码内阈值计算14个手工制作的计算机视觉形态学特征,并根据从世界卫生组织辅助工具汇编的厚涂片特定知识库,用视觉证据和计算测量结果查询GPT-4o。主要输出是一个结构化解释,确定哪些形态学特征支持检测到的物种以及为什么排除竞争物种。通过四个自动指标对解释进行验证:知识库一致性(KBC)、计算机视觉声明忠实性(CCF)、区分度得分(DS)和大语言模型作为评判(LLMj)。一个带有物种感知否定过滤的句子级语义评分规则解决了临床散文和知识库术语之间的词汇不匹配问题。在跨越四种疟原虫物种和白细胞的139张厚涂片图像的737个检测中,寄生虫类的平均KBC为0.91,平均DS为0.99,平均CCF为0.97,而每个规则的CCF细分证实了视觉语言模型基于计算机视觉的声明与它们引用的测量结果一致。

英文摘要

Malaria diagnosis in endemic regions depends on species-level identification of Plasmodium parasites in thick blood smears, but deep learning detectors classify detections without providing morphological evidence for their predictions, limiting the ability of microscopists to audit those predictions at the case level. We present SGMCE (Segment-Grounded Morphological Concept Explanation), a post-hoc explanation framework that requires no additional training, no morphological annotations, and no labelled explanation data, yet produces per-detection natural-language explanations anchored in thick-smear morphology. For each detection, SGMCE extracts mask-guided crop thumbnails, computes fourteen handcrafted computer-vision morphological features (shape, colour, chromatin, haemozoin pigment) using adaptive within-mask thresholds, and queries GPT-4o with both visual evidence and computed measurements, conditioned on a thick-smear-specific knowledge base compiled from the World Health Organization bench aids. The primary output is a structured explanation identifying which morphological features support the detected species and why the competing species are excluded. Explanations are validated by four automatic metrics: Knowledge-Base Consistency (KBC), CV-Claim Faithfulness (CCF), Discriminativeness Score (DS), and LLM-as-Judge (LLMj). A sentence-level semantic scoring rule with species-aware negation filtering resolves the vocabulary mismatch between clinical prose and knowledge-base terms. Across 737 detections from 139 thick-smear images spanning four Plasmodium species and white blood cells, parasite-class mean KBC is 0.91, mean DS is 0.99, and mean CCF is 0.97, while a per-rule CCF breakdown confirms that the CV-grounded claims made by the vision-language model are consistent with the measurements they cite.

URL PDF HTML 收藏
2605.04344 2026-07-21 stat.ML cs.LG math.ST stat.TH 版本更新

Perturbation is All You Need for Extrapolating Language Models

扰动是语言模型外推所需的一切

Zetai Cen, Jin Zhu, Xinwei Shen, Chengchun Shi

机构 * School of Mathematics, University of Bristol(布里斯托大学数学系) School of Mathematics, University of Birmingham(伯明翰大学数学系) Department of Statistics, University of Washington(华盛顿大学统计系) Department of Statistics, London School of Economics and Political Science(伦敦政治经济学院统计系)

AI总结 本文针对大语言模型外推问题,提出基于扰动的方法,先转换前缀为语义邻域再进行下一个token预测,构建分层模型。通过建立五个属性发展外推性理论,经合成与真实数据评估,该方法提升支持域外预测性能,为语言建模提供实用路径。

Comments 59 pages

详情
AI中文摘要

本文通过预加性噪声模型重新诠释大语言模型,发展了一种外推统计理论。与基于精确前缀的标准自回归下一个token预测不同,我们引入基于扰动的过程,先将前缀转换为语义邻域,再基于此扰动变体进行下一个token预测,产生具有预加性噪声结构的分层模型。在此框架下,通过建立所提过程的适应性、收缩性、鲁棒性、外推性和双重鲁棒性五个属性,发展了严格的外推性理论。我们用合成和真实语言数据评估了所提过程的有限样本性能。结果表明,该方法持续改善支持域外预测,同时保持支持域内性能的竞争力,证明扰动为语言建模提供了实用途径。

英文摘要

This paper develops a statistical theory of extrapolation for large language models, by reinterpreting them through pre-post-additive noise models. In contrast to the standard autoregressive next-token prediction based on an exact prefix, we introduce a perturbation-based procedure that first transforms the prefix into a semantic neighbour and then conditions on this perturbed variant for next-token prediction. This yields a hierarchical model with a pre-post-additive noise structure. Within this framework, we develop a rigorous theory of extrapolability, namely, the capacity of a model class to make reliable predictions for token sequences that lie outside the empirical support of the training corpus, by establishing five properties of the proposed procedure: adaptivity, contractivity, robustness, extrapolability, and double robustness. We evaluate the finite sample performance of the proposed procedure using both synthetic and real world language data. Results show that the proposed method consistently improves out-of-support prediction while maintaining competitive in-support performance, demonstrating that perturbation offers a practical route to language modelling.

URL PDF HTML 收藏
2607.13718 2026-07-21 cs.CR cs.AI 版本更新

How Agents Ask for Permission: User Permissions for AI Agents, from Interfaces to Enforcement

智能体如何请求许可:人工智能智能体的用户权限,从接口到执行

Alexandra E. Michael, Franziska Roesner

机构 * University of Washington(华盛顿大学)

AI总结 研究人工智能智能体系统中用户级权限处理问题,通过调查21个提案构建分类法,分析五个商业智能体并与文献系统比较,确定主题与空白,为智能体权限系统研究提供参考。

Comments 15 pages, 4 figures

详情
AI中文摘要

随着人工智能智能体的普及,用户越来越多地面临此类系统带来的风险。提示注入攻击和幻觉可能导致智能体向第三方泄露私人信息。作为自主系统,智能体还存在未经用户意图或授权执行敏感任务(如银行交易)的更严重风险。认识到这一挑战,智能体安全社区已为安全智能体系统提出了众多建议。这项工作大多集中在产品级方法上,即智能体系统开发者为所有用户确定并应用相同的安全策略和权限。然而,不同用户有不同需求和偏好,因此智能体人工智能系统需要支持用户级权限策略。为了解人工智能智能体系统中用户级权限是如何处理的,我们调查了21个智能体权限系统提案。通过这项综述,我们构建了一个分类法,用于说明不同系统如何在用户界面和内部指定用户级权限策略;从用户输入中推导内部策略;并在运行时执行这些策略。然后,我们分析了五个著名的商业智能体,并将它们的权限处理与文献中的智能体权限系统进行比较。我们确定了文献和商业智能体中的几个高层次主题,以及未来工作需要填补的多个空白。

英文摘要

As AI agents gain prevalence, users are increasingly exposed to the risks such systems entail. Prompt injection attacks, as well as hallucination, can cause agents to leak private information to third parties. As autonomous systems, agents also present the more active danger of performing sensitive tasks, such as bank transactions, without the user's intent or authorization. Recognizing this challenge, the agentic security community has developed numerous proposals for secure agentic systems. Much of this work has focused on product-level approaches, where agentic system developers determine and apply the same security policies and permissions to all users. Yet different users have different needs and preferences, necessitating support for user-level permissions policies in agentic AI systems. To understand how user-level permissions are handled in AI agent systems, we survey 21 proposals for agent permissions systems. From this review, we construct a taxonomy of how different systems specify user-level permissions policies, both at the user interface and internally; derive internal policies from user input; and enforce those policies at run-time. We then analyze five prominent commercial agents and compare their permissions handling to agentic permissions systems in the literature. We identify several high-level themes across the literature and commercial agents, as well as multiple gaps where future work is needed.

URL PDF HTML 收藏
2605.29448 2026-07-21 cs.LG cs.AI cs.CV cs.IT math.IT 版本更新

How Much Is a Dataset Worth? Scaling Laws, the Vendi Score, and Matrix Spectral Functions

数据集值多少钱?缩放定律、Vendi分数与矩阵谱函数

Jeff A. Bilmes, Gantavya Bhatt, Arnav M. Das

机构 * Department of Electrical & Computer Engineering(电气与计算机工程系) Paul G. Allen School of Computer Science & Engineering(保罗·G·艾伦计算机科学与工程学院) University of Washington(华盛顿大学)

AI总结 本文通过子模性理论统一了神经缩放定律与Vendi分数,提出矩阵谱函数作为广义数据评估框架,并开发了基于割线方程的快速优化算法,在ImageNet-1K规模上实现了约35,000倍加速,实验表明设施选址函数在预测子集价值方面表现最佳。

Comments 75 pages

详情
AI中文摘要

神经缩放定律通过数据集大小评估数据,而Vendi分数使用量子熵衡量数据集价值。我们证明常见的神经缩放定律目标和Vendi分数都是子模的。进一步,我们表明Vendi分数是一类更广泛的子模目标(称为矩阵谱函数)的特例,这还包括行列式点过程(DPP)目标以及许多其他目标。我们还引入了弱矩阵单调函数,并展示了它们如何导致弱子模矩阵谱函数,从而产生一系列实用的数据评估目标。我们开发了基于割线方程的更新方法,避免了贪心优化过程中的重复特征分解,将$m$维嵌入的边际增益评估相对于预言机查询减少了$O(m)$因子。这实现了平均约35,000倍的实证加速,使得在ImageNet-1K规模的数据集上直接优化Vendi分数成为可能。由此,我们比较了多个目标在固定大小、类别平衡和固定训练预算条件下预测训练子集对保留测试性能价值的能力,包括Vendi分数、DPP、设施选址以及三种新的矩阵谱变体。在多个数据集上,设施选址表现最佳。直接优化还揭示,虽然Vendi分数在中等分数范围内具有预测性,但将目标推向更高值可能使其成为下游性能的糟糕代理。我们还发现,均匀随机选择的固定大小子集(无论是否类别平衡)在评估分数和保留性能上都表现出显著的集中性。最后,我们表明大小、类别平衡和训练预算单独并不决定数据价值:即使控制这些因素,性能范围也从好到差平滑变化。

英文摘要

Neural scaling laws appraise data through dataset size, while the Vendi Score uses quantum entropy to measure dataset value. We show both that common neural-scaling-law objectives and the Vendi Score are submodular. We further show that the Vendi Score is a special case of a broader class of submodular objectives that we call matrix spectral functions. This also includes determinantal (DPP) objectives, as well as many others. We also introduce weakly matrix monotone functions and show how they lead to weakly submodular matrix spectral functions, yielding a broad family of practical objectives for data appraisal. We develop secular-equation-based updates that avoid repeated eigendecompositions during greedy optimization, reducing marginal-gain evaluation for $m$-dimensional embeddings by an $O(m)$ factor relative to oracle queries. This yields an average empirical speedup of about 35,000x, making direct optimization of the Vendi Score feasible on ImageNet-1K-scale datasets. Thus enabled, we compare how well several objectives predict the value of training subsets for held-out test performance under fixed-size, class-balanced, and fixed training-budget regimes, including the Vendi Score, DPPs, facility location, and three new matrix spectral variants. Across multiple datasets, facility location performs the best. Direct optimization also reveals that, while the Vendi Score is predictive over moderate score ranges, pushing the objective to higher values can make it a poor downstream performance proxy. We also find that uniformly at random fixed-size subsets, both unconstrained and class-balanced, are remarkably concentrated in both appraisal scores and held-out performance. Finally, we show that size, class balance, and training budget do not alone determine data value: even when controlling for these factors, performance ranges smoothly from good to bad.

URL PDF HTML 收藏
2505.20161 2026-07-21 cs.LG cs.AI cs.CL 版本更新

Prismatic Synthesis: Gradient-based Data Diversification Boosts Generalization in LLM Reasoning

棱柱形合成:基于梯度的数据多样化提升语言模型推理中的泛化能力

Jaehun Jung, Seungju Han, Ximing Lu, Skyler Hallinan, David Acuna, Shrimai Prabhumoye, Mostafa Patwary, Mohammad Shoeybi, Bryan Catanzaro, Yejin Choi

机构 * NVIDIA Research(NVIDIA研究部) University of Washington(华盛顿大学) University of Southern California(南加州大学)

AI总结 研究语言模型训练数据多样性对泛化的作用,提出基于梯度熵的G - Vendi指标,进而构建棱柱形合成框架生成多样合成数据,有效提升模型性能,在多个基准测试中表现优于依赖更大数据生成器的模型。

详情
AI中文摘要

语言模型中的有效泛化关键取决于训练数据的多样性。现有多样性指标常依赖与模型行为脱节的表面启发式方法,难以达成目标。为此研究何种训练数据多样性驱动语言模型泛化及如何衡量与增强它。通过超300次训练运行的大规模实证分析表明,数据多样性可有力预测语言模型推理中的泛化。引入G - Vendi指标,基于模型诱导梯度的熵量化多样性,表现优于其他方法。在此基础上提出棱柱形合成框架,通过针对梯度空间中代表性不足区域生成多样合成数据。实验结果显示,随着合成数据规模增加,棱柱形合成持续提升模型性能,在多个基准测试中显著优于依赖更大数据生成器的现有模型。

英文摘要

Effective generalization in language models depends critically on the diversity of their training data. Yet existing diversity metrics often fall short of this goal, relying on surface-level heuristics that are decoupled from model behavior. This motivates us to ask: What kind of diversity in training data actually drives generalization in language models -- and how can we measure and amplify it? Through large-scale empirical analyses spanning over 300 training runs, carefully controlled for data scale and quality, we show that data diversity can be a strong predictor of generalization in LLM reasoning -- as measured by average model performance on unseen out-of-distribution benchmarks. We introduce G-Vendi, a metric that quantifies diversity via the entropy of model-induced gradients. Despite using a small off-the-shelf proxy model for gradients, G-Vendi consistently outperforms alternative measures, achieving strong correlation (Spearman's $ρ\approx 0.9$) with out-of-distribution (OOD) performance on both natural language inference (NLI) and math reasoning tasks. Building on this insight, we present Prismatic Synthesis, a framework for generating diverse synthetic data by targeting underrepresented regions in gradient space. Experimental results show that Prismatic Synthesis consistently improves model performance as we scale synthetic data -- not just on in-distribution test but across unseen, out-of-distribution benchmarks -- significantly outperforming state-of-the-art models that rely on 20 times larger data generator than ours. For example, PrismMath-7B, our model distilled from a 32B LLM, outperforms R1-Distill-Qwen-7B -- the same base model trained on proprietary data generated by 671B R1 -- on 6 out of 7 challenging benchmarks.

URL PDF HTML 收藏
2607.15396 2026-07-20 cs.CV cs.AI 新提交

Partial Information Decomposition as a Multi-Contrast 3D MRI Selection Strategy for Resource-Constrained Deep Neural Network Training in Brain Tumor Segmentation

部分信息分解作为一种多对比度3D MRI选择策略,用于脑肿瘤分割中资源受限的深度神经网络训练

Agamdeep Chopra, Mehmet Kurt

机构 * University of Washington(华盛顿大学)

AI总结 研究针对脑肿瘤分割中多对比度3D MRI分割计算量大的问题,采用部分信息分解框架对输入对排序选最优用于训练,实验表明该方法选出的T1c+T2-FLAIR是强双输入配置,证明了基于PID预训练选择的实用价值。

详情
AI中文摘要

当使用所有可用序列时,多对比度3D MRI分割在计算上要求很高。我们评估了一个预训练的部分信息分解框架,该框架根据输入对关于区域肿瘤负担的冗余、独特和协同信息对其进行排序,并选择排名最高的对用于下游训练。应用于T1n、T1c、T2w和T2-FLAIR MRI时,该框架选择了T1c+T2-FLAIR。然后,我们使用不同的输入配置训练了11个结构相同的轻量级3D U-Net。在一个独立测试队列中,T1c+T2-FLAIR是最强的双输入配置,在平均Dice中排名第二(所有四个输入为0.676,而所有四个输入为0.687)。对全输入模型的独立Shapley分析也确定T2-FLAIR和T1c是最有影响的输入,它们的成对交互作用最强。这些发现证明了基于PID的预训练选择在昂贵的3D模型开发之前识别紧凑、信息丰富的MRI输入集的实用价值。

英文摘要

Multi-contrast 3D MRI segmentation can be computationally demanding when all available sequences are used. We evaluate a pre-training Partial Information Decomposition framework that ranks input pairs according to their redundant, unique, and synergistic information about regional tumor burden and selects the highest-ranked pair for downstream training. Applied to T1n, T1c, T2w, and T2-FLAIR MRI, the framework selected T1c+T2-FLAIR. We then trained eleven architecturally identical lightweight 3D U-Nets using different input configurations. On an independent test cohort, T1c+T2-FLAIR was the strongest two-input configuration and ranked second overall in mean Dice (0.676 versus 0.687 for all four inputs). Independent Shapley analysis on the full-input model also identified T2-FLAIR and T1c as the most influential inputs and their pairwise interaction as the strongest. These findings demonstrate the practical value of PID based pre-training selection for identifying compact, informative MRI input sets before costly 3D model development.

URL PDF HTML 收藏
2605.22759 2026-07-20 cs.AI 版本更新

Towards a General Intelligence and Interface for Wearable Health Data

迈向可穿戴健康数据的通用智能与接口

Girish Narayanswamy, Maxwell A. Xu, A. Ali Heydari, Samy Abdel-Ghaffar, Marius Guerard, Kara Vaillancourt, Zhihan Zhang, Jake Garrison, Levi Albuquerque, Dimitris Spathis, Hong Yu, Hamid Palangi, Xuhai "Orson" Xu, David G. T. Barrett, Joseph Breda, Jed McGiffin, Yubin Kim, Yuwei Zhang, Naghmeh Rezaei, Samuel Solomon, Karan Ahuja, Tim Althoff, Jake Sunshine, Ming-Zher Poh, Benjamin Yetton, Ari Winbush, Nicholas B. Allen, James M. Rehg, Isaac Galatzer-Levy, Yun Liu, John Hernandez, Anupam Pathak, Conor Heneghan, Yuzhe Yang, Ahmed A. Metwally, Pushmeet Kohli, Mark Malhotra, Shwetak Patel, Xin Liu, Daniel McDuff

机构 * Google Research(谷歌研究) Google DeepMind(谷歌DeepMind) University of Washington(华盛顿大学) University of Oregon(俄勒冈大学) University of Illinois Urbana-Champaign(伊利诺伊大学厄巴纳-香槟分校)

AI总结 提出一个基于超过一万亿分钟无标签传感器数据预训练的可穿戴健康基础模型,通过联合扩展模型容量和预训练数据量,在35项健康预测任务上实现系统性性能提升,并利用LLM代理自动搜索下游预测头,集成到个人健康代理中以提高相关性和安全性。

Comments Narayanswamy and Xu are co-first authors. McDuff and Liu are co-last authors

详情
AI中文摘要

虽然无处不在的可穿戴传感器捕获了大量的行为和生理信息,但有效地将这些信号转化为个性化的健康见解具有挑战性。具体来说,由于高度的表型多样性以及个体基线健康、生理和生活方式因素的差异,将低层传感器数据转换为能够表征高层状态的表示是困难的。此外,收集带有健康结果注释的可穿戴数据既费力又昂贵,而回顾性注释实际上不可行,导致高质量标签数据的稀缺。为了克服这些限制,我们提出了一个可穿戴健康基础模型,该模型在来自五百万参与者的大型队列中超过一万亿分钟的无标签传感器信号上进行了预训练。我们证明了模型容量和预训练数据量的联合扩展在35项健康预测任务(涵盖心血管、代谢、睡眠和心理健康以及生活方式选择和人口统计因素)的多样化评估中带来了系统性的性能提升。我们发现这种人群规模的表示解锁了标签高效的少样本学习和稳健的日常指标估计的生成能力。为了进一步利用这种学习到的表示,我们部署了一个LLM代理教室来自动搜索基于模型嵌入构建的下游预测头空间,显示出随着LLM模型容量增加而广泛性能提升。最后,我们展示了将这些下游预测器集成到个人健康代理中如何能够支持更相关、更具上下文感知和更安全的模型响应,并通过来自一组临床医生的1,860个评分进行了验证。

英文摘要

While ubiquitous wearable sensors capture a wealth of behavioral and physiological information, effectively transforming these signals into personalized health insights is challenging. Specifically, converting low-level sensor data into representations capable of characterizing higher-level states is difficult due to high phenotypic diversity and variation in individual baseline health, physiology, and lifestyle factors. Moreover, collecting wearable data paired with health outcome annotations is laborious and expensive, and retrospective annotation remains practically unfeasible, contributing to a scarcity of data with high-quality labels. To overcome these limitations, we propose a foundation model for wearable health that is pretrained on more than one trillion minutes of unlabeled sensor signals drawn from a large cohort of five million participants. We demonstrate that the joint scaling of model capacity and pretraining data volume leads to systematic improvements in performance, as evaluated on a diverse set of 35 health prediction tasks, spanning cardiovascular, metabolic, sleep, and mental health, as well as lifestyle choices and demographic factors. We find that this population scale representation unlocks label-efficient few-shot learning and generative capabilities for robust daily metric estimation. To further leverage this learned representation, we deploy a classroom of LLM agents to autonomously search the space of downstream predictive heads built on the model embeddings, showing broad performance improvements that increase with LLM model capacity. Finally, we show how integrating these downstream predictors into a Personal Health Agent can support model responses that are more relevant, contextually aware, and safe, and we validate this via 1,860 ratings from a cohort of clinicians.

URL PDF HTML 收藏
2607.15271 2026-07-17 cs.CV cs.GR cs.LG 新提交

Online Neural Space Time Memory for Dynamic Novel View Synthesis

用于动态新视角合成的在线神经时空记忆

Baback Elmieh, Lynn Tsai, Zeman Li, Srinivas Kaza, Tiancheng Sun, Gabor Csapo, Ali Behrouz, Yuan Deng, Stephen Lombardi, Steven M. Seitz, Xuan Luo

机构 * University of Washington(华盛顿大学) Google(谷歌)

AI总结 研究多视图流视频在线新视角合成问题,提出解耦记忆更新与应用频率的方法,通过跨视图注意力管理变形,引入辅助记忆损失和记忆缓存策略,实现实时、领先性能及微小尺度在线记忆。

Comments 15 pages. Preprint. Project page with demos and video results: https://nst-mem.github.io

详情
AI中文摘要

从多视图流视频进行在线新视角合成面临一个基本权衡:在严格的实时约束下运行时,既要维护持久的长时记忆以重建暂时遮挡的区域。虽然测试时训练(TTT)提供了强大的记忆机制,但标准模型要求在每一帧基于梯度更新记忆以适应动态场景中变化的运动。大量记忆更新的计算成本排除了实时应用,且可能导致长上下文的不稳定。鉴于记忆更新比记忆应用要求更高且视频内容大多冗余,我们提出解耦这两个过程的频率。我们的方法在逐帧应用记忆时进行周期性记忆更新,使用跨视图注意力管理先前记忆状态与当前帧之间的变形。为锁定历史上下文,我们引入两个关键机制:辅助记忆损失强制场景的持久内化,以及记忆缓存策略规范活动权重以防灾难性漂移。我们的方法在具有动态人体运动的场景以及微小尺度的在线记忆方面展示了实时的、领先的性能。

英文摘要

Online novel view synthesis from multi-view streaming videos faces a fundamental trade-off: maintaining a persistent, long-horizon memory to reconstruct temporarily occluded regions while operating under strict real-time constraints. While Test-Time Training (TTT) offers a powerful memory mechanism, standard models mandate gradient-based memory updates at every frame to adapt to the changing motion in dynamic scenes. The computational cost of heavy memory updates precludes real-time application and can lead to instability over long contexts. Given that memory updates are more demanding than memory application and video content is largely redundant, we propose to decouple the frequencies of these two processes. Our approach performs periodic memory updates while applying the memory on a per-frame basis, using cross-view attention to manage deformations between the prior memory state and the current frame. To lock in the historical context, we introduce two critical mechanisms: an auxiliary Memory Loss that forces persistent internalization of the scene, and a Memory Caching strategy that regularizes active weights against catastrophic drift. Our method demonstrates real-time, state-of-the-art performance on scenes with dynamic human motion as well as minute-scale online memorization.

URL PDF HTML 收藏
2607.15267 2026-07-17 cs.AI cs.CL 新提交

Pretraining Data Can Be Poisoned through Computational Propaganda

预训练数据可通过计算宣传被下毒

Victoria Graf, Hannaneh Hajishirzi, Noah A. Smith, David Kohlbrenner, Kyle Lo

机构 * University of Washington(华盛顿大学) Allen Institute for Artificial Intelligence(艾伦人工智能研究所)

AI总结 研究发现预训练数据可通过公共讨论界面被下毒,引入HalfLife方法衡量恶意内容,探索在网络规模下毒预训练语料库的可行性,证明估计毒注入重要性,确立第三方网页内容为攻击语言模型预训练的可能载体。

详情
AI中文摘要

毒害预训练数据会给语言模型引入难以检测和缓解的有害行为。以往毒害预训练数据的工作大多利用维基百科等既定数据源,未体现预训练语料库的大规模和异质性,且忽视了中毒数据与数据处理管道的交互。本文通过现有网络规模内容注入机制——公共讨论界面,证明了在这种有限设置之外对预训练数据进行中毒攻击是可行的。此外,为衡量网络爬虫和数据处理后是否包含恶意内容,引入了HalfLife,一种用于估计基于网络爬虫的语言模型训练数据中对抗性内容包含情况的新颖分析方法。利用HalfLife探索通过开放讨论界面在网络规模下毒预训练语料库的可行性。分析表明估计预训练数据中是否包含毒注入的重要性,并将第三方网页内容确立为攻击语言模型预训练的可能载体。

英文摘要

Poisoning pretraining data can introduce harmful behaviors to LMs that are difficult to detect and mitigate. Prior work on poisoning pretraining data has largely exploited established data sources such as Wikipedia, which do not represent the large scale and heterogeneity typical of pretraining corpora, and has ignored the interaction between poisoned data and data curation pipelines. We demonstrate that poisoning attacks on pretraining data are feasible beyond this limited setting through an existing web-scale content injection mechanism: public discussion interfaces. Additionally, to measure whether malicious content is included after web crawling and data curation, we introduce HalfLife, a novel analysis for estimating adversarial content inclusion in web-crawl based LM training data. We use HalfLife to explore the feasibility of poisoning pretraining corpora at web scale through open discussion interfaces. Our analysis demonstrates the importance of estimating whether poison injections are included in pretraining data, and establishes third-party webpage content as a possible vector for attacking language model pretraining.

URL PDF HTML 收藏
2607.15265 2026-07-17 cs.CV cs.AI cs.MM cs.SD 新提交

SceneBind: Binding What and Where Across Vision, Audio and Language

SceneBind:跨视觉、音频和语言绑定“什么”与“哪里”

Mingfei Chen, Zijun Cui, Ruoke Zhang, Hyeonggon Ryu, Eli Shlizerman

机构 * University of Washington(华盛顿大学) University of Texas at Dallas(德克萨斯大学达拉斯分校) Hankuk University of Foreign Studies(韩国外国语大学)

AI总结 研究提出SceneBind全模态场景表示,结合全局语义与对象中心语义空间插槽解决空间结构缺失问题,还提出匹配方案。通过构建新数据集及训练协议进行训练评估,兼容预训练编码器,实现先进检索并能零样本转移到下游任务。

Comments Project website: https://scenebind.github.io/

详情
AI中文摘要

我们提出了SceneBind,一种对现实场景的全模态表示,具有跨视觉、音频和语言的联合语义和3D空间理解。现有全模态编码器在实例级语义(即存在什么)方面表现出色,但往往缺乏明确的空间结构(即其位置)。SceneBind通过将每个场景表示为语义空间实体来解决这一差距,结合全局语义嵌入和以对象为中心的语义空间插槽。我们还提出了SceneBind匹配,一种整合全局场景相似度与对象对齐的语义空间匹配方案,支持跨模态场景检索和对象定位。为训练和评估SceneBind,我们精心构建了一个具有结构化语义和空间注释的新型真实世界双耳视听数据集,并提出了一种跨模态对齐语义和空间信号的训练协议。SceneBind与大规模预训练语义编码器兼容,仅添加少量额外令牌即可进行轻量级空间建模。它实现了最先进的场景和空间检索,同时能够强大地零样本转移到下游任务,如视听定位。

英文摘要

We present SceneBind, an omni-modal representation of realistic scenes with joint semantic and 3D spatial understanding across vision, audio and language. Existing omni-modal encoders excel at instance-level semantics (i.e., what is present), but often lack explicit spatial structure (i.e., where it is). SceneBind addresses this gap by representing each scene as a semantic-spatial entity, combining a global semantic embedding with object-centric semantic-spatial slots. This representation explicitly captures object-level semantics, spatial attributes, and uncertainty. We further propose SceneBind Matching, a semantic-spatial matching scheme that integrates global scene similarity with object alignment, supporting cross-modal scene retrieval and object grounding. To train and evaluate SceneBind, we curate a novel real-world binaural audio-visual dataset with structured semantic and spatial annotations, and propose a training protocol for aligning semantic and spatial signals across modalities. SceneBind is compatible with large-scale pretrained semantic encoders, adds lightweight spatial modeling with only a few additional tokens. It achieves state-of-the-art scene and spatial retrieval while enabling strong zero-shot transfer to downstream tasks such as audio-visual localization.

URL PDF HTML 收藏
2607.15077 2026-07-17 cs.LG 新提交

An Introduction to Sparse Identification of Nonlinear Dynamics for Engineering Applications

工程应用中的非线性动力学稀疏识别介绍

Yao Cheng Li, Ana Larrañaga, Steven L. Brunton, Urban Fasel

机构 * Department of Aeronautics, Imperial College London(伦敦帝国理工学院航空系) Department of Mechanical Engineering, University of Washington(华盛顿大学机械工程系) NSF AI Institute in Dynamic Systems, University of Washington(华盛顿大学动态系统领域美国国家科学基金会人工智能研究所)

AI总结 介绍工程应用中非线性动力学稀疏识别(SINDy)方法,通过对候选非线性项库稀疏回归解决代理建模局限性,教程介绍该方法及扩展,经案例研究表明其易实现且灵活,是工程应用有价值的识别工具。

Comments 15 pages, 4 figures

详情
AI中文摘要

许多工程问题涉及控制方程表征不佳或仅部分已知的现象。神经网络等代理建模技术虽能捕捉系统行为,但需大量难以获取的训练数据集,且模型物理可解释性有限。稀疏识别非线性动力学(SINDy)方法通过对候选非线性项库进行稀疏回归,从小规模数据集中恢复可解释的控制方程,解决了上述两个局限性。本教程介绍了SINDy方法,并逐步介绍其主要扩展,从抗噪声弱形式和基于集成的变体到约束和可参数化公式。本文及配套教程分为三个部分:第一部分介绍标准SINDy算法并逐步扩展,让无先验知识读者能跟随步骤并将方法应用于自身问题;其余两部分给出详细案例研究,一是无人机系统识别,二是混沌热虹吸换热器。通过这些例子,旨在证明SINDy易于实现且足够灵活,可作为先进工程应用的有价值识别工具。

英文摘要

Many engineering problems involve phenomena whose governing equations are poorly characterized or only partially known. Surrogate modeling techniques such as neural networks can capture the behavior of these systems, but they typically demand large training datasets that are difficult to obtain in engineering contexts and yield models with limited physical interpretability. The Sparse Identification of Nonlinear Dynamics (SINDy) method addresses both limitations by performing sparse regression over libraries of candidate nonlinear terms, recovering interpretable governing equations from comparatively small datasets. Although SINDy has been demonstrated extensively on canonical benchmark systems, its application to practical engineering problems is less widely documented. This tutorial introduces the SINDy method and progressively builds toward its main extensions, from noise-robust weak-form and ensembling-based variants to constrained and parametrizable formulations. The paper and the accompanying tutorial (available at https://github.com/paullililili/SINDy4Engineers) is organized in three parts: the first introduces the standard SINDy algorithm and progressively extends it, inviting readers without prior knowledge to follow each step and adapt the methods to their own problems; the remaining two parts present detailed case studies on (1) the system identification of an unmanned aerial vehicle and (2) a chaotic thermosyphon heat exchanger. Through these examples, we aim to demonstrate that SINDy is simple to implement yet flexible enough to serve as a valuable identification tool for advanced engineering applications.

URL PDF HTML 收藏
2607.14703 2026-07-17 cs.CV cs.AI 新提交

Pretraining Multiple Instance Learning Networks with Multi-Teacher Distillation from Pathology Slide Foundation Models

基于病理切片基础模型的多教师蒸馏预训练多实例学习网络

Mingxi Fu, Jiawen Li, Renao Yan, Jiali Hu, Qiehe Sun, Tian Guan, Yonghong He

机构 * Tsinghua Shenzhen International Graduate School, Tsinghua University(清华大学深圳国际研究生院) University of Washington(华盛顿大学) City University of Hong Kong (Dongguan)(香港城市大学(东莞)) Medical Optical Technology R&D Center, Research Institute of Tsinghua(清华研究院医学光学技术研发中心) Jinfeng Laboratory(金凤实验室)

AI总结 针对计算病理学中多实例学习存在的问题,提出基于蒸馏的预训练框架,利用两个基础模型作为教师,引入角分散归一化蒸馏损失,将蒸馏权重用于下游适应,实验表明该方法在少样本场景中优势明显,能提升计算效率。

详情
AI中文摘要

多实例学习(MIL)已成为计算病理学中全切片图像(WSI)分析的主要范式。现有MIL聚合器通常针对每个下游任务从头开始训练,依赖有限的切片级标签同时学习聚合机制和下游判别表示,存在优化不稳定、过拟合和可迁移性有限等问题。本文提出基于蒸馏的MIL预训练框架,利用两个切片级基础模型TITAN和CARE作为教师,将其表示知识转移到多种MIL架构中。引入角分散归一化蒸馏损失平衡不同教师的监督,蒸馏权重用作下游适应的初始化。在15个基准数据集上进行系统评估,结果表明预训练总体上优于从头训练,尤其在少样本场景中,同时保持轻量级MIL模型的计算效率。

英文摘要

Multiple instance learning (MIL) has become the main paradigm for whole-slide image (WSI) analysis in computational pathology. However, existing MIL aggregators are still typically trained from scratch for each downstream task, relying on limited slide-level labels to learn both aggregation mechanisms and downstream discriminative representations simultaneously. As a result, they often suffer from unstable optimization, overfitting, and limited transferability. Similar to pretrained ResNet and Vision Transformer models in natural image learning, MIL also requires reusable pretrained initialization. However, high-quality slide-level pretraining data remain scarce, and MIL models are usually lightweight and weakly supervised, making large-scale pretraining difficult in practice. To address this challenge, we propose a distillation-based pretraining framework for MIL, which leverages two slide-level foundation models, TITAN and CARE, as teachers to transfer their representational knowledge into a diverse set of MIL architectures. To effectively balance supervision from different teachers, we further introduce an angular dispersion normalized distillation loss. The distilled weights are then used as initialization for downstream adaptation. We conduct systematic evaluations on 15 benchmark datasets under both linear probing and full-parameter fine-tuning, and further validate its advantages in few-shot scenarios. Experimental results show that pretraining generally improves MIL aggregators over from scratch training, especially in linear-probing and few-shot settings, while maintaining the computational efficiency of lightweight MIL models. Code is available at https://github.com/fu0201/MIL_Pretrained.

URL PDF HTML 收藏
2607.14611 2026-07-17 cs.CR cs.AI cs.MA 新提交

Bad Memory: Evaluating Prompt Injection Risks from Memory in Agentic Systems

糟糕的记忆:评估智能体系统中内存引发的提示注入风险

Soham Gadgil, David Alexander, Sai Sunku, Franziska Roesner

机构 * University of Washington(华盛顿大学)

AI总结 研究基于内存的智能体系统中提示注入攻击,利用沙盒合成工作区评估两个系统四个模型,发现虽难用外部内容重写内存文件,但已植入的有效载荷可攻击当前及未来会话,揭示持久内存改变威胁模型并推动相关防御研究。

Comments Preprint

详情
AI中文摘要

一类不断发展的智能体系统通过内存文件、行为偏好和知识库在会话间维持持久状态。这虽使智能体更有用且能自我改进,但也为提示注入创造了新的攻击面,恶意指令可嵌入持久文件并影响未来行为。本文利用沙盒合成工作区研究基于内存的智能体系统中的提示注入攻击。我们评估了Anthropic Claude Code和OpenAI Codex这两个智能体系统的四个模型:Claude Haiku 4.5、Claude Opus 4.7、GPT - 5.2和GPT - 5.5。结果表明,虽难以用不可信外部内容使智能体重写自身内存文件,但已植入文件的有效载荷可成功攻击当前及未来会话。攻击成功率和有效载荷持久性在不同系统、模型、对抗目标和多会话攻击序列中差异很大。这些发现表明持久内存改变了提示注入的威胁模型,并促使开发在不消除智能体有益适应性的情况下保护内存更新的防御措施。

英文摘要

A growing class of agentic systems maintain persistent state across sessions through memory files, behavioral preferences, and knowledge bases. While this makes agents more useful and self-improving, it also creates a new attack surface for prompt injections in which malicious instructions can be embedded within persistent files and influence future behavior. In this work, we study prompt injection attacks in memory-based agentic systems using a sandboxed synthetic workspace. We evaluate two agentic systems, Anthropic Claude Code and OpenAI Codex, across four models: Claude Haiku 4.5, Claude Opus 4.7, GPT-5.2, and GPT-5.5. Our results show that although it is difficult to make an agent overwrite its own memory files using untrusted external content, payloads already planted in those files can successfully attack current and future sessions. Attack success and payload persistence vary substantially across systems, models, adversarial goals, and multi-session attack sequences. These findings show that persistent memory changes the threat model for prompt injection and motivate defenses that protect memory updates without removing useful agent adaptation.

URL PDF HTML 收藏
2607.14418 2026-07-17 cs.LG econ.GN q-fin.EC 新提交

Adaptive Ad Load Design for Sponsored Search Markets: Evidence, Theory, and Deployment

赞助搜索市场的自适应广告加载设计:证据、理论与部署

Mohammad Rashid, Hema Yoganarasimhan

机构 * University of Washington(华盛顿大学)

AI总结 研究赞助搜索市场广告加载设计权衡,通过安卓应用商店实验发现增加广告加载量对收入、转化率和参与度的影响及异质性,设计并部署自适应算法e-LAAL,在生产部署中改善收益与转化率权衡,优于静态基准。

Comments 54 pages

详情
AI中文摘要

广告加载设计是赞助搜索中供应方的核心决策。更多赞助位虽能增加收入,但可能排挤自然搜索结果并降低用户体验。我们在安卓应用商店进行大规模随机现场实验,超五百万用户接触一至六个赞助位。增加广告加载量最多可使收入提高43%,但总搜索转化率最多降低5%,每日参与度最多降低2.2%。不同查询的效果存在异质性。基于此,我们设计并部署了新型自适应算法e-LAAL,它结合了LAAL和静态探索策略,还提供了有限时间动态遗憾保证。在面向2230万用户和7760万次搜索的平台级生产部署中,e-LAAL改善了收益与转化率的权衡,优于统一和历史查询相关的静态基准。

英文摘要

Ad-load design is a central supply-side decision in sponsored search: more sponsored slots can raise revenue, but may crowd out organic results and degrade user outcomes. We study this trade-off using a large-scale randomized field experiment on an Android app store, where over five million users are exposed to one through six sponsored slots. Increasing ad load raises revenue by up to 43%, but reduces total search conversions by up to 5% and daily engagement by up to 2.2%. These average effects mask substantial heterogeneity: additional slots generate large revenue gains for high-ad-conversion queries, but little or negative marginal revenue for low-conversion queries. The trade-off also shifts within query as advertiser composition changes, such as brand-advertiser presence. Motivated by these findings, we design and deploy a novel adaptive algorithm -- exploration-augmented Locally Adaptive Ad Load (e-LAAL). e-LAAL combines LAAL, a model-free query-level decision rule that updates ad-load recommendations using recent outcomes, with static exploration arms that maintain support and provide fixed-policy counterfactual benchmarks. We provide a finite-time dynamic-regret guarantee for the e-LAAL architecture. In a platform-level production deployment serving 22.3 million users and 77.6 million searches, e-LAAL improves the empirical revenue--conversion trade-off relative to deployed static benchmarks and outperforms uniform and historical query-dependent static benchmarks.

URL PDF HTML 收藏
2607.13938 2026-07-16 cs.RO 新提交

Discriminative Barrier Functions for Safe Adversarial Imitation Learning from Observation

用于从观察中进行安全对抗模仿学习的判别障碍函数

Anubhav Vishwakarma, Bhaumik Mehta, Caleb Hsu, Byron Boots, Karen Leung, Tyler Han

机构 * University of Washington(华盛顿大学)

AI总结 研究针对逆强化学习不安全及控制障碍函数设计难的问题,通过将奖励函数候选限制在CBF空间,实现安全在线控制与经验改进,能从无标签观察中恢复障碍函数,模拟实验显示其安全性能提升,并研究了不同IRL方法的权衡。

Comments 20 pages, 5 figures

详情
AI中文摘要

逆强化学习(IRL)算法是从专家示范中学习和泛化的强大工具,但通常依赖无约束探索,对实际部署不安全。同时,控制障碍函数(CBF)可保证控制系统安全,但其解析设计耗时且深奥。本文通过在IRL中将奖励函数候选限制在CBF空间来共同解决这些限制,实现具有持续经验改进的安全在线控制。关键是,该框架能直接从无标签专家观察中数据驱动恢复障碍函数。实验表明,恢复的障碍函数对专家数据中完全不存在的不安全状态具有鲁棒性,在模拟导航环境中安全性能优于标准IRL基线,并研究了基于规划与基于策略的IRL方法在模拟和现实世界避障任务中的权衡。

英文摘要

Inverse Reinforcement Learning (IRL) algorithms are powerful tools for learning from and generalizing expert demonstrations, but they often rely on unconstrained exploration, rendering them unsafe for real-world deployment. Meanwhile, Control Barrier Functions (CBFs) can guarantee the safety of control systems, but the analytical design of CBFs can be time-consuming and esoteric. In this work, we address these limitations jointly by constraining reward function candidacy during IRL to the space of CBFs, yielding a formulation that exhibits safe online control with continuous experiential improvement. Crucially, this framework enables the data-driven recovery of barrier functions directly from unlabeled expert observations. We demonstrate that the recovered barrier function is robust to unsafe states entirely absent from the expert data. Furthermore, we benchmark our method against standard IRL baselines in a simulated navigation environment, demonstrating improved safety performance. Finally, we investigate the trade-offs of planning-based versus policy-based IRL methods across both simulation and a real world obstacle avoidance task.

URL PDF HTML 收藏
2607.13591 2026-07-16 cs.CL cs.AI 新提交

Memory as a Controlled Process: Learned Adaptive Memory Management for LLM Agents

作为受控过程的记忆:为大语言模型智能体学习自适应内存管理

Eric Hanchen Jiang, Zhi Zhang, Yuchen Wu, Levina Li, Dong Liu, Xiao Liang, Rui Sun, Yubei Li, Edward Sun, Haozheng Luo, Zhaolu Kang, Aylin Caliskan, Kai-Wei Chang, Ying Nian Wu

机构 * University of California Los Angeles(加利福尼亚大学洛杉矶分校) University of Washington(华盛顿大学) Northwestern University(西北大学)

AI总结 研究LLM智能体内存管理问题,提出MemCon框架将内存操作建模为马尔可夫决策过程,通过在线策略自适应管理内存,该框架与后端无关,实验表明其在多基准测试中优于基线,提升任务成功率并减少令牌消耗。

详情
AI中文摘要

大语言模型(LLM)智能体越来越依赖外部内存系统来积累跨任务经验。然而,几乎所有现有方法,从图结构内存到反思洞察存储,都通过固定的、手工设计的启发式方法访问内存。我们认为,这种静态的内存观点是智能体学习的核心瓶颈,因为最佳内存行为本质上依赖于上下文。我们提出了“作为受控过程的记忆”(MemCon)框架,将内存操作建模为马尔可夫决策过程,并学习一种在线策略,以自适应地决定何时、检索什么以及检索多少,何时注入提炼的计划,以及何时进行合并或遗忘。MemCon与后端无关,通过任务级二进制反馈学习,无需预训练和额外的LLM调用。在6个基准测试、3个智能体框架和3个LLM主干上,MemCon在任务成功率上比多个内存基线高出15.2分,同时减少了5%-20%的令牌消耗。

英文摘要

Large Language Model (LLM) agents increasingly rely on external memory systems to accumulate experience across tasks. Yet nearly all existing approaches, from graph-structured memories to reflective insight stores, access memory through fixed, hand-designed heuristics. We argue that this static view of memory is a core bottleneck for agentic learning because optimal memory behavior is fundamentally context-dependent. The early stages of the tasks, benefit from minimal retrieval because memory is sparse; recurring goal types benefit from plan reuse rather than generic nearest-neighbor lookup; stuck agents benefit from re-retrieval with alternative queries; and across long task streams, the memory store itself must be consolidated and pruned to remain useful. We present Memory as a Controlled Process (MemCon), a framework that models memory operations as a Markov Decision Process and learns an online policy that adaptively decides when, what, and how much to retrieve, when to inject a distilled plan, and when to consolidate or forget. MemCon is backend-agnostic: it wraps any existing memory implementation, learns from task-by-task binary feedback with no pretraining and no additional LLM calls, and uses a lightweight tabular contextual bandit with UCB exploration that converges within tens of tasks. Across 6 benchmarks, 3 agent frameworks, and 3 LLM backbones, MemCon consistently outperforms multiple memory baselines by up to 15.2 points in task success while reducing token consumption by 5--20%.

URL PDF HTML 收藏
2607.13558 2026-07-16 cs.AI 新提交

Multi-Agent Collaborative Reasoning with Tool-Augmented Evidence for Urban Region Profiling

基于工具增强证据的多智能体协作推理用于城市区域剖析

Xixuan Hao, Yutian Jiang, Jiabo Liu, Yihang Yang, Guangyin Jin, Song Gao, Yuxuan Liang

机构 * The Hong Kong University of Science and Technology (Guangzhou)(香港科技大学(广州)) University of Washington(华盛顿大学) Chang’an University(长安大学) University of Wisconsin - Madison(威斯康星大学麦迪逊分校)

AI总结 研究针对城市区域剖析问题,提出UrbanAgent框架,通过多智能体协作推理解决跨模态不一致,将指标预测扩展为闭环过程,经实验验证其性能优于现有基线,在未见城市设置中有强泛化性。

Comments Accepted by KDD 2026

详情
AI中文摘要

城市区域剖析是城市计算中的核心问题,支持人口估计、经济评估和环境监测等应用。现有方法通常将此任务表述为多模态表示学习,融合卫星图像、兴趣点、文本描述和3D建筑信息等异构城市数据到潜在嵌入中进行预测。但这些方法大多由相关性驱动,假设跨模态一致性,依赖静态管道,限制了其在异构或未见城市区域的鲁棒性。我们提出UrbanAgent,一个将城市区域剖析重新构建为推理驱动推理问题的智能体框架。UrbanAgent为每个数据模态实例化一个独立智能体,进行结构化多智能体协作推理以明确解决跨模态不一致性,而非将其吸收到单一表示中。此外,UrbanAgent将指标预测扩展为主动证据获取和迭代推理的闭环过程,使智能体能够通过经强化学习优化的外部知识的工具增强检索来验证不确定推理。在全球城市数据集上进行的碳排放、GDP和人口估计的广泛实验表明,UrbanAgent始终优于现有基线,R2平均提高8.1%,并在未见城市设置中表现出强大的泛化性能。

英文摘要

Urban region profiling constitutes a core problem in urban computing, supporting applications such as population estimation, economic assessment, and environmental monitoring. Existing methods typically formulate this task as multimodal representation learning, fusing heterogeneous urban data, e.g., satellite imagery, points of interest, textual descriptions, and 3D building information, into latent embeddings for prediction. However, these approaches are largely correlation-driven, assume cross-modal consistency, and rely on static pipelines, which limit their robustness in heterogeneous or unseen urban regions. We propose UrbanAgent, an agentic framework that reframes urban region profiling as a reasoning-driven inference problem. UrbanAgent instantiates an independent agent for each data modality and performs structured multi-agent collaborative reasoning to explicitly address cross-modal inconsistencies rather than absorbing them into a single representation. In addition, UrbanAgent extends indicator prediction as a closed-loop process of active evidence acquisition and iterative reasoning, enabling agents to verify uncertain inferences through tool-augmented retrieval of external knowledge optimized via reinforcement learning. Extensive experiments on global urban datasets for Carbon emissions, GDP, and Population estimation show that UrbanAgent consistently outperforms existing baselines, achieving an average improvement of 8.1% in R2, and exhibiting strong generalization performance in unseen-city settings.

URL PDF HTML 收藏
2607.13498 2026-07-16 cs.LG 新提交

Factorized Spectral Representations for Reinforcement Learning

用于强化学习的因式分解谱表示

Junyi Wu, Dan Li

机构 * University of Washington(华盛顿大学)

AI总结 该研究聚焦强化学习,提出FaStR方法,通过对转移核的三模张量CP分解,用噪声对比目标拟合,产生单独编码器形成谱表示。其因式分解形式缩小假设类,在高维运动任务中效果好,状态编码器可跨执行器移位转移。

详情
AI中文摘要

从交互数据中学习世界的紧凑模型是高效样本深度强化学习的核心。谱表示方法通过将转移核视为矩阵,以状态-动作对和下一个状态为两侧,通过自监督对比目标学习低秩分解,已成为连续控制中表示学习的主导范式。我们进一步拓展这一观点。转移核自然是关于状态、动作和下一个状态的三模张量,CP分解为每个模式给出一个特征图。我们提出FaStR,它通过噪声对比目标拟合这种分解,产生单独的状态、动作和下一个状态编码器,共同形成单个谱表示。因式分解形式产生更小的假设类,表示学习所需的样本大小按与状态和动作维度中较小者成比例的因子缩小。实证上,FaStR在动力学与因式分解结构一致的高维运动任务上取得最大收益,并且学习到的状态编码器在执行器移位时完整转移,只需重新训练动作编码器。

英文摘要

Learning a compact model of the world from interaction data is central to sample-efficient deep reinforcement learning. Spectral representation methods have become the leading paradigm for representation learning in continuous control by taking a matrix view of the transition kernel, with state-action pairs on one side and next states on the other, and learning a low-rank factorization through self-supervised contrastive objectives. We take this view one step further. The transition kernel is naturally a three-mode tensor over states, actions, and next states, and a CP decomposition gives one feature map per mode. We propose FaStR, which fits this decomposition with a noise contrastive objective, producing separate state, action, and next-state encoders that together form a single spectral representation. The factored form yields a smaller hypothesis class, and the sample size needed for representation learning shrinks by a factor that scales with the smaller of the state and action dimensions. Empirically, FaStR delivers its largest gains on high-dimensional locomotion tasks whose dynamics align with the factored structure, and the learned state encoder transfers intact across actuator shift while only the action encoder is retrained.

URL PDF HTML 收藏
2607.04113 2026-07-16 cs.LG cs.NA math.NA 版本更新

Asymptotic Preservation and Uniform Accuracy of Diffusion and Flow-Matching Samplers

扩散与流匹配采样器的渐近保持后验分析

Shiheng Zhang

机构 * University of Washington(华盛顿大学)

AI总结 研究将最小标准差视为奇异摄动参数,通过后验审计确定固定步长采样器的渐近保持性,分析不同时钟在终端层的稳定性及准确性,在特定模型上验证确定性和随机采样器特性,并用于EDM CIFAR-10检查点分析。

详情
AI中文摘要

扩散和流匹配采样器将学习到的概率流常微分方程从大噪声尺度积分到小终端下限σ_min,此时得分僵硬且流形成边界层。我们将σ_min视为奇异摄动参数,确定哪些固定步长采样器是渐近保持的,将标准作为后验审计:具有σ_min均匀系数的残差泛函,可在预训练检查点上计算,无需真实得分或精确轨迹。在终端层,σ时钟中的欧拉方法、确定性DDIM更新是唯一的层精确离散化,λ时钟仅在步长h≤h_star = 1 + W(1/e)时稳定,均匀σ^2热时钟在距数据σ_min无关的距离处停滞。在两个可解模型上,确定性采样器保持一阶均匀准确性,无log(1/σ_min)因子,对数完全归因于随机采样器的伊藤项,其路径KL与常微分方程的预算相比缩放为Λ^2/N,而常微分方程的预算为O(Λ^2/N^2),其中Λ = log(σ_max/σ_min)。在EDM CIFAR-10检查点上,一次测量的光谱可预测跨步数、调度和噪声水平的留出残差预算,无需针对每个配置重新拟合,并在M_1 = 1.00±0.01处校准伊藤系数。时钟决定稳定性;噪声而非几何结构导致对数出现。

英文摘要

Diffusion and Gaussian-interpolant flow-matching samplers approach data through a terminal noise floor $\varepsilon$, a singular limit for manifold-supported or rank-deficient data. We study two properties of a complete sampler specification, comprising its update rule, time grid, and terminal rule. Asymptotic preservation (AP) means a stable and consistent zero-noise discretization with a step count bounded independently of $\varepsilon$. Uniform accuracy (UA) of order $p$ means that, at numerical resolution $h$, the endpoint $W_2$ error is $O(h^p)$ with a floor-independent constant. Bounded log-noise stepping fails AP because its step count diverges. Stopping a stable base solver at a positive switching scale $a$ and appending one map fitted to the analytic normal mode restores AP. On smooth compact boundaryless manifolds, the standard map has exact-input error $O(a^2-\varepsilon^2)$ and sharp zero-floor error $Θ(a^2)$. A base solver with a floor-uniform order-$p$ estimate on the resolved interval retains that order when $a=O(h^{p/2})$, provided the terminal transfer factor remains bounded. Along exact trajectories, the posterior-mean identity $D(x(σ),σ)=x(σ)-σx'(σ)$ cancels the linear terminal defect and enables higher-order fitted maps. A three-evaluation Hermite construction is uniformly third order for exact switching-scale input over $0\le\varepsilon\le a$, and a seven-evaluation construction is fourth order at zero. We classify representative diffusion and flow-matching specifications by AP and UA. On EDM and Rectified Flow checkpoints, a paired decomposition separates base-integration from terminal-completion error and predicts held-out same-seed endpoint errors.

URL PDF HTML 收藏
2605.20689 2026-07-16 cs.CL cs.AI cs.IR cs.LG 版本更新

DIVE: Embedding Compression via Self-Limiting Gradient Updates

DIVE: 通过自限制梯度更新实现嵌入压缩

Dongfang Zhao

机构 * University of Washington Tacoma School of Engineering and Technology(华盛顿大学塔可姆分校工程与技术学院)

AI总结 本文提出DIVE方法,通过自限制的三元组损失和头级NT-Xent对比损失解决嵌入压缩中因标注数据稀缺导致的过拟合问题,提升了检索性能。

详情
AI中文摘要

大型语言模型的高维嵌入对向量搜索系统造成了显著的存储和计算成本。最近的嵌入压缩方法,包括Matryoshka-Adaptor(EMNLP 2024)、Search-Adaptor(ACL 2024)和SMEC(EMNLP 2025),通过轻量级残差适配器实现降维,但其训练目标在标注数据稀缺时导致严重过拟合,使检索性能低于冻结基线。我们提出DIVE(通过隐式视图集合进行降维),一种压缩适配器,通过两种机制解决这一失败。首先,一个自限制的基于hinge的三元组损失在三元组满足边距约束时产生零梯度,限制应用于预训练嵌入空间的总扰动。其次,头级NT-Xent对比损失将每个嵌入的多个学习投影视为隐式视图,提供密集的自监督梯度,补偿小数据集上三元组信号的稀疏性。在六个BEIR数据集上,DIVE在每个数据集和每个评估的压缩比上均优于所有三个基线适配器,具有14M参数的开源实现。

英文摘要

High-dimensional language-model embeddings increase storage and search costs, while supervised compressors can overfit when relevance labels are scarce. We present DIVE (Dimensionality reduction with Implicit View Ensembles), a residual compression adapter codesigned with a self-limiting hinge loss, geometry distillation, and head-wise NT-Xent over implicit coordinate views. The hinge stops updating satisfied ranking constraints, while the dense objectives stabilize the compressed representation; only the first head is retained at inference. Under query-disjoint evaluation with two LLM2Vec backbones, five BEIR benchmarks, 128d and 256d outputs, and six baselines, DIVE is the strongest adapter on all five primary benchmarks. It also outperforms PCA and an autoencoder in comparisons against unsupervised compressors.

URL PDF HTML 收藏
2607.12886 2026-07-15 cs.AI 新提交

A Multi-Agent System for Autonomous, Fine-Tuning-Free Clinical Symptom Detection: Development and Validation Study

一种用于自主、无需微调的临床症状检测的多智能体系统:开发与验证研究

Cameron Cagan, Pedram Fard, Jiazi Tian, Jingya Cheng, Shawn N. Murphy, Hossein Estiri

机构 * Massachusetts General Hospital(麻省总医院) University of Washington(华盛顿大学)

AI总结 研究针对临床症状检测中信息难结构化及现有方法不足的问题,提出多智能体系统Pythia,无需人工提示工程或微调,能自主优化提取提示。通过与词汇表比较,验证其在临床记录症状提取上的有效性及推广性,优于部分传统方法。

详情
AI中文摘要

临床记录包含许多使患者就医的体征和症状,但这些信息很少进入结构化字段。现有提取方法要么依赖产生误报的上下文无关规则,要么依赖需要大量微调的监督模型。我们提出了Pythia,一个多智能体系统,它能自主编写和优化临床概念的提取提示,无需人工提示工程或微调。Pythia在本地托管的开放权重模型上运行,将临床记录保存在本地基础设施上,并根据开发集的敏感性和特异性选择提示。我们将Pythia与一个精心策划的词汇表在400份代表387名患者的临床记录中的72种体征和症状上进行了比较。每个概念的开发集(n = 300)和验证集(n = 100)独立划分。Pythia的平均敏感性为0.76,特异性为0.95,而词汇表分别为0.82和0.76,在62个直接可比概念中的20个概念上,Pythia在这两个指标上匹配或超过了词汇表。对于词汇表将每份记录都标记为阳性的14个概念,Pythia通过要求是现在时态、患者归因的发现而不是对术语的任何文本提及,恢复了0.97的平均特异性。特异性从开发集转移到验证集时,在不同患病率下退化最小,而敏感性转移在患病率低于5%时减弱,在患病率低于2%时平均差距达到0.25。在相同开发集上按每个概念微调的BERT分类器平均敏感性为0.23,对于患病率低于约5%的概念,敏感性降至零。这些发现表明,自主、无需微调的提示优化可以产生症状提取提示,能从开发集有效推广到验证集,同时仍可在本地基础设施上部署。

英文摘要

Clinical notes contain many of the signs and symptoms that bring patients to care, yet this information rarely reaches structured fields. Existing extraction approaches either rely on context-insensitive rules that generate false positives or on supervised models that require substantial fine-tuning. We present Pythia, a multi-agent system that autonomously writes and optimizes extraction prompts for clinical concepts without manual prompt engineering or fine-tuning. Running on a locally hosted open-weights model, Pythia keeps clinical notes on local infrastructure and selects prompts using development-set sensitivity and specificity. We compared Pythia with a curated lexicon across 72 signs and symptoms from 400 clinical notes representing 387 patients. Development (n=300) and validation (n=100) sets were partitioned independently for each concept. Pythia achieved mean sensitivity of 0.76 and specificity of 0.95, compared with 0.82 and 0.76 for the lexicon, and matched or exceeded the lexicon on both metrics for 20 of 62 directly comparable concepts. For 14 concepts where the lexicon labeled every note positive, Pythia recovered mean specificity of 0.97 by requiring a present-tense, patient-attributed finding rather than any textual mention of a term. Specificity transferred from development to validation with minimal degradation across prevalences, whereas sensitivity transfer weakened below 5% prevalence, reaching a mean gap of 0.25 below 2% prevalence. A BERT classifier fine-tuned per concept on the same development set achieved mean sensitivity of 0.23 and collapsed to zero sensitivity for concepts below roughly 5% prevalence. These findings suggest that autonomous, fine-tuning-free prompt optimization can produce symptom extraction prompts that generalize effectively from development to validation while remaining deployable on local infrastructure.

URL PDF HTML 收藏
2607.12520 2026-07-15 cs.AI 新提交

The Model Knows Your Project, Not You: Measuring Recognition in LLMs with NameRank

模型了解你的项目,而非你本人:使用NameRank衡量大语言模型中的识别度

Bojie Li, Noah Shi

机构 * Pine AI(松树人工智能公司) University of Washington(华盛顿大学)

AI总结 研究用NameRank衡量大语言模型对实体的识别度,通过对多实体多模型探测及独立评判获取分数,发现识别关注可索引工件,奥运资质与知名奖项情况不同,独立创作者中工具与创造者排名有别等,并指出文献计量法难测识别度等结论。

详情
AI中文摘要

前沿模型在任何检索步骤之前,从自身权重中回忆起的关于个人或工具的信息,往往会塑造人类看到的首个描述,这使得参数化语料库的存在成为一个测量问题。引用能解释约三分之一模型是否识别研究人员的情况;我们针对剩余部分构建了NameRank,这是一个[0,1]的识别分数。对54个群组中的4685个实体,用一个开放式问题在36个模型上进行探测,由独立评判员根据精心策划的黄金标准给出二元判定。合成空实体得分接近零,判定追踪实体而非模型。研究发现:识别关注的是有名称、可索引的工件,而非资质或头衔。每个奥运式资质都低于在职研究人员基线,在知名奖项层面排名反转。对于独立创作者,工具的排名高于其创造者,传播的资质是命名方法或获奖论文。作为著名工件的众多署名贡献者之一,所得认可几乎为零。没有文献计量法能很好地预测识别度;高引用密度机构在相同引用量下比同行识别度更高;在258个新闻事件中,识别取决于峰值显著性而非持续性。自我报告探测显示自我反省读取的是语料库而非自身知识。

英文摘要

What a frontier model recalls about a person or tool from its own weights -- before any retrieval step -- often shapes the first description a human sees, making that parametric corpus presence a measurement problem. Citations explain about a third of whether a model recognizes a researcher; we target the residual and build NameRank, a [0,1] recognition score: each of 4,685 entities in 54 cohorts is probed with one open-ended question across 36 models, and an independent judge returns a binary verdict against a curated gold -- did the model state a specific, non-guessable fact about this exact entity? -- so hallucination, context echo, and guesses earn nothing. Synthetic-null entities hold the floor near zero, and verdicts track the entity, not the model. One thesis organizes the findings: recognition is paid to named, indexable artifacts, not to credentials or titles. Every Olympic-style credential sits below a working-researcher baseline, because no named artifact ships with the medal, yet the ranking inverts at the marquee tier, where Nobel, Turing, and Fields laureates saturate the panel. For independent creators the tool out-ranks its maker, and the credential that does propagate is a named method or awarded paper. Being one of many named contributors to a celebrated artifact, by contrast, earns almost nothing -- the authors listed on a flagship model report or system card sit near the recognition floor -- because recognition attaches to the artifact's own distinctive name, not to the roster behind it. No bibliometric predicts recognition well; top-density institutions out-recognize peers at matched citations; and on 258 news events recognition loads on peak salience, not persistence. A self-report probe shows introspection reads a corpus prior, not its own knowledge.

URL PDF HTML 收藏
2607.12227 2026-07-15 cs.AI 新提交

Rethinking the Evaluation of Harness Evolution for Agents

重新思考智能体的 harness 进化评估

Yike Wang, Huaisheng Zhu, Zhengyu Hu, Yige Yuan, Zhengyu Chen, Shakti Senthil, Hannaneh Hajishirzi, Yulia Tsvetkov, Pradeep Dasigi, Teng Xiao

机构 * Allen Institute for AI(艾伦人工智能研究所) University of Washington(华盛顿大学)

AI总结 研究重新审视大语言模型智能体的自动 harness 进化评估,通过在可比条件下与基线比较及在保留任务上测试,发现其不总优于简单方法且泛化有限,对其有效性提出质疑,强调需更公平评估协议和基准。

详情
AI中文摘要

我们重新审视了大语言模型智能体的自动 harness 进化评估。现有 harness 进化方法使用单元测试用例来搜索 harness 配置,然后在相同的公共基准上报告最终性能。此协议引发了两个基本问题。首先,harness 进化本身是一个迭代搜索过程,应在匹配的反馈和推理预算下与简单的任务级搜索基线进行比较。其次,由于搜索和最终评估共享相同的基准,报告的收益可能会过度拟合该特定任务集。为解决这些问题,我们进行了广泛评估,在可比的反馈和推理预算下将 harness 进化与简单的测试时缩放和发现基线进行比较,并在保留任务上评估进化后的 harness 以评估发现的改进是否具有通用性。在使用 GPT - 5.4 和 Claude Opus 4.6 对 Terminal - Bench 2.1 进行的实验表明,自动 harness 进化并不总是优于简单的测试时缩放方法,并且泛化能力有限。我们的结果对自动 harness 进化的有效性提出了重要问题,并强调了对自动 harness 设计采用更公平评估协议和基准的必要性。

英文摘要

We revisit the evaluation of automatic harness evolution for LLM agents. Existing harness evolution methods use unit test cases to search for harness configurations and then report final performance on the same public benchmark. This protocol raises two fundamental concerns. First, harness evolution is itself an iterative search procedure that repeatedly evaluates and revises candidate harnesses using task feedback. As in agentic test-time scaling, it should therefore be compared with simple task-level search baselines under matched feedback and inference budgets to determine whether its gains arise from improved harness design or from additional search alone. Second, because the search and the final evaluation share the same benchmark, the reported gains risk overfitting to that specific task set. To address these concerns, we conduct an extensive evaluation comparing harness evolution with simple test-time scaling and discovery baselines under comparable feedback and inference budgets, and also evaluate evolved harnesses on held-out tasks to assess whether the discovered improvements generalize. Experiments on Terminal-Bench 2.1 with GPT-5.4 and Claude Opus 4.6 show that automatic harness evolution does not consistently outperform simple test-time scaling methods and exhibits limited generalization. Our results raise important questions about the effectiveness of automatic harness evolution and highlight the need for fairer evaluation protocols and benchmarks for automatic harness design. Our code is available at https://github.com/rethinking-harness-evolution.

URL PDF HTML 收藏
2606.24267 2026-07-15 cs.CL cs.AI 版本更新

Pigeonholing: how bad prompts hurt models, causing collapse and mistakes

鸽笼效应:不良提示导致模型崩溃和犯错

Hyunji Nam, Keertana Chidambaram, Dorottya Demszky, Natasha Jaques

机构 * Stanford University(斯坦福大学) University of Washington(华盛顿大学)

AI总结 研究不良上下文导致大语言模型性能下降和模式崩溃的“鸽笼效应”,发现重复错误答案、收敛于狭窄答案集等问题,并提出RLVR合成错误缓解方法。

Comments 10 pages

详情
AI中文摘要

虽然上下文学习通常被证明在大语言模型(LLMs)中有效,但不良上下文可能导致性能下降和模式崩溃,我们称之为“鸽笼效应”。**非故意不良**上下文可能在没有恶意越狱意图的情况下发生:例如,用户要求模型证明一个不正确的数学定理,或未能纠正模型有错误的代码。具体来说,我们在两种场景下研究“鸽笼效应”:(1)当用户提出解决方案时,以及(2)当对话上下文包含助手之前的(错误)回答时。我们在10个可验证和开放式任务上使用10个不同模型进行的实验表明,鸽笼效应以多种方式表现:(1)重复上下文中的错误答案(导致38-40%的性能下降),(2)在编码和文本生成中收敛于狭窄的答案集而不探索替代方案,以及(3)在有争议的话题上转变立场以与用户或助手之前的说法保持一致。我们发现,鸽笼效应几乎随着对话轮次数量的增加而单调恶化(当重复错误从1次增加到5次时,性能额外下降14%以上),并且即使提供的示例是正确的,鸽笼效应诱导的模式崩溃也可能发生。作为缓解的一步,我们提出了带有合成错误的RLVR,与普通RLVR基线相比,在不良上下文下将模型性能提高了43-60%。

英文摘要

While in-context learning is generally shown to be effective in Large Language Models (LLMs), bad contexts can cause performance degradation and mode collapse, a phenomenon we call "pigeonholing." **Unintentionally bad** contexts can happen without malicious jailbreaking intents: For example, a user asks the model to justify an incorrect math theorem or fails to correct the model's buggy code. Specifically, we investigate ``pigeonholing" in two scenarios: (1) when the user suggests a solution, and (2) when the conversation context includes the assistant's previous (incorrect) responses. Our experiments across 10 verifiable and open-ended tasks with 10 different models show that pigeonholing manifests in several ways: (1) repeating the incorrect answers from context (leading to 38-40% performance drop), (2) converging on a narrow set of answers in coding and text generation without exploring alternatives, and (3) flipping stance on controversial topics to align with the user or the assistant's previous claims. We find that pigeonholing worsens almost monotonically with the number of conversation turns (performance drops by additional 14+% as repeated mistakes increase from 1 to 5), and pigeonholing-induced mode collapse can happen even when the provided example is correct. As a step toward mitigation, we propose RLVR with synthetic errors which improves models by 43-60% under bad contexts compared to vanilla RLVR baselines.

URL PDF HTML 收藏
2602.09907 2026-07-15 cs.HC cs.AI cs.CY 版本更新

Self-Regulated Reading with AI Support: An Eight-Week Study with Students

具有AI支持的自主阅读:一项为期八周的学生研究

Yue Fu, Joel Wester, Niels Van Berkel, Alexis Hiniker

机构 * University of Washington(华盛顿大学) University of Copenhagen(哥本哈根大学) Aalborg University(奥胡斯大学)

AI总结 本研究探讨了AI支持下学生自主阅读的认知过程,发现学生在阅读中表现出从理解到推理的认知发展,但受效率驱动,倾向于使用AI生成摘要来筛选阅读内容。

详情
AI中文摘要

大学生越来越多地使用AI聊天机器人来支持学术阅读,但缺乏对这些互动如何塑造其阅读体验和认知参与的细致理解。我们与15名本科生进行了一项为期八周的纵向研究,这些学生在课程中使用AI支持指定阅读。我们收集了838个提示,分布在239次阅读会话中,并开发了一种编码方案,将提示分为四个认知主题:解码、理解、推理和元认知。理解提示占主导地位(59.6%),推理(29.8%)、元认知(8.5%)和解码(2.1%)较少。大多数会话(72%)恰好包含三个提示,即阅读任务的最低要求。在会话中,学生表现出从理解向推理的自然认知进步,但这种进步被截断。在八周内,学生的参与模式保持稳定,但存在显著的个体差异。定性分析揭示了意图-行为的差距:学生认识到有效的提示需要努力,但很少应用这种知识,效率成为主要驱动因素。学生还根据兴趣和学术压力战略性地分配参与,表现出一种新的通过AI阅读而非使用AI的模式:使用AI生成的摘要作为主要材料来筛选哪些部分值得深入关注。我们讨论了AI阅读系统在支撑持续认知参与方面的设计启示。

英文摘要

College students increasingly use AI chatbots to support academic reading, yet we lack granular understanding of how these interactions shape their reading experience and cognitive engagement. We conducted an eight-week longitudinal study with 15 undergraduates who used AI to support assigned readings in a course. We collected 838 prompts across 239 reading sessions and developed a coding schema categorizing prompts into four cognitive themes: Decoding, Comprehension, Reasoning, and Metacognition. Comprehension prompts dominated (59.6%), with Reasoning (29.8%), Metacognition (8.5%), and Decoding (2.1%) less frequent. Most sessions (72%) contained exactly three prompts, the required minimum of the reading assignment. Within sessions, students showed natural cognitive progression from comprehension toward reasoning, but this progression was truncated. Across eight weeks, students' engagement patterns remained stable, with substantial individual differences persisting throughout. Qualitative analysis revealed an intention-behavior gap: students recognized that effective prompting required effort but rarely applied this knowledge, with efficiency emerging as the primary driver. Students also strategically triaged their engagement based on interest and academic pressures, exhibiting a novel pattern of reading through AI rather than with it: using AI-generated summaries as primary material to filter which sections merited deeper attention. We discuss design implications for AI reading systems that scaffold sustained cognitive engagement.

URL PDF HTML 收藏
2312.17670 2026-07-15 cs.CV cs.LG q-bio.QM q-bio.TO 版本更新

The TopCoW Challenge -- Topology-Aware Circle of Willis Segmentation for CT and MR Angiography

TopCoW挑战——用于CT和MR血管造影的拓扑感知Willis环分割

Kaiyuan Yang, Fabio Musio, Yihui Ma, Norman Juchler, Johannes C. Paetzold, Rami Al-Maskari, Luciano Höher, Hongwei Bran Li, Ibrahim Ethem Hamamci, Anjany Sekuboyina, Suprosanna Shit, Houjing Huang, Chinmay Prabhakar, Ezequiel de la Rosa, Bastian Wittmann, Diana Waldmannstetter, Florian Kofler, Fernando Navarro, Martin J. Menten, Ivan Ezhov, Daniel Rueckert, Iris N. Vos, Ynte M. Ruigrok, Birgitta K. Velthuis, Hugo J. Kuijf, Pengcheng Shi, Wei Liu, Ting Ma, Maximilian R. Rokuss, Yannick Kirchhoff, Fabian Isensee, Klaus Maier-Hein, Chengcheng Zhu, Huilin Zhao, Philippe Bijlenga, Julien Hämmerli, Catherine Wurster, Laura Westphal, Jeroen Bisschop, Elisa Colombo, Hakim Baazaoui, Hannah-Lea Handelsmann, Andrew Makmur, James Hallinan, Amrish Soundararajan, Benedikt Wiestler, Jan S. Kirschke, Evamaria O. Riedel, Roland Wiest, Emmanuel Montagnon, Laurent Letourneau-Guillon, Kwanseok Oh, Dahye Lee, Orhun Utku Aydin, Adam Hilbert, Jana Rieger, Dimitrios Rallios, Satoru Tanioka, Alexander Koch, Dietmar Frey, Abdul Qayyum, Moona Mazher, Steven Niederer, Nico Disch, Julius C. Holzschuh, Dominic LaBella, Francesco Galati, Daniele Falcetta, Maria A. Zuluaga, Chaolong Lin, Haoran Zhao, Zehan Zhang, Minghui Zhang, Xin You, Hanxiao Zhang, Guang-Zhong Yang, Yun Gu, Sinyoung Ra, Jongyun Hwang, Hyunjin Park, Junqiang Chen, Marek Wodzinski, Henning Müller, Nesrin Mansouri, Florent Autrusseau, Cansu Yalcin, Rachika E. Hamadache, Clara Lisazo, Joaquim Salvi, Adrià Casamitjana, Xavier Lladó, Uma Maria Lal-Trehan Estrada, Valeriia Abramova, Luca Giancardo, Arnau Oliver, Paula Casademunt, Adrian Galdran, Matteo Delucchi, Oscar Camara, Jialu Liu, Haibin Huang, Yue Cui, Zehang Lin, Yusheng Liu, Shunzhi Zhu, Tatsat R. Patel, Adnan H. Siddiqui, Vincent M. Tutino, Maysam Orouskhani, Huayu Wang, Mahmud Mossa-Basha, Yuki Sato, Sven Hirsch, Susanne Wegener, Bjoern Menze

机构 * Department of Quantitative Biomedicine, University of Zurich, Zurich, Switzerland Institute of Computational Life Sciences, Zurich University of Applied Sciences (ZHAW), Waedenswil, Switzerland Department of Neuroradiology, University Hospital of Zurich, Zurich, Switzerland Department of Neurosurgery, Zhongnan Hospital of Wuhan University, Wuhan, China Department of Radiology at Weill Cornell Medicine, Cornell University, New York, USA Institute for Tissue Engineering School of Computation, Information Technology, Technical University of Munich, Germany Athinoula A. Martinos Center for Biomedical Imaging, Harvard Medical School, Boston, USA School of Medicine Health, TUM Klinikum, Technical University of Munich, Germany Munich Center for Machine Learning, Munich, Germany Department of Computing, Imperial College London, London, UK Image Sciences Institute, UMC Utrecht, Utrecht, The Netherlands Department of Neurology Neurosurgery, University Medical Center Utrecht, Utrecht, The Netherlands Department of Radiology, University Medical Center Utrecht, Utrecht, The Netherlands Electronic \& Information Engineering School, Harbin Institute of Technology (Shenzhen), China Peng Cheng Laboratory, Shenzhen, China Division of Medical Image Computing, German Cancer Research Center (DKFZ), Heidelberg, Germany Faculty of Mathematics Computer Science, Heidelberg University, Germany Helmholtz Imaging, German Cancer Research Center, Heidelberg, Germany Data Science School for Health, Karlsruhe/Heidelberg, Germany Learning Group, Department of Radiation Oncology, Heidelberg University Hospital Department of Radiology, University of Washington, Seattle, WA, USA Department of Radiology, Ren Ji Hospital, Shanghai Jiao Tong University School of Medicine, Shanghai, China Department of Clinical Neurosciences, Division of Neurosurgery, Geneva University Hospitals, Geneva, Switzerland Department of Neurology, University Hospital of Zurich, Zurich, Switzerland Department of Physiology, University of Toronto, Canada Department of Neurosurgery, University Hospital of Zurich, Zurich, Switzerland Department of Diagnostic Imaging, National University Hospital, Singapore University of Chicago, USA Department of Diagnostic Interventional Neuroradiology, University Hospital Berne University of Berne, Berne, Switzerland Centre de Recherche du Centre Hospitalier de l’Université de Montréal (CRCHUM), Montréal, Québec, Canada DEEPNOID Inc., Seoul, South Korea Department of Artificial Intelligence, Korea University, Seoul, South Korea Charité Lab for AI in Medicine (CLAIM), Charité Universitätsmedizin Berlin, Berlin, Germany Lung Institute, Faculty of Medicine, Imperial College London, London, UK Centre for Medical Image Computing, Department of Computer Science, University College London, London, UK Department of Radiation Oncology, Duke University Medical Center, Durham, NC, USA Institute of Medical Technology, Peking University Health Science Center, Beijing, China Hangzhou Genlight MedTech Co., Ltd., China Institute of Medical Robotics, Shanghai Jiao Tong University, Shanghai, China Department of Automation, Shanghai Jiao Tong University, Shanghai, China Department of Artificial Intelligence, Sungkyunkwan University, Seoul, South Korea Department of Electrical Computer Engineering, Sungkyunkwan University, Seoul, South Korea Shanghai MediWorks Precision Instruments Co., Ltd., China Institute of Informatics, HES-SO Valais-Wallis, Switzerland Department of Measurement Electronics, AGH University of Krakow, Poland Laboratoire de Thermique et Energie de Nantes (LTeN), Université Nantes, Polytech’Nantes, Nantes, France Research Institute of Computer Vision Center for Precision Health, McWilliams School of Biomedical Informatics, University of Texas Health Science Center at Houston, USA Physense, BCN-Medtech, Department of Communication Information Technologies, Universitat Pompeu Fabra, Barcelona, Spain Department of Mathematical Modeling Machine Learning, University of Zurich, Zurich, Switzerland Laboratory of Brain Atlas Brain-inspired Intelligence, Institute of Automation, Chinese Academy of Sciences, Beijing, China School of Artificial Intelligence, University of Chinese Academy of Sciences, Beijing, China School of Computer Information Engineering, Xiamen University of Technology, Xiamen, China Vascular Research Center, University at Buffalo, NY, USA Department of Pathology Anatomical Sciences, University at Buffalo, NY, USA Department of Neurosurgery, University at Buffalo, NY, USA LPIXEL Inc., Tokyo, Japan

AI总结 组织TopCoW基准挑战,发布含125对MRA和CTA扫描的注释数据集,参与者提交CoW分割和变体分类算法,经评估,最佳算法在多任务中表现出色,证明CoW分割算法对下游临床应用有可解释性效用。

Comments Summary paper for the TopCoW Challenge: 4 figures, 1 table, and supplementary material in appendix. Accepted for publication in NEJM AI. Datasets and best-performing algorithm Dockers are available at https://zenodo.org/records/15692630 and https://zenodo.org/records/15665435

详情
AI中文摘要

Willis环(CoW)是连接大脑主要循环的重要动脉网络。其血管结构被认为会影响严重神经血管疾病的风险、严重程度和结果。然而,表征高度可变的CoW解剖结构仍然是一项人工且耗时的专家任务。CoW通常通过磁共振血管造影(MRA)和计算机断层血管造影(CTA)这两种非侵入性血管造影成像方式进行成像,但带注释的CoW解剖结构数据集很少,也没有用于比较CoW分割算法的既定基准。我们组织了TopCoW基准挑战,并发布了一个带注释的CoW数据集,其中包含来自同一患者的125对MRA和CTA扫描。使用虚拟现实技术创建了13个血管成分的体素级注释,并由临床专家进行了验证。参与者提交了CoW分割和变体分类算法,我们在包含来自五个以上中心的226次扫描的内部和外部测试集上进行了评估。该基准包括体素级分割、CoW成分检测、CoW变体分类和两个临床应用任务。我们收到了来自六大洲250多名参与者的提交。表现最佳的团队在几乎所有测试集中,CoW分割的Dice分数超过90%,关键血管成分检测的F1分数超过80%,CoW变体分类中的平衡准确率超过70%。最佳算法还通过准确分类胎儿型大脑后动脉并定位与CoW解剖结构相关的动脉瘤,支持了临床相关的下游任务。这个基准证明了CoW分割算法在一些具有可解释性的下游临床应用中的效用。

英文摘要

The Circle of Willis (CoW) is an important network of arteries connecting major circulations of the brain. Its vascular architecture is believed to influence the risk, severity, and outcome of serious neurovascular diseases. However, characterizing the highly variable CoW anatomy remains a manual and time-consuming expert task. The CoW is commonly imaged by two non-invasive angiographic imaging modalities, magnetic resonance angiography (MRA) and computed tomography angiography (CTA), yet few datasets with annotated CoW anatomy exist, and there have been no established benchmarks for comparing CoW segmentation algorithms. We organized the TopCoW benchmark challenge alongside the release of an annotated CoW dataset with 125 paired MRA and CTA scans from the same patients. Voxel-level annotations for 13 vessel components were created using virtual reality technology and verified by clinical experts. Participants submitted algorithms for CoW segmentation and variant classification, which we evaluated on internal and external test sets comprising 226 scans from over five centers. The benchmark includes voxel-level segmentation, CoW component detection, CoW variant classification, and two clinical application tasks. We received submissions from over 250 participants across six continents. Top-performing teams achieved over 90% Dice scores for CoW segmentation, over 80% F1 scores for detecting key vessel components, and over 70% balanced accuracy in CoW variant classification across nearly all test sets. The best algorithms also supported clinically relevant downstream tasks by accurately classifying fetal-type posterior cerebral arteries and localizing aneurysms in relation to CoW anatomy. This benchmark demonstrated the utility of CoW segmentation algorithms for some downstream clinical applications with explainability.

URL PDF HTML 收藏