arXivDaily arXiv每日学术速递 周一至周五更新

高校专区

Northeastern University(东北大学)

至 收录 1090
2607.17548 2026-07-21 cs.HC cs.AI 新提交

Human-in-the-Loop User Feedback Affects Perceived Accuracy and Trust, but Task Subjectivity Matters

人在回路中的用户反馈会影响感知准确性和信任,但任务主观性很重要

Donald R. Honeycutt, Mahsan Nourani, Eric D. Ragan

机构 * University of Florida(佛罗里达大学) Northeastern University(东北大学)

AI总结 研究人在回路的用户反馈对智能系统的影响,通过三项对照实验,发现在客观反馈情境下提供反馈会降低用户对系统的信任和准确性感知,主观反馈则无此负面偏差,强调设计智能系统时考虑用户反馈对信任影响的重要性。

详情
AI中文摘要

虽然机器学习可以生成比人类手动生成更复杂的模型,但纳入人类输入通常可以提高性能。在许多情况下,经常使用系统的最终用户会自然了解其缺陷并希望能改变系统行为。征求最终用户反馈可随时间显著改进模型,但也会影响一些未被充分理解的人为因素。为此进行了三项对照实验,研究交互式反馈收集在客观和主观反馈领域对用户印象的影响。结果表明,在有客观正确答案的情况下,提供人在回路反馈会降低参与者对系统的信任和对系统准确性的感知,而在主观反馈情况下则未观察到这种负面偏差。此外,在客观情境中参与者对系统的不信任随时间增加,而在主观情境中则不然。这些结果凸显了在设计智能系统时考虑不同类型最终用户反馈对用户信任影响的重要性。

英文摘要

While ML can produce complex models beyond those that a human could produce manually, incorporating human input can often improve performance beyond purely data-driven models. While this feedback could come from system designers or domain experts, in many cases, the end users who regularly use the system will naturally develop an understanding of its flaws and desire the ability to change the system's behavior based on their knowledge. While soliciting feedback from end users can result in significant model improvement over time, introducing these feedback techniques can also affect several human factors-such as trust or perception of system accuracy-that are not yet fully understood and have different effects reported in the existing literature. Therefore, we sought to build on the existing research to further explore how the act of providing feedback can affect user understanding of an intelligent system and its accuracy in different contexts. We present three controlled experiments that study the effects of interactive feedback collections on user impressions in domains with objective and subjective feedback. The results show that in a context where there is an objectively correct answer, providing HITL feedback lowered both participants' trust in the system and their perception of system accuracy, regardless of whether the system accuracy improved in response to their feedback. However, when the feedback being provided involved subjective opinion, no such negative bias was observed. Furthermore, in the objective context, participants distrusted the system over time, whereas participants in the subjective context mistrusted the system over time. These results highlight the importance of considering the effects of allowing different types of end-user feedback on user trust when designing intelligent systems.

URL PDF HTML 收藏
2607.17213 2026-07-21 cs.RO 新提交

Retriever: Composing Closed-Loop Asynchronous Robot Programs

Retriever:组合闭环异步机器人程序

Linfeng Zhao, Haojie Huang, Jiayuan Mao, Weiyu Liu, Mykel Kochenderfer, Lawson L. S. Wong

机构 * Stanford University(斯坦福大学) MIT(麻省理工学院) Northeastern University(东北大学)

AI总结 研究构建长期运行机器人智能体的闭环管道问题,提出Retriever,它涵盖异步决策模型等整个堆栈,将智能体表示为有状态因果流函数图,编译到支持多后端的运行时,可系统调试和确定性重放,通过案例研究等进行评估。

Comments Project website: http://retriever.systems; Package open-source website: http://openretriever.org

详情
AI中文摘要

构建长期运行的机器人智能体需要组合闭环管道,其组件运行在不同时钟且延迟可变。当前系统常采用临时并发和发布/订阅约定,导致时间和输入消费语义隐含,行为依赖调度且难重现、调试和复用。现有解决方案多只解决部分问题。本文提出Retriever,涵盖整个堆栈,包括异步决策模型、编程模型、运行时和示例闭环智能体管道。它将智能体表示为在显式运行时钟上执行的有状态因果流函数图,通过连续时间流上的异步环境-智能体循环形式化此观点,表明有限内存因果策略可由这些算子组合表示。Retriever将这些图编译到支持多个后端的运行时,实现跨运行环境的系统调试和从记录的异步数据进行确定性重放。我们通过实际机器人案例研究以及对运行时开销和确定性重放行为的控制研究对Retriever进行评估。

英文摘要

Building long-horizon robot agents requires composing closed-loop pipelines -- perception, belief update, planning, and control -- whose components run at different clocks and with variable latency. Today, these systems are often assembled with ad-hoc concurrency and pub/sub conventions that make timing and input-consumption semantics implicit, yielding schedule-dependent behavior that is hard to reproduce, debug, and reuse. Current solutions typically solve parts of this problem at either the algorithmic or the systems layer, but not both. In this work, we propose Retriever, which spans the entire stack: an asynchronous decision model, a programming model, a runtime, and an example closed-loop agent pipeline. Retriever represents an agent as a graph of stateful causal stream functions executed on explicit run clocks. We formalize this view via an asynchronous environment-agent loop over continuous-time streams and show that finite-memory causal policies can be represented by compositions of these operators. Retriever compiles these graphs into a runtime that supports multiple backends, enabling systematic debugging across running environments and deterministic replay from logged asynchronous data. We evaluate Retriever through a real-robot case study together with controlled studies of runtime overhead and deterministic replay behavior.

URL PDF HTML 收藏
2607.16513 2026-07-21 cs.CY cs.AI cs.HC 新提交

How Formerly Incarcerated People Envision Technologies for Prison Parole

曾经被监禁的人如何设想用于监狱假释的技术

Saiph Savage, Jesse Nava, Wanqing Iris Zhou, Hwijoon Lee

机构 * Northeastern University(东北大学) San Diego State University (SDSU)(圣地亚哥州立大学) Brandeis University(布兰迪斯大学)

AI总结 研究探讨曾被监禁者对假释技术的设想,通过调查发现他们将其视为驾驭权力的资源,而非瓦解监狱系统的工具,进而提出应将假释计算工具从监视转向以人为本、助力驾驭监狱权力结构的系统。

Comments 25 pages, 1 figure. To appear in Proceedings of the 2026 ACM Conference on Fairness, Accountability, and Transparency (FAccT '26)

详情
AI中文摘要

人工智能驱动的算法和自动化工具越来越多地融入惩教领域,影响假释资格、释放决定和监视。这些工具常被视为解决效率和偏见的客观必然方案。但这些计算系统很少有受司法影响个人的参与,可能无法满足被监禁者实际需求。我们调查了31名曾经被监禁的人关于他们的假释经历及对支持假释准备技术的设想。参与者不认为计算工具是瓦解监狱系统的手段,而是驾驭权力的资源,如将复杂假释概念转化为文化上熟悉的术语等。我们认为这些设想指向为战略能动性设计的技术,工具应帮助被监禁者及其家人驾驭现有权力结构追求自由。最后我们提出应将假释的计算工具从监视转向支持人们驾驭监狱权力结构的以人为本的系统。

英文摘要

AI-driven algorithms and automated tools are increasingly embedded in the correctional landscape, shaping parole eligibility,release decisions, and surveillance. These tools are also often framed as objective, inevitable solutions to inefficiency andbias. Yet, these computational systems are rarely designed with input from justice-impacted individuals, which means theymight fail to address the real needs of incarcerated people. To address this gap, we surveyed 31 formerly incarcerated peopleabout their parole experiences and their visions for technologies that could support parole preparation. Contrary to dominantassumptions, participants did not imagine computational tools as instruments to dismantle the prison system, but as resourcesfor navigating power: translating complex parole concepts into culturally familiar terms, documenting personal transformationin board-legible ways, and recognizing the often-invisible labor of families. We argue that these imaginaries point towardtechnologies designed for strategic agency, where tools help incarcerated individuals and their families build the capacity tonavigate existing power structures in pursuit of freedom. We conclude by reframing computational tools for parole awayfrom surveillance and toward human-centered systems that support people in navigating carceral power structures.

URL PDF HTML 收藏
2607.16230 2026-07-21 cs.LG cs.AI 新提交

RouteCost: A Production-Inspired Multi-Stage Framework for Pre-Order Shipping Cost Estimation in E-Commerce

RouteCost:一种受生产启发的多阶段框架,用于电子商务中的预订单运输成本估计

Xianling Zeng, Zihan Yu, Sichen Zhao, Yalun Qi, Zhiming Xue

机构 * Northeastern University(东北大学)

AI总结 研究电子商务中预订单运输成本估计问题,提出受生产启发的多阶段框架RouteCost,将其分解为多步骤,通过路线加权期望公式汇总成本估计,在大量订单等数据上提高了预测质量和校准,保留了路线级可解释性。

详情
AI中文摘要

在电子商务中,准确的预订单运输成本估计很重要,因为它会影响价格展示、利润规划和转化率。实际中,运输成本不仅受距离影响,还受目的地需求组合、计费重量、尺寸定价、附加费触发因素以及诸如货物合并等潜在运营影响。静态查找方法会遗漏重要的变化来源,而整体回归器可能利用强但非因果的相关性。我们提出了RouteCost,这是一个受生产启发的多阶段框架,将问题分解为时间感知需求预测、费用卡告知的基线定价、第二阶段残差校正和基于代理的箱式合并推理。通过路线加权期望公式汇总路线级成本估计,以生成产品级运输成本预测。在超过250,000个订单、260种产品和18个月的订单历史数据上,该框架提高了预测质量和总体校准,同时保留了路线级的可解释性。

英文摘要

Accurate pre-order shipping cost estimation is important in e-commerce because it affects price presentation, margin planning, and conversion. In practice, shipping cost is shaped not only by distance but also by destination demand mix, billable weight, dimensional pricing, surcharge triggers, and latent operational effects such as shipment consolidation. Static lookup methods therefore miss important sources of variation, while monolithic regressors may exploit strong but non-causal correlations. We propose RouteCost, a production-inspired multi-stage framework that decomposes the problem into time-aware demand forecasting, fee-card-informed baseline pricing, Stage 2 residual correction, and proxy-based box-consolidation inference. Route-level cost estimates are aggregated through a route-weighted expectation formulation to produce product-level shipping cost predictions. Across over 250,000 orders, 260 products, and 18 months of order history, the framework improves predictive quality and aggregate calibration while preserving route-level interpretability.

URL PDF HTML 收藏
2607.16197 2026-07-21 cs.AI 新提交

Some Large Language Models Exhibit Consistent Risk Attitudes

一些大语言模型表现出一致的风险态度

Bowen Sun, Rui Min, Yuxi Wang, Brian Odegaard, Qi Wang, Jing Du

机构 * University of Florida(佛罗里达大学) Northeastern University(东北大学)

AI总结 研究大语言模型在不确定性下的风险态度,引入跨域框架,应用于六个大语言模型和100名人类参与者,发现多数大语言模型有任务内一致性、跨域排序稳定性,且风险态度分布趋向受限,为评估和调整AI系统奠定基础。

Comments 38

详情
AI中文摘要

随着人工智能系统在开放式、高风险环境中部署,一个关键维度仍未得到衡量:感知风险如何转化为行动。我们测试大语言模型在不确定性下是否表现出系统且一致的风险态度。我们引入一个跨域框架,将情境风险信念与分类决策解耦,并将其应用于六个代表性大语言模型和100名人类参与者,涉及空间导航、临床分诊和财务分配任务。使用回归模型,我们提取每个智能体的信念到决策映射,并量化风险敏感性和风险态度偏差。我们发现大多数测试的大语言模型表现出:(i)强大的任务内一致性,即在固定任务域内从情境信念到风险决策的稳定映射;(ii)跨域排序稳定性,在不同任务中保持相对风险态势;(iii)相对于更广泛的人类基线,趋向于受限的风险态度分布。这些结果揭示了风险态度是大语言模型行为中一个稳定且先前未被表征的维度,为评估和调整开放式决策中的人工智能系统奠定了基础,并激发了对这些内在行为倾向起源的进一步研究。

英文摘要

As artificial intelligence systems are deployed in open-ended, high-stakes settings, a critical dimension remains unmeasured: how perceived risk is translated into action. We test whether large language models (LLMs) exhibit systematic and consistent risk attitudes under uncertainty. We introduce a cross-domain framework that decouples contextual risk belief from categorical decision, and apply it to six representative LLMs and 100 human participants across spatial navigation, clinical triage, and financial allocation tasks. Using regression models, we extract each agents belief-to-decision mapping and quantify risk sensitivity and risk attitude bias. We find that most tested LLMs exhibit (i) robust intra-task consistency, indicating stable mappings from contextual belief to risk decision within a fixed task domain; (ii) cross-domain rank-order stability, preserving relative risk posture across tasks; and (iii) a convergence toward a restricted risk-attitude distribution relative to the broader human baseline. These results reveal risk attitude as a stable and previously uncharacterized dimension of LLM behavior, establishing a foundation for evaluating and aligning AI systems in open-ended decision-making and motivating further investigation into the origins of these intrinsic behavioral dispositions.

URL PDF HTML 收藏
2607.16238 2026-07-21 cs.LG cs.AI physics.comp-ph physics.flu-dyn 新提交

Diffusion-corrected Autoregressive Fourier Neural Operator for Droplet Evolution Prediction

用于液滴演化预测的扩散校正自回归傅里叶神经算子

Jinghao Cao, Minsung Kang, Hongyue Sun, Chi Zhou, Jihoon Chung, Xubo Yue, Sanchoy Das, Bo Shen

机构 * New Jersey Institute of Technology(新泽西理工学院) University of Georgia(佐治亚大学) University at Buffalo(纽约州立大学布法罗分校) Hanyang University(汉阳大学) Northeastern University(东北大学)

AI总结 研究材料喷射中液滴演化预测难题,提出DiffARFNO两阶段框架,结合自回归傅里叶 - MIONet与DDIM校正器,经实验验证该方法显著优于现有模型,能为长期预测提供高保真结果。

Comments 11 figures, 4 tables

详情
AI中文摘要

预测材料喷射(即喷墨打印,IJP)中的液滴演化对于维持打印质量至关重要。然而,由于误差累积和过程变量的复杂耦合,长期预测仍然具有挑战性。在这项工作中,我们引入了扩散校正自回归傅里叶神经算子(DiffARFNO),这是一个两阶段框架,它将自回归傅里叶 - MIONet与条件去噪扩散隐式模型(DDIM)校正器相结合。傅里叶 - MIONet被训练为粗预测器并自回归地用于长期预测。在第二阶段,基于DDIM的条件校正器通过有效的迭代去噪在每个滑动窗口内细化粗预测。通过将傅里叶 - MIONet的粗预测与恢复精细细节的DDIM校正器相结合,DiffARFNO旨在为长期预测提供高保真预测。在ANSYS Fluent的液滴数据集上进行的大量实验表明,DiffARFNO明显优于现有的最先进模型。

英文摘要

Predicting droplet evolution in material jetting, or Inkjet Printing (IJP), is essential for maintaining printing quality. However, long-horizon forecasts remain challenging due to error accumulation and the complex coupling of process variables. In this work, we introduce the Diffusion-corrected Auto-Regressive Fourier Neural Operator (DiffARFNO), a two-stage framework that combines an autoregressive Fourier-MIONet with a conditional Denoising Diffusion Implicit Model (DDIM) corrector. Fourier-MIONet is trained as a coarse predictor and deployed autoregressively for long-horizon forecasting. In the second stage, a DDIM-based conditional corrector refines the coarse prediction within each sliding window through efficient iterative denoising. By combining coarse predictions from Fourier-MIONet with a DDIM corrector that restores fine details, DiffARFNO aims to provide high-fidelity predictions for long-horizon forecasts. Extensive experiments on droplet datasets from ANSYS Fluent demonstrate that DiffARFNO significantly outperforms existing state-of-the-art models.

URL PDF HTML 收藏
2607.14407 2026-07-21 cs.IT cs.AI cs.LG math.IT 版本更新

Decision Making Needs Uncertainty Quantification [Lecture Notes]

决策需要不确定性量化[讲义笔记]

Osvaldo Simeone

机构 * Institute for Intelligent Networked Systems, Northeastern University London(智能网络系统研究所,伦敦大学东北学院)

AI总结 研究决策中不确定性量化问题,从基本原理出发,在已知和未知环境分布下,探讨风险中性与厌恶型智能体所需不确定性表示形式,确定解决认知不确定性的三种互补方法,强调可靠决策需匹配不确定性表示及效用保证。

详情
AI中文摘要

许多信号处理系统最终目的是行动。当决策者(或智能体)采取行动所依据的状态变量不确定时,不确定性的表示方式决定了智能体的表现及表现的可信度。本讲义笔记从基本原理出发,在单一决策理论框架内,阐述了智能体的目标、知识与足以实现最优行动的不确定性表示形式之间的联系。首先,在已知环境分布下,风险中性智能体需要状态的后验分布,风险厌恶智能体可依赖预测集和最坏情况决策规则且不失最优性。接着探讨环境未知的情况,确定了三种解决认知不确定性的互补方法:固定预测器的校准、基于分布鲁棒优化的可信(模糊)集以及对模型参数的贝叶斯推断。共同要点是可靠决策需要与决策目标和智能体知识概况相匹配的不确定性表示,以及对智能体实际获得效用的保证。

英文摘要

Many signal processing systems ultimately exist to {act}. Whenever the state variable that determines the action to be taken by a decision maker, or agent, is uncertain, the way that uncertainty is represented decides how well the agent performs and how much its performance can be trusted. This lecture note develops, from first principles and within a single decision-theoretic setting, the link between the {objective} and the knowledge of an agent and the form of uncertainty representation that is sufficient to act optimally. To start, assuming a known environment distribution, we show that a risk-neutral agent needs the posterior distribution over the state, whereas a risk-averse agent can rely without loss of optimality on a {prediction set} and a worst-case decision rule. We then turn to the case in which the environment is unknown, and identify three complementary approaches to address the resulting epistemic uncertainty: calibration of a fixed predictor, credal (ambiguity) sets with distributionally robust optimization, and Bayesian inference over model parameters. The common thread is that reliable decisions require an uncertainty representation matched to the decision objective and to the knowledge profile of the agent, together with a guarantee that certifies the utility the agent will actually obtain.

URL PDF HTML 收藏
2607.06623 2026-07-21 cs.LG cs.AI cs.SY eess.SY 版本更新

LLM-Guided Task-Semantic Field Factorization for Industrial Process Forecasting

用于工业过程预测的大语言模型引导的任务语义场分解

Youcheng Zong, Runda Jia, Mingxuan Ren, Dakuo He

机构 * College of Information Science and Engineering, Northeastern University(东北大学信息科学与工程学院)

AI总结 针对工业过程预测中标记数据稀缺等问题,提出大语言模型引导的任务语义场分解框架TSF,通过构建任务语义场,结合传统时间序列主干进行训练和推理,在多任务上降低平均绝对误差,增加参数少且推理开销小。

详情
AI中文摘要

流程工业依赖时间序列预测和软传感来估计难以在线测量的质量变量。标记数据稀缺,操作模式频繁变化,为每种情况重新训练模型或重建对齐管道成本高昂。本文提出了任务语义场分解(TSF),这是一个大语言模型引导的框架。TSF在训练前从任务协议和变量文档构建任务语义场,仅将大语言模型用于离线语义构建。在线训练和推理仍使用传统时间序列主干。在多个复杂工业预测和软传感任务上,TSF在改进设置下平均将平均绝对误差降低6.4%,最大降幅达25.5%。它仅增加约1800 - 3000个参数,额外在线推理开销小于0.008毫秒/步。这些结果表明TSF将现有过程文档转化为跨主干和语义生成器的可测量预测增益,同时保持轻量级以便部署。

英文摘要

Process industries rely on time-series forecasting and soft sensing to estimate quality variables that are hard to measure online. Labeled data are scarce, operating regimes change frequently, and retraining models or rebuilding alignment pipelines for each scenario is costly. Such settings often provide variable tables and process documents that record variable names, units, physical meanings, and process roles. However, standard time-series backbones usually treat inputs as anonymous numerical columns. Existing text-enhanced methods also rarely make the semantic-logical relations between input variables and the prediction target available to the model within each numerical window. To address this problem, this article proposes Task-Semantic Field Factorization (TSF), a large language model (LLM)-guided framework. TSF builds a task-semantic field from task protocols and variable documents before training and uses the LLM only for offline semantic construction. Online training and inference are handled by conventional time-series backbones. During training and inference, the current numerical window activates variable semantics, so semantic information participates in each prediction and supports adaptation to different prediction targets and operating shifts. Across multiple complex industrial forecasting and delayed soft-sensing tasks, TSF reduces MAE by 3.6\% on average. Across all dataset--backbone pairs, the macro-average reduction is 2.9\%, with a maximum reduction of 24.9\%. It adds only about 0.7--4.3k parameters, with less than 8\,$μ$s/sample of additional online inference overhead. These results show that TSF turns existing process documents into measurable forecasting gains across backbones and semantic generators while remaining lightweight for deployment.

URL PDF HTML 收藏
2604.08879 2026-07-21 cs.CL 版本更新

GRASP: Grounded CoT Reasoning with Dual-Stage Optimization for Multimodal Sarcasm Target Identification

GRASP:基于双阶段优化的 grounded CoT 推理用于多模态讽刺目标识别

Faxian Wan, Xiaocui Yang, Yifan Cao, Shi Feng, Daling Wang, Yifei Zhang

机构 * School of Computer Science and Engineering, Northeastern University(东北大学计算机科学与工程学院)

AI总结 GRASP通过整合视觉定位与显式Co-T推理,解决多模态讽刺目标识别中的细粒度定位问题,提出MSTI-MAX数据集和双阶段联合优化策略,实验表明其在多模态讽刺目标识别中表现更优。

详情
AI中文摘要

超越传统二元分类范式,多模态讽刺目标识别(MSTI)要求精确定位细粒度目标如文本短语和视觉区域。现有方法依赖隐式跨模态对齐,可解释性差且细粒度定位效果有限。为解决这些问题,我们提出GRASP,一种结合视觉定位与显式Chain-of-Thought(CoT)推理的框架,以突破黑箱MSTI。我们构建了MSTI-MAX数据集,缓解类别不平衡并丰富多模态讽刺线索。我们引入Grounded CoT推理,将讽刺相关视觉区域显式锚定在推理轨迹中,并促使模型在预测最终分类标签和讽刺目标前阐述推理过程。此外,我们采用双阶段结果监督联合优化策略:具有坐标感知加权损失的监督微调,随后是细粒度目标策略优化。大量实验表明,GRASP在多模态讽刺目标识别中优于现有基线,且LLM-as-a-Judge评估量化了内部推理链的质量。我们的数据集和源代码将在GitHub上发布。

英文摘要

Moving beyond the traditional binary classification paradigm of Multimodal Sarcasm Detection, Multimodal Sarcasm Target Identification (MSTI) presents a more formidable challenge, requiring precise localization of fine-grained targets such as textual phrases and visual regions. Existing approaches predominantly rely on implicit cross-modal alignment, offering limited interpretability and suboptimal fine-grained localization. To address these limitations, we propose GRASP, Grounded Chain-of-Thought ReAsoning with Dual-Stage Optimization for Multimodal Sarcasm Prediction and Target Identification, a framework that integrates visual grounding with explicit Chain-of-Thought (CoT) reasoning to move beyond black-box MSTI. Specifically, we curate MSTI-MAX, a refined dataset that mitigates class imbalance and enriches multimodal sarcasm cues. We introduce Grounded CoT reasoning, which explicitly anchors sarcasm-related visual regions within the reasoning trajectory and prompts the model to articulate rationales before predicting the final classification labels and sarcasm targets. Furthermore, we employ a dual-stage outcome-supervised joint optimization strategy: Supervised Fine-Tuning with a coordinate-aware weighted loss, followed by Fine-Grained Target Policy Optimization. Extensive experiments demonstrate that GRASP outperforms existing baselines in fine-grained sarcasm target identification across modalities, and an LLM-as-a-Judge evaluation quantitatively measures the quality of internal reasoning chains. Our dataset and source code will be released on GitHub.

URL PDF HTML 收藏
2604.02142 2026-07-21 cs.RO cs.MA 版本更新

PRO-SPECT: Probabilistically Safe Scalable Planning for Energy-Aware Coordinated UAV-UGV Teams in Stochastic Environments

PRO-SPECT:面向能量感知的协调无人机-无人地面车辆团队在随机环境中的概率安全可扩展规划

Roger Fowler, Cahit Ikbal Er, Benjamin Johnsenberg, Yasin Yazicioglu

机构 * Khoury Department of Computer Sciences at Northeastern University(东北大学Khoury计算机科学系) The Charles Stark Draper Laboratory, Inc.(查尔斯·斯塔克·德雷珀实验室公司) Department of Mechanical and Industrial Engineering at Northeastern University(东北大学机械与工业工程系) Departments of Mechanical and Industrial Engineering and Electrical and Computer Engineering at Northeastern University(东北大学机械与工业工程系和电气与计算机工程系)

AI总结 本文提出PRO-SPECT算法,用于在随机环境中实现无人机和无人地面车辆团队的能量感知规划,通过概率安全约束确保任务成功,支持离线和在线重新规划。

Comments Accepted to the IEEE/RSJ International Conference on Intelligent Robots and Systems 2026 (IROS 2026)

详情
AI中文摘要

我们考虑在随机环境中无人机和无人地面车辆团队的能量感知规划问题。无人机必须在最小时间内访问一组空点,同时遵守能量约束,依赖无人地面车辆作为移动充电站。与之前假设确定性旅行时间或使用固定鲁棒性边界不同,我们模型旅行时间作为随机变量,并将整个任务中失败(能量耗尽)的概率限制在用户指定的风险水平内。我们将问题建模为混合整数规划问题,并提出PRO-SPECT算法,一种多项式时间算法,生成风险受限的计划。该算法支持离线规划和在线重新规划,使团队能够适应干扰同时保持风险界限。我们提供了关于解决方案可行性和时间复杂度的理论结果。我们还通过数值比较和模拟展示了方法的性能。

英文摘要

We consider energy-aware planning for an unmanned aerial vehicle (UAV) and unmanned ground vehicle (UGV) team operating in a stochastic environment. The UAV must visit a set of air points in minimum time while respecting energy constraints, relying on the UGV as a mobile charging station. Unlike prior work that assumed deterministic travel times or used fixed robustness margins, we model travel times as random variables and bound the probability of failure (energy depletion) across the entire mission to a user-specified risk level. We formulate the problem as a Mixed-Integer Program and propose PRO-SPECT, a polynomial-time algorithm that generates risk-bounded plans. The algorithm supports both offline planning and online re-planning, enabling the team to adapt to disturbances while preserving the risk bound. We provide theoretical results on solution feasibility and time complexity. We also demonstrate the performance of our method via numerical comparisons and simulations.

URL PDF HTML 收藏
2607.15970 2026-07-20 cs.CR cs.LG 新提交

Code-Poisoning Property Inference Attacks

代码中毒属性推理攻击

Xukun Luan, Yuhui Gong, Gang Zhang, Zixuan Huang, Yuanguo Bi, Xuesong Li, Jinyan Liu

机构 * School of Computer Science and Technology, Beijing Institute of Technology(计算机科学与技术学院,北京理工大学) School of Computer Science and Engineering, Northeastern University(计算机科学与工程学院,东北大学)

AI总结 研究针对机器学习模型训练数据隐私泄露问题,提出代码中毒属性推理攻击(CPPIA),克服现有工作局限。通过恶意代码提供者使数据持有者下载中毒代码训练模型,对手借此嵌入属性查询模型泄露隐私,该方法攻击准确率高、计算轻量,评估证明其通用性和有效性。

详情
AI中文摘要

蓬勃发展的代码托管平台和编码代理使初学者即使拥有私有数据也能利用现有代码快速构建定制机器学习(ML)模型。ML模型的训练数据常被视为私有财产,面临信息泄露风险。属性推理攻击(PIA)旨在暴露训练集的全局属性信息。本文提出代码中毒属性推理攻击(CPPIA),克服了现有工作的四个局限。考虑恶意代码提供者,数据持有者下载中毒代码后用私有数据训练模型并向公众发布仅标签的API,对手在训练时将属性嵌入秘密样本并随后查询训练模型来泄露隐私。CPPIA攻击准确率达100%且不降低模型准确率,计算轻量级且无需影子模型。通过四个数据集、八个模型架构、十八个属性及三种防御机制评估了攻击性能,证明了CPPIA的通用性和有效性。

英文摘要

The flourishing code hosting platforms and coding agents enable even beginners with private data to build tailored Machine Learning (ML) models using available code quickly. The training data for ML models, often regarded as private property (e.g., clinical records, transaction information), is at significant risk of information leakage. Property Inference Attacks (PIAs), as a significant type of privacy attack, aim to expose global property information of the training set. In this paper, we present Code-Poisoning Property Inference Attack (CPPIA), the first code-level PIA, which overcomes four limitations of existing works: insufficient attack performance, severe degradation of model accuracy, high computational overhead, and failure under defenses. We consider malicious code providers from code hosting platforms (GitHub) and coding agents (Codex). Upon downloading the poisoned code, data holders train models with their private data without professional auditing, subsequently releasing label-only APIs to the public. The adversary embeds the properties into secret samples during training and queries the trained model on these samples later to leak privacy. CPPIA offers 100\% attack accuracy without degrading model accuracy. It is also computationally lightweight and requires no shadow models. We evaluate the attack performance across four datasets, eight model architectures, eighteen properties, and under three defense mechanisms, demonstrating the universality and effectiveness of CPPIA.

URL PDF HTML 收藏
2605.26772 2026-07-20 cs.AI cs.LG 交叉投稿

Beyond a Single Direction: Chain-of-Thought Disrupts Simple Steering of Refusal

超越单一方向:思维链破坏简单的拒绝引导

Kia-Jüng Yang, Dominik Meier, Jiachen Zhao, Terry Ruas, Bela Gipp

机构 * University of Göttingen, Germany(哥廷根大学,德国) Northeastern University, Boston, MA, USA(东北大学,波士顿,马萨诸塞州,美国)

AI总结 本文研究大型推理模型(LRM)中拒绝行为的机制,发现思维链(CoT)与激活共同编码拒绝信号,使得仅通过激活引导难以逆转拒绝,但通过两阶段干预(激活引导下重新生成CoT)可显著提高逆转率。

详情
AI中文摘要

大型推理模型(LRM)在生成最终输出之前会生成思维链(CoT)轨迹,引入动态内部状态,可能使拒绝等控制机制复杂化。与指令调优的LLM不同,后者的拒绝由单一方向子空间介导,而LRM中的拒绝还依赖于CoT。在DeepSeek-R1-Distill-LLaMA-8B中,当CoT保持不变时,激活引导仅在39%的情况下逆转拒绝,但完全移除CoT可将此比例提高到70%,表明CoT积极强化拒绝。在两阶段干预中,模型在激活引导下重新生成其CoT,拒绝在94%的情况下被逆转,而即使移除引导,生成的CoT本身仍保留48%的效果。这表明CoT可以独立携带和重建顺从信号。这些发现表明,LRM中的拒绝由残差流激活和CoT共同编码。这种联合编码使得LRM对仅激活层面的干预更具鲁棒性,但使CoT暴露于可能的替代表面攻击。

英文摘要

Large reasoning models (LRMs) generate chain-of-thought (CoT) traces before producing final outputs, introducing a dynamic internal state that may complicate control mechanisms such as refusal. Unlike instruction-tuned LLMs, where refusal is mediated by a single directional subspace, refusal in large reasoning models (LRMs) additionally depends on the CoT. In DeepSeek-R1-Distill-LLaMA-8B, activation steering reverses refusal in only 39% of cases when the CoT is kept fixed, but removing the CoT entirely increases this to 70%, indicating that the CoT actively reinforces refusal. In a two-stage intervention where the model regenerates its CoT under activation steering, refusal is reversed in 94% of cases, while the resulting CoT alone retains 48% of this effect even after steering is removed. This suggests that the CoT can carry and reconstruct the compliance signal independently. These findings indicate that refusal in LRMs is jointly encoded in residual stream activations and CoT. This joint activation makes LRM more robust against activation-level interventions alone, but exposes CoT to a possible alternative surface attack.

URL PDF HTML 收藏
2604.03984 2026-07-20 cs.CV 版本更新

High-Fidelity Mural Restoration via a Unified Hybrid Mask-Aware Transformer

高保真壁画修复的统一混合掩码感知变换器

Jincheng Jiang, Qianhao Han, Chi Zhang, Zheng Zheng

机构 * Northeastern University (Toronto Campus)(东北大学(多伦多校区)) The Bishop Strachan School(主教斯特拉坎学校)

AI总结 本文提出混合掩码感知变换器HMAT,通过结合掩码感知动态过滤与Transformer瓶颈,实现壁画的高保真修复,同时通过掩码条件风格融合模块和教师强迫解码器提升修复效果。

Comments 17 pages, 3 figures

详情
AI中文摘要

古代壁画是重要的文化遗存,但许多因环境暴露、材料老化和人类活动而严重退化。修复这些艺术品极具挑战性,需重建缺失结构并严格保护未受损区域。本文提出了混合掩码感知变换器(HMAT),一个统一的高保真壁画修复框架。HMAT整合了掩码感知动态过滤用于稳健的局部纹理建模,以及Transformer瓶颈用于长距离结构推断。为进一步处理退化形态的多样性,我们引入了掩码条件风格融合模块,动态指导生成过程。此外,设计了教师强迫解码器与硬门控跳跃连接,以在有效区域强制保真度并在缺失区域聚焦重建。我们在DHMural数据集和精选的九色 Deer 数据集上评估HMAT,在不同退化水平下进行测试。实验结果表明,所提方法在与最先进方法相比时表现竞争,同时产生更具结构一致性和视觉忠实性的修复结果。这些发现表明,HMAT为文化遗存壁画的数字修复提供了有效解决方案。

英文摘要

Ancient murals are valuable cultural artifacts, but many have suffered severe degradation due to environmental exposure, material aging, and human activity. Restoring these artworks is challenging because it requires both reconstructing large missing structures and preserving authentic, undamaged regions. We present the Hybrid Mask-Aware Transformer (HMAT), a unified framework for high-fidelity mural restoration that addresses both structural completion and authentic-region preservation. HMAT integrates Mask-Aware Dynamic Filtering for robust local texture modeling with a Transformer bottleneck for long-range structural inference, enabling recovery of continuous line patterns and coherent mural structures under irregular damage. To handle diverse degradation morphologies, we introduce a mask-conditional style fusion module that adapts the generative process according to the shape and extent of missing regions. We also propose a fidelity-oriented training objective that combines hole-normalized reconstruction, discriminator feature matching, and high-receptive-field perceptual supervision to improve damaged-region fidelity, texture consistency, and boundary quality. In addition, we analyze a Teacher-Forcing Decoder with hard-gated skip connections as a feature-space boundary-conditioning strategy. Experiments show that HMAT matches or outperforms representative convolutional, transformer-based, and edge-guided inpainting baselines, with especially strong gains in perceptual realism and severe-mask settings. Ablation studies further identify the proposed objective, MADF-based mask-aware encoding, and mask-conditioned synthesis as the main contributors to restoration quality. These results demonstrate that HMAT provides an effective and competitive solution for cultural heritage mural restoration.

URL PDF HTML 收藏
2603.22507 2026-07-20 cs.RO cs.MA 版本更新

Energy-Aware Collaborative Exploration for a UAV-UGV Team

面向能耗的无人机-地面车辆协同探索

Cahit Ikbal Er, Saikiran Juttu, Yasin Yazicioglu

机构 * Department of Mechanical and Industrial Engineering at Northeastern University(东北大学机械与工业工程系) Departments of Mechanical and Industrial Engineering and Electrical and Computer Engineering at Northeastern University(东北大学机械与工业工程系和电气与计算机工程系)

AI总结 本文提出一种面向能耗的无人机-地面车辆协同探索框架,通过密度感知分层概率路标实现稀疏耦合空地路标,将路线选择建模为耦合导向问题以最大化信息增益,同时满足会合约束。

Comments Accepted to the IEEE/RSJ International Conference on Intelligent Robots and Systems 2026 (IROS 2026)

详情
AI中文摘要

我们提出了一种面向能耗的无人机-地面车辆协同探索框架,用于在未知环境中操作的无人机-地面车辆团队。其中,无人机的能耗约束被建模为最大飞行时间限制。无人机执行一系列受能耗限制的探索路线,同时地面车辆在地面进行探索并充当移动充电站。在共享时间预算下强制会合,以确保车辆在每次路线结束前在无人机达到飞行时间限制前相遇。我们使用密度感知分层概率路标(PRM)构建稀疏耦合空地路标,并将路线选择建模为耦合导向问题(OPs),以在满足会合约束的情况下最大化信息增益。所生成的路线是在碰撞验证的路标边路上构建的。我们通过模拟研究、基准比较和实际实验验证了我们的方法。

英文摘要

We present an energy-aware collaborative exploration framework for a UAV-UGV team operating in unknown environments, where the UAV's energy constraint is modeled as a maximum flight-time limit. The UAV executes a sequence of energy-bounded exploration tours, while the UGV simultaneously explores on the ground and serves as a mobile charging station. Rendezvous is enforced under a shared time budget so that the vehicles meet at the end of each tour before the UAV reaches its flight-time limit. We construct a sparsely coupled air-ground roadmap using a density-aware layered probabilistic roadmap (PRM) and formulate tour selection over the roadmap as coupled orienteering problems (OPs) to maximize information gain subject to the rendezvous constraint. The resulting tours are constructed over collision-validated roadmap edges. We validate our method through simulation studies, benchmark comparisons, and real-world experiments.

URL PDF HTML 收藏
2409.10897 2026-07-20 cs.LG cs.SE 版本更新

AutoSpec: Automated Generation of Neural Network Specifications

AutoSpec:神经网络规范的自动生成

Shuowei Jin, Taobo Liao, Anuj Kalia, Xenofon Foukas, Huan Zhang, Cheng Tan, Z. Morley Mao, Francis Y. Yan

机构 * University of Michigan(密歇根大学) Microsoft Research(微软研究院) Northeastern University(东北大学) University of Illinois Urbana-Champaign(伊利诺伊大学厄巴纳-香槟分校)

AI总结 针对神经网络在学习增强系统中对模型安全鲁棒性需求,AutoSpec提出首个自动生成和评估神经网络规范的综合框架,通过基于树算法、统计认证框架及评估框架,经实验验证其性能优于手动定义规范和现有基线算法。

详情
AI中文摘要

神经网络在学习增强系统中的日益广泛应用凸显了对模型安全性和鲁棒性的需求,尤其是在安全关键领域。虽然神经网络验证的最新进展提供了最坏情况行为的形式保证,但现有方法要求用户手动定义模型规范,这一过程容易出错、不完整且耗时。本文提出了AutoSpec,这是首个用于为学习增强系统自动生成和评估神经网络规范的综合框架。AutoSpec引入了一种基于树的算法,该算法自适应地划分输入空间以生成与模型行为一致的规范集,以及一个统计认证框架,为每个规范提供严格的准确性保证。我们还提出了一个有原则的评估框架,该框架定义了规范准确性和覆盖范围的可解释指标,为未来研究建立了基准。在四个不同应用中的实验表明,AutoSpec优于手动定义的规范和现有的基线算法,比人工定义的规范提高F1分数高达53%,比最强基线提高73%。

英文摘要

The increasing adoption of neural networks in learning-augmented systems highlights the growing need for model safety and robustness, especially in safety-critical domains. While recent advances in neural network verification offer formal guarantees on worst-case behavior, existing approaches require users to manually define model specifications, an error-prone, incomplete, and time-consuming process. In this paper, we present AutoSpec, the first comprehensive framework for automatically generating and evaluating neural network specifications for learning-augmented systems. AutoSpec introduces a tree-based algorithm that adaptively partitions the input space to generate specification sets aligned with model behavior, as well as a statistical certification framework that provides rigorous accuracy guarantees for each specification. We also propose a principled evaluation framework that defines interpretable metrics for specification accuracy and coverage, establishing a benchmark for future research. Experiments across four diverse applications show that AutoSpec outperforms both manually defined specifications and existing baseline algorithms, improving the F1 score by up to 53% over human-defined specifications and 73% over the strongest baseline.

URL PDF HTML 收藏
2607.15180 2026-07-17 cs.LG cs.SY eess.SY 新提交

RTS Smoother-Guided Learning of Physics-Based Neural Differential Models

基于RTS平滑器引导的物理神经网络微分模型学习

Ahmet Demirkaya, Georgios Stratis, Tales Imbiriba, Zachary D. Danziger, Deniz Erdogmus

机构 * Northeastern University(东北大学) University of Massachusetts Boston(马萨诸塞大学波士顿分校) Emory University(埃默里大学)

AI总结 针对部分状态变量可测、动力学方程部分未知的情况,提出混合神经-物理框架,交替进行状态和参数估计,利用RTS平滑器和反向传播,能从测量中学习缺失的ODE组件,提升潜在状态重建和长期预测能力。

详情
AI中文摘要

常微分方程(ODEs)广泛用于物理、生物、神经科学和生理学中的动力系统建模,但在许多应用中,动力学的一些方程未知,只有部分状态变量可测量。我们提出了一种混合神经-物理框架,其中ODE的已知部分保持显式,缺失部分由神经网络表示。该方法包括两个阶段,在状态估计和参数估计之间交替迭代,直到满足预定标准。具体而言,第一步,将模型参数视为已知,使用Rauch-Tung-Striebel(RTS)平滑器从可用测量中推断潜在状态;第二步,将平滑后的轨迹视为已知,通过反向传播估计神经网络参数。我们在部分状态观测下的线性、非线性和刚性动力学的基准系统上评估了该方法。在这些设置中,该方法从不完整测量中学习缺失的ODE组件,同时利用并保留可解释的机制结构,改善潜在状态重建和长期预测。

英文摘要

Ordinary differential equations (ODEs) are widely used to model dynamical systems in physics, biology, neuroscience, and physiology, but in many applications some equations of the dynamics are unknown and only a subset of the state variables are measured. We propose a hybrid neural--physics framework in which the known components of the ODE are kept explicit and the missing components are represented by a neural network. The proposed method consists of two stages where we alternate between state and parameter estimation and iterate until a predetermined criterion is met. Specifically, in the first step, we treat the model parameters as being known and we infer the latent states from the available measurements using a Rauch--Tung--Striebel (RTS) smoother. In the second stage, we treat the smoothed trajectories as being known and use them to estimate the neural networks' parameters through backpropagation. We evaluate the method on benchmark systems spanning linear, nonlinear, and stiff dynamics under partial state observation. Across these settings, the proposed method learns missing ODE components from incomplete measurements while exploiting and retaining interpretable mechanistic structure and improving latent-state reconstruction and long-horizon prediction.

URL PDF HTML 收藏
2607.14499 2026-07-17 cs.AI 新提交

Contextualized Evaluation of Vision Language Models through Dynamic, Multi-turn Interactions

通过动态多轮交互对视觉语言模型进行情境化评估

Yijiang Li, Huiqi Zou, Bingyang Wang, Ziang Xiao

机构 * UC San Diego(加州大学圣地亚哥分校) Northeastern University(东北大学) Georgia Institute of Technology(佐治亚理工学院) Johns Hopkins University(约翰·霍普金斯大学)

AI总结 研究多模态大语言模型现实有效性问题,提出CEDI框架,通过三方交互、多轮半结构化对话及多种策略评估,应用于视觉幻觉,发现能揭示更多接近实际情况的幻觉,凸显其对MLLMs能力评估的作用。

详情
AI中文摘要

多模态大语言模型(MLLMs)在基准测试中取得了显著进展,但其在现实世界中的有效性仍不确定。这种差距源于受控静态环境中的基准测试与现实世界应用的动态、交互和情境化性质之间的根本错位。为弥合这一差距,我们提出了CEDI(通过动态多轮交互对MLLMs进行情境化评估)框架,将评估重新构建为被评估模型、自动考官和评分者之间的三方交互。考官通过基于任务的图形表示进行多轮半结构化对话。通过导航状态空间转换,CEDI部署从澄清请求到对抗性探测等各种策略,以获取性能证据。我们将CEDI应用于视觉幻觉。多个模型、不同设置、数据集和领域的实证结果表明,情境化、交互式评估不仅比传统静态评估揭示出更多幻觉,而且揭示出的幻觉更接近实际用例中出现的幻觉。我们还表明,幻觉往往会通过自我强化的对话历史在长语境中累积,并且模型特别容易受到需要拒绝前提或拒绝的问题的影响。这些发现共同凸显了CEDI是朝着对MLLMs能力进行现实、系统和生态有效评估迈出的一步。

英文摘要

Multi-modal Large Language Models (MLLMs) have made substantial advances on benchmarks, yet their real-world effectiveness remains uncertain. This gap stems from the fundamental misalignment between benchmarks in controlled, static settings and the dynamic, interactive, and contextualized nature of real-world applications. To bridge this gap, we propose CEDI (Contextualized Evaluations of MLLMs through Dynamic, multi-round Interactions), a framework that recasts evaluation as a three-party interaction between an evaluatee model, an automated examiner, and a grader. The examiner conducts multi-turn, semi-structured conversation guided by a graph-based representation of the task. By navigating state-space transitions, CEDI deploys diverse strategies, from clarification requests to adversarial probes, to elicit performance evidence. We apply CEDI to visual hallucinations. Empirical results across multiple models, diverse settings, datasets, and domains show that contextualized, interactive evaluations reveal not only significantly more hallucinations than conventional static evaluation but also ones that more closely resemble those arising in practical use cases. We further show that hallucinations often accumulate over long contexts, through self-reinforcing dialogue history, and models are particularly vulnerable to questions requiring premise rejection or refusal. Together, these findings highlight CEDI as a step toward realistic, systematic, and ecologically valid assessments of MLLMs' capabilities. Code is available at github.com/williamium3000/cedi.

URL PDF HTML 收藏
2607.14152 2026-07-17 cs.HC cs.AI 新提交

"Trust Junk" Leads to Unjustified Support for Highly Discriminatory Predictive Models

“信任垃圾”导致对高度歧视性预测模型的不合理支持

Michael Correll, Lucy Havens, Mahsan Nourani

机构 * Northeastern University(东北大学)

AI总结 研究发现在可解释人工智能中,模型解释里提供的准确但多余或无关的数据会使人们对明显歧视性和不公平的模型产生不合理信任,提示XAI设计者和开发者要注意工作中的修辞及可视化带来的潜在问题。

详情
AI中文摘要

数据可视化的说服力可能会出错:例如,在可解释人工智能(XAI)环境中,可视化会导致对预测模型的过度信任。本文通过众包研究表明,在模型解释中提供准确(但多余或无关)的数据,实际上会导致对模型产生不合理的信任和其他积极信念,即使该模型明显具有歧视性和不公平性。研究结果表明,XAI设计者和开发者需要考虑其工作中隐含或明确的修辞,并警惕可视化可能赋予模型不应有的信任。

英文摘要

The persuasive power of data visualizations can go awry: for instance, in an explainable AI (XAI) context, visualizations can produce over-trust of predictive models. In this paper, we use a crowdsourced study to show that providing accurate (but superfluous or irrelevant) data in a model explanation can, in fact, result in unjustified trust and other positive beliefs about a model, even when the model is patently discriminatory and unfair. Our results suggest that XAI designers and developers need to consider the implicit or explicit rhetorics of their work, and beware of the potential of visualizations to imbue models with unearned trust.

URL PDF HTML 收藏
2607.14099 2026-07-17 cs.CL cs.AI cs.CV 新提交

Just Keep Prompting: Evaluating Repetitive Socratic Prompting in VLMs

只需持续提示:评估视觉语言模型中的重复苏格拉底式提示

Shayda Moezzi, Bishoy Galoaa, Lorena Genua, Taskin Padir, Sarah Ostadabbas

机构 * Northeastern University(东北大学)

AI总结 研究视觉语言模型在反复提示下的稳定性,引入JKP多轮评估框架,用三种策略对模型提问,在STAR基准子集上评估GPT-4o、Gemini 2.5 Pro和Qwen3-VL-30B,发现重复提示利弊兼具且因模型而异,揭示了模型压力反应概况。

详情
AI中文摘要

在现实世界中部署视觉语言模型(VLM)不仅需要强大的视觉推理能力,还需要在持续对话压力下保持稳定。我们引入了Just Keep Prompting(JKP),这是一个多轮评估框架,用于衡量当用户反复挑战、质疑或反驳模型答案时VLM的认知稳定性。JKP使用三种策略对模型进行多达10轮的后续提问:对抗性否定(重复拒绝)、纯粹的苏格拉底式询问(重复要求重新评估确定性)和上下文感知苏格拉底式总结(在要求重新考虑之前反映模型先前的理由)。我们在STAR基准的一个子集上对GPT-4o、Gemini 2.5 Pro和Qwen3-VL-30B进行了720次多轮运行的评估。从第0轮到第10轮,总体准确率变化不大,但轨迹级分析显示出显著的不稳定性:正确答案倒退,错误答案恢复,许多运行显示出答案反复翻转。重复提示的好处有限,而且往往起到破坏稳定的作用,而不是推理辅助作用。这种影响强烈依赖于模型:Qwen3-VL-30B最终准确率最高,但在直接矛盾下会自信地给出错误答案;Gemini 2.5 Pro相对稳定,但token成本高;GPT-4o最脆弱且波动大。这些发现表明,多轮VLM评估不仅捕捉了额外的推理,还捕捉了压力反应概况:模型在反复挑战下如何权衡视觉基础、校准和对话合规性。

英文摘要

Deploying Vision-Language Models (VLMs) in real-world settings requires not only strong visual reasoning but also stability under sustained conversational pressure. We introduce Just Keep Prompting (JKP), a multi-turn evaluation framework that measures VLM epistemic stability when users repeatedly challenge, question, or contradict a model's answer. JKP probes models for up to 10 follow-up turns using three strategies: Adversarial Negation (repeated rejection), Pure Socratic Interrogation (repeated calls to reassess certainty), and Context-Aware Socratic Summarization (reflecting the model's prior rationale back before asking for reconsideration). We evaluate GPT-4o, Gemini 2.5 Pro, and Qwen3-VL-30B on a subset of the STAR benchmark across 720 multi-turn runs. Aggregate accuracy changes modestly from Turn 0 to Turn 10, but trajectory-level analysis reveals substantial instability: correct answers regress, wrong answers recover, and many runs exhibit repeated answer flipping. Repeated prompting has bounded upside and often acts as a destabilizer rather than a reasoning aid. The effect is strongly model-dependent: Qwen3-VL-30B achieves the highest final accuracy but becomes confidently wrong under direct contradiction; Gemini 2.5 Pro is comparatively stable but token-expensive; GPT-4o is the most brittle and oscillatory. These findings reveal that multi-turn VLM evaluation captures not just additional reasoning but pressure-response profiles: how models trade off visual grounding, calibration, and conversational compliance under repeated challenge.

URL PDF HTML 收藏
2607.12771 2026-07-17 cs.LG cs.CE cs.CL q-bio.BM 版本更新

Learning Mechanistic Reasoning for Chemical Reactions with Large Language Models

利用大语言模型学习化学反应的机理推理

Xingyu Dang, Haocheng Tang, Junmei Wang, Yanjun Li

机构 * Princeton University(普林斯顿大学) University of Pittsburgh(匹兹堡大学) Northeastern University(东北大学) University of Florida(佛罗里达大学)

AI总结 研究利用大语言模型学习化学反应机理推理,构建大规模推理数据集及福山基准,通过机理感知训练微调Qwen3 - 30B - A3B,在福山基准集A上精确路径匹配超FlowER模型,增强了语言模型的化学推理能力。

详情
AI中文摘要

反应机理由解释化学转化的基本反应的逐步序列组成。因此,学习机理逻辑对于增强大语言模型(LLMs)的基本化学智能至关重要。反应机理的逐步推导与推理LLMs的推理范式自然契合。然而,当前化学LLMs主要强调用于产物预测和逆合成的粗粒度名称反应,常导致物理不一致和幻觉。相比之下,用于机理推断的专门小规模生成模型通常在不同化学空间中泛化能力受限。为克服这些限制,我们构建了一个新颖的大规模反应机理推理数据集。此外,我们建立了福山基准,这是一个源自福山《高等有机反应机理》一书的具有挑战性的基准,用于严格评估模型在分层机理推理上的性能。我们微调后的Qwen3 - 30B - A3B在福山基准集A上实现了8.3%的精确路径匹配,超过了专门的FlowER模型(5.1%),表明机理感知训练显著增强了语言模型中的化学推理能力。

英文摘要

Reaction mechanisms consist of the step-by-step sequences of elementary reactions that explain chemical transformations. Learning the mechanism logic is therefore essential for enhancing the fundamental chemical intelligence of large language models (LLMs). The stepwise deduction of reaction mechanism aligns naturally with the reasoning paradigms of reasoning LLMs. However, current chemical LLMs primarily emphasize coarse-grained name reactions for product prediction and retrosynthesis, often leading to physical inconsistencies and hallucinations. In contrast, specialized small-scale generative models for mechanism inference typically suffer from restricted generalization capacity across diverse chemical spaces. To overcome these limitations, we built a novel, large-scale reasoning dataset of reaction mechanisms. Furthermore, we established the FukuyamaBench, a difficult benchmark derived from Fukuyama's Advanced Organic Reaction Mechanism book, to rigorously evaluate model performance on hierarchical mechanism reasoning. Our fine-tuned Qwen3-30B-A3B achieves 8.3% exact pathway match on FukuyamaBench Set~A, surpassing the specialized FlowER model (5.1%), demonstrating that mechanism-aware training substantially enhances chemical reasoning in language models.

URL PDF HTML 收藏
2607.13818 2026-07-16 cs.RO 新提交

Learning Robust Execution in Robotic Manipulation with Agentic Reinforcement Learning

通过智能强化学习在机器人操作中学习鲁棒执行

Xiaopeng Zhang, Yueyang Weng, Qi Liu, Yongjin Mu, Yanjie Li

机构 * School of Inteligence Science and Engineering, the Harbin Institute of Technology Shenzhen(哈尔滨工业大学(深圳)智能科学与工程学院) Faculty of Robot Science and Engineering, Northeastern University(东北大学机器人科学与工程学院)

AI总结 针对机器人操作面临的挑战,提出用两个互补指标评估执行质量,构建智能强化学习框架,通过高级决策恢复有效执行,在LIBERO基准测试中提升了执行成功率,增强了执行鲁棒性。

详情
AI中文摘要

机器人操作因不确定性、长期执行和复合误差面临根本挑战,易导致执行不稳定和任务失败。近期视觉语言动作(VLA)模型虽有强泛化能力,但缺乏评估执行稳定性及恢复偏离标称行为的机制。本文提出两个互补指标评估运行时执行质量,以及一个智能强化学习框架,通过高级决策恢复有效执行。该框架基于执行历史推理并选择执行模式调节执行过程,执行退化时触发恢复机制使任务继续。在LIBERO基准测试中评估,标准设置下成功率提高达13.7%,干扰设置下提高达39.2%,显著增强了执行鲁棒性。

英文摘要

Robotic manipulation poses fundamental challenges due to uncertainty, long-horizon execution, and compounding errors, which can easily destabilize execution and lead to task failure. Although recent vision-language-action (VLA) models exhibit strong generalization, they typically lack explicit mechanisms to assess execution stability and to recover when execution deviates from its nominal behavior. In this paper, we propose: (1) two complementary metrics to assess execution quality at runtime, and (2) an agentic reinforcement learning framework that learns to restore effective execution through high-level decision-making rather than directly learning low-level actions. In this framework, an agentic policy reasons over recent execution history and selects among a small set of execution modes to regulate the execution process. Under execution degradation, it triggers appropriate recovery mechanisms to restore the robot to previously visited nominal states, enabling the task to continue. We evaluate the proposed method on the LIBERO benchmark, achieving up to a 13.7% improvement in success rate under standard settings and up to a 39.2% improvement under disturbance settings, demonstrating substantially enhanced execution robustness.

URL PDF HTML 收藏
2607.13497 2026-07-16 cs.RO 新提交

Layered Risk Mapping for Autonomous Patient Transport in Expeditionary Medical Facilities

用于远征医疗设施中自主患者运输的分层风险映射

Lorena Maria Genua, Sarvesh Prajapati, Damla Leblebicioglu, Taşkın Padır

机构 * Institute for Experiential Robotics, Northeastern University(体验式机器人研究所,东北大学)

AI总结 针对远征医疗设施中自主患者运输面临的复杂导航挑战,提出分层风险映射框架,融合多种环境危害,经配对蒙特卡洛评估及实际验证,有效降低碰撞率、提高障碍物清除率,满足相关运营模式规划要求。

详情
AI中文摘要

在远征医疗设施中,常规患者运输带来了个人防护装备消耗、人员转移和感染风险增加等多重负担,在高峰情况下难以为继。虽然自动轮椅可分担此运营负荷,但在高度非结构化和动态环境中患者运输的安全关键特性带来了复杂导航挑战。为此,我们提出了一个分层风险映射框架,通过噪声或融合模型将四种异构环境危害(地形坡度、静态和动态障碍物以及语义可通行性)融合到统一的概率成本表面。在配对蒙特卡洛评估中,风险知情融合将碰撞率从超过73%降至32%以下,相对于无风险基线,障碍物清除率增加了一倍多。此外,在所有测试的危害密度下,噪声或实现了最高的障碍物清除率和最低条件峰值风险。我们还在商业电动轮椅上针对室内和室外部署的三个代表性任务配置文件验证了该框架,表明此架构成功满足了这一以前未解决的运营模式的规划要求。

英文摘要

In expeditionary medical facilities, routine patient transport imposes a compounding burden of personal protective equipment consumption, staff diversion, and elevated infection risk that becomes unsustainable under surge conditions. While autonomous wheelchairs could absorb this operational load, the safety-critical nature of patient transit within these highly unstructured and dynamic environments poses complex navigational challenges. To address this, we present a layered risk mapping framework that fuses four heterogeneous environmental hazards (terrain slope, static and dynamic obstacles, and semantic traversability) into a unified probabilistic cost surface via a Noisy-OR fusion model. In a paired Monte-Carlo evaluation, risk-informed fusion reduces collision rates from over 73% to under 32% and more than doubles obstacle clearance relative to a risk-unaware baseline. Additionaly, Noisy-OR achieves the highest clearance to obstacles and the lowest conditional peak risk across all tested hazard densities. We further validate the framework on a commercial powered wheelchair across three representative mission profiles in indoor and outdoor deployments, demonstrating that this architecture successfully meets the planning requirements of this previously unaddressed operational regime.

URL PDF HTML 收藏
2606.02911 2026-07-16 cs.CL 版本更新

The Ghost Annotator: a Framework to Explore Human Label Variation in Content Moderation through Conformal Prediction

幽灵标注者:通过共形预测探索内容审核中人类标签变异的框架

Mirko Lai, Alessandra Urbinati, Simona Frenda, Fabiana Vernero, Marco Antonio Stranisci

机构 * Laboratory for the Modeling of Biological and Socio-technical Systems, Northeastern University(生物与社会技术系统建模实验室,东北大学) Heriot-Watt University(赫瑞-沃顿大学) aequa-tech Università del Piemonte Orientale(皮埃蒙特东方大学) Università degli Studi di Torino(托斯卡纳大学)

AI总结 提出结合共形预测与协同过滤式标注者表征的框架,通过幽灵预测度量和幽灵标注者表征量化模型预测与所有人类标注的分歧,并发现模型在标注者分歧时不确定性增加,但大型模型对无人类对齐文本更自信,且存在结构性人口统计偏差。

Comments The publishing of this preprint is contextual with the ACL ARR cycle system. After an encouraging review in January we revised and submit the paper on Arxiv. However, a new batch of reviewers raised additional issues that will lead to significant revisions of the experimental setting. Therefore, we decide to withdraw the manuscript

详情
AI中文摘要

当前研究主要关注模型性能,而对不确定性估计的关注相对较少,特别是在LLM越来越多地用于生成标注数据的场景中。我们引入了一个框架,将共形预测与协同过滤式的标注者表征相结合,以建模LLM相对于人类标注者的行为,并分析一致与分歧的模式。利用非一致性分数,我们引入了幽灵预测度量和幽灵标注者表征,以量化模型预测与所有可用人类标注不一致的情况。我们计算余弦相似度度量,以探索模型行为在不同社会人口统计轴上的差异。我们在四个内容审核数据集上评估了四种不同规模和家族的LLM。我们的发现表明,虽然所有模型的不确定性随着标注者分歧的增加而增加,但较大的模型在对与任何人类标注不一致的文本进行分类时往往更自信。最后,幽灵标注者框架揭示了一致且稳健的人口统计错位模式,表明可能存在源于预训练语料库的结构性偏见。

英文摘要

Current research primarily focuses on model performance, while comparatively less attention has been devoted to uncertainty estimation, particularly in settings where LLMs are increasingly used to generate annotated data. We introduce a framework combining conformal prediction with Collaborative Filtering-style annotators' representation to model LLM behavior in relation to human annotators and to analyze patterns of agreement and disagreement. Using Non-Conformity Scores, we introduce the Ghost Prediction metric and the Ghost Annotator representation to quantify cases in which model predictions diverge from all available human annotations. We compute cosine similarity measures to explore differences in model behavior across sociodemographic axes. We evaluated four LLMs of different size and families across four content moderation datasets. Our finding shows that while we find that all models uncertainty increases with annotator disagreement, larger models tend to be more confident in the classification of texts that are not aligned with any human annotation. Finally, the Ghost Annotator framework reveals a consistent and robust pattern of demographic misalignment, suggesting a structural bias likely rooted in pretraining corpora.

URL PDF HTML 收藏
2512.10607 2026-07-16 cs.CV 版本更新

Track and Caption Any Motion: Open-Vocabulary Spatiotemporal Captioning via Trajectory-Conditioned Generation

跟踪并描述任何运动:通过轨迹条件生成实现开放词汇时空字幕

Bishoy Galoaa, Sarah Ostadabbas

机构 * Northeastern University(东北大学)

AI总结 研究提出TCAM框架,无需文本查询和区域提示,通过字幕感知重采样器在点粒度结合跟踪与语言,用现有分割注释训练,能描述视频中运动、定位时间及轨迹,优于密集视频字幕基线,匹配相关方法,为运动驱动视频理解提供新途径。

详情
AI中文摘要

我们提出了TCAM(跟踪并描述任何运动),这是一个生成框架,无需文本查询和区域提示就能观看视频,确定运动物体,用开放词汇描述每个运动,定位其时间,并指出承载该运动的精确轨迹。两条成熟的工作线使这成为可能但未解决:密集点跟踪器能高精度跟踪像素但不输出语言,视频语言模型仅在有查询时且仅从无法分辨哪些像素移动的剪辑级特征生成流畅描述。对象级字幕器缩小了差距但仍基于检测框或掩码推理,未触及单个轨迹。TCAM通过字幕感知重采样器在点粒度上结合跟踪和语言,少量可学习查询交叉关注密集点轨迹令牌并将其提炼为固定长度的运动上下文来条件化语言解码器。解码器单次生成整个视频的事件,每个事件有自由形式字幕、开始和结束时间以及指向其参考轨迹的指针。训练仅使用现有分割注释监督字幕质量、指针掩码对齐和指针多样性。在超过50K个剪辑上,TCAM优于密集视频字幕基线,且在不使用查询的情况下与专用的基于查询的定位和点跟踪方法相匹配,表明轨迹条件生成是通往运动驱动视频理解的直接途径。

英文摘要

We present TCAM (Track and Caption Any Motion), a generative framework that watches a video and with no text query and no region prompt decides what is moving, describes each motion in open vocabulary, locates it in time, and points to the exact trajectories that carry it. Two mature lines of work make this possible yet leave it unsolved: dense point trackers follow pixels with sub-object precision but emit no language, while video-language models produce fluent descriptions only when handed a query and only from clip-level features that cannot resolve which pixels move. Object-level captioners narrow the gap but still reason over detector boxes or masks, never reaching individual trajectories. TCAM couples tracking and language at point granularity through a Caption-Aware Resampler, where a small set of learnable queries cross-attends to dense point trajectory tokens and distills them into a fixed-length motion context that conditions a language decoder. The decoder generates an entire video's events in a single pass, each with a free-form caption, a start and end time, and a pointer to the trajectories it refers to, for sequential events and several subjects active at once. Training uses only existing segmentation annotations, with no extra event labeling, to supervise caption quality, pointer-mask alignment, and pointer diversity. On over 50K clips, TCAM outperforms dense video captioning baselines and matches dedicated, query-based grounding and point-tracking methods despite using no query, showing that trajectory-conditioned generation is a direct route to motion-driven video understanding.

URL PDF HTML 收藏
2607.08072 2026-07-15 cs.CV cs.RO 版本更新

Post-Training in End-to-End Autonomous Driving

端到端自动驾驶中的训练后处理

Ruining Yang, Muxing Wang, Yixiao Chen, Tongfei Guo, Yi Xu, Can Cui, Zichong Yang, Yitian Zhang, Ziran Wang, Yun Fu, Lili Su

机构 * Northeastern University(东北大学) Purdue University(普渡大学)

AI总结 探讨端到端自动驾驶中训练后处理技术,针对传统方法不足,将现有文献按监督形式分四类,阐述各分类的能力、局限与挑战,助力系统理解该领域并推动相关研究。

详情
AI中文摘要

将多模态输入直接映射到未来轨迹/操纵的端到端模型在自动驾驶中已成为日益突出的研究范式,包括视觉-语言-动作模型和轨迹生成规划器。与传统机器学习应用不同,自动驾驶车辆在安全关键且交互密集的环境中运行,传统的专家示范开环模仿不足以确保可靠性。小的执行错误会随时间累积,训练数据中恢复行为稀缺,逐点标签也无法捕捉安全和驾驶舒适性等长期目标。这些限制促使转向训练后技术,以进一步完善驾驶策略。本综述通过定义其范围并将现有文献按所使用的监督形式分为四个主要类别,对自动驾驶的训练后处理给出统一观点。针对每个类别,讨论其能力、局限性和开放挑战。旨在促进对这一新兴领域的系统理解,并激发未来关于可靠高效的自动驾驶训练后处理的研究。

英文摘要

End-to-end models that map multimodal inputs directly to future trajectories/maneuvers have emerged as an increasingly prominent research paradigm in autonomous driving. This class of models includes both Vision-Language-Action models and trajectory-generative planners. Unlike classic machine learning applications, autonomous vehicles operate in safety-critical and interaction-intensive environments where traditional open-loop imitation of expert demonstrations is not sufficient to ensure reliability. In particular, small execution errors can accumulate over time, while recovery behaviors are scarce in training data. In addition, long-horizon objectives such as safety and driving comfort are not captured by pointwise labels either. These limitations have motivated a shift toward post-training techniques, which further refine driving policies beyond pure imitation. This survey presents a unified view of post-training for autonomous driving by defining its scope and organizing the existing literature into four major families based on the form of supervision they use. For each family, we discuss its capabilities, limitations, and open challenges. We aim to facilitate a systematic understanding of this emerging area and stimulate future research on reliable and efficient post-training for autonomous driving.A collection of related papers is available at https://github.com/RYNing/Awesome-Post-Training-In-Autonomous-Driving-Papers.

URL PDF HTML 收藏
2506.21324 2026-07-15 cs.NE cs.LG 版本更新

Stochastic Quantum Spiking Neural Networks with Quantum Memory and Local Learning

具有量子记忆和局部学习的随机量子脉冲神经网络

Jiechen Chen, Bipin Rajendran, Osvaldo Simeone

机构 * Department of Engineering, King’s College London(伦敦国王学院工程系) Institute for Intelligent Networked Systems, Northeastern University London(伦敦东北大学智能网络系统研究所)

AI总结 本文提出了一种具有量子记忆和局部学习的随机量子脉冲神经网络模型,通过单次 shot 实现事件驱动的随机脉冲生成,并通过硬件友好的局部学习规则进行训练,从而在固定参数数量下优于传统和量子脉冲神经网络。

Comments Published in IEEE Journal on Selected Areas in Communications

详情
AI中文摘要

神经形态计算和量子计算近年来 emerged 作为推动人工智能发展的有前景的范式,各自提供互补的优势。基于脉冲神经元的神经形态系统通过稀疏、事件驱动的计算高效处理时间序列数据,仅在输入事件时消耗能量。而量子计算则在量子态空间中随着量子比特数量的增长呈指数级增长,量子态允许在基态之间叠加以及子系统之间的纠缠。结合这些范式的混合方法开始展现出潜力,但现有的量子脉冲模型有重要的局限性。值得注意的是,它们在单个量子比特上实现经典记忆机制,需要重复测量来估计放电概率,同时依赖于传统的反向传播进行训练。在本文中,我们提出了一种新的随机量子脉冲(SQS)神经元模型,以解决这些挑战。SQS神经元使用多量子比特量子电路实现具有内部量子记忆的脉冲单元,能够在推理过程中单次 shot 中实现事件驱动的随机脉冲生成。此外,我们研究了SQS神经元的网络,称为SQS神经网络(SQSNN),并证明它们可以通过硬件友好的局部学习规则进行训练,从而消除对全局经典反向传播的依赖。所提出的SQSNN模型通过使用传统和神经形态数据集的实验表明,在固定总体可训练参数数量的情况下,优于先前的量子脉冲神经网络以及经典模型。

英文摘要

Neuromorphic and quantum computing have recently emerged as promising paradigms for advancing artificial intelligence, each offering complementary strengths. Neuromorphic systems built on spiking neurons excel at processing time series data efficiently through sparse, event-driven computation, consuming energy only upon input events. Quantum computing, on the other hand, operates on state spaces that grow exponentially in dimension with the number of qubits -- as a consequence of tensor-product composition -- with quantum states admitting superposition across basis states and entanglement between subsystems. Hybrid approaches combining these paradigms have begun to show potential, but existing quantum spiking models have important limitations. Notably, they implement classical memory mechanisms on single qubits, requiring repeated measurements to estimate firing probabilities, while relying on conventional backpropagation for training. In this paper, we propose a novel stochastic quantum spiking (SQS) neuron model that addresses these challenges. The SQS neuron uses multi-qubit quantum circuits to realize a spiking unit with internal quantum memory, enabling event-driven probabilistic spike generation in a single shot during inference. Furthermore, we study networks of SQS neurons, dubbed SQS neural networks (SQSNN), and demonstrate that they can be trained via a hardware-friendly local learning rule, eliminating the need for global classical backpropagation. The proposed SQSNN model is shown via experiments with both conventional and neuromorphic datasets to improve over previous quantum spiking neural networks, as well as over classical counterparts, when fixing the overall number of trainable parameters, highlighting its potential for event-driven applications such as neuromorphic integrated sensing and communications (N-ISAC).

URL PDF HTML 收藏
2512.04144 2026-07-15 cs.AI 版本更新

RippleBench: Capturing Ripple Effects Using Existing Knowledge Repositories

RippleBench: 利用现有知识库捕捉涟漪效应

Roy Rinberg, Usha Bhalla, Igor Shilov, Flavio P. Calmon, Rohit Gandikota

机构 * Harvard University(哈佛大学) Imperial College London(伦敦帝国学院) Northeastern University(东北大学)

AI总结 提出RippleBench-Maker自动管道,从知识库检索语义邻居生成选择题,评估八种遗忘方法在Llama3-8B-Instruct上的涟漪效应,发现准确率下降随语义距离衰减且跨模型一致。

详情
AI中文摘要

针对语言模型的目标干预,如遗忘或模型编辑,旨在修改特定信息,但其效果往往传播到相关的、非预期的领域(例如,删除病毒学内容可能降低对过敏任务的性能);这些副作用通常被称为涟漪效应。我们引入RippleBench-Maker,一个自动管道,从知识库中检索任何源概念的语义邻居,并生成不同语义距离的多选题。我们使用WikiRAG(一个基于英文维基百科的开源RAG系统)实例化该框架,构建RippleBench-WMDP-Bio(584个种子主题,352,961个问题),并在Llama3-8B-Instruct上评估八种遗忘方法。所有八种方法在遗忘目标附近准确率下降最大,并随语义距离衰减,每种方法具有不同的传播曲线。我们在Mistral-7B、Zephyr-7B和Yi-34B上复现了这些发现;跨模型的差值曲线几乎相同,表明涟漪效应是遗忘方法的属性而非基础模型。我们通过一项包含四个实验的Mechanical Turk研究(5,200+次响应,61名工作者)验证了所有主要管道阶段。我们发布所有代码、数据和基础设施。

英文摘要

Targeted interventions on language models, such as unlearning or model editing, aim to modify specific information, but their effects often propagate to related, unintended areas (e.g., removing virology content may degrade performance on allergies); these side-effects are commonly referred to as the ripple effect. We introduce RippleBench-Maker, an automatic pipeline that retrieves semantic neighbors of any source concept from a knowledge repository and generates multiple-choice questions at varying semantic distances. We instantiate this framework using WikiRAG, an open-source RAG system over English Wikipedia, to construct RippleBench-WMDP-Bio (584 seed topics, 352,961 questions), and evaluate eight unlearning methods on Llama3-8B-Instruct. All eight exhibit accuracy drops that are largest near the unlearned target and decay with semantic distance, each with a distinct propagation profile. We replicate these findings across Mistral-7B, Zephyr-7B, and Yi-34B; cross-model delta curves are nearly identical, suggesting ripple effects are a property of the unlearning method rather than the base model. We validate all major pipeline stages using a four-experiment Mechanical Turk study (5,200+ responses, 61 workers). We release all code, data, and infrastructure.

URL PDF HTML 收藏
2502.11554 2026-07-15 cs.HC cs.AI cs.CL cs.CY cs.ET 版本更新

Toward Metaphor-Fluid Conversation Design for Voice User Interfaces

面向语音用户界面的隐喻灵活对话设计

Smit Desai, Jessie Chin, Dakuo Wang, Benjamin Cowan, Michael Twidale

机构 * Northeastern University(东北大学) University of Illinois, Urbana-Champaign(伊利诺伊大学厄巴纳-香槟分校) University College Dublin(都柏林大学)

AI总结 研究针对语音用户界面现有隐喻设计不足,提出隐喻灵活设计方法,通过动态调整隐喻表示适应不同情境。经实验对比默认VUI,该方法提升了用户相关感受,但因个体差异需个性化,挑战了传统设计范式,展现新方法潜力。

Comments Accepted at ACM CUI'26

详情
AI中文摘要

隐喻在塑造语音用户界面(VUI)的用户体验中起着关键作用,但现有设计往往依赖静态、以人类为中心的隐喻,无法适应不同情境和用户需求。本文介绍了隐喻灵活设计,一种基于对话使用情境动态调整隐喻表示的新方法。将其与默认VUI比较,在研究1中,隐喻被映射到四个关键使用情境,揭示了对特定任务隐喻设计的不同偏好。研究2表明,隐喻灵活VUI通过更好地符合用户对不同情境的期望增强了采用意愿、愉悦感和喜爱度。然而,隐喻偏好的个体差异凸显了个性化的必要性。这些发现挑战了VUI设计的一刀切范式,证明了隐喻灵活设计在创建更具适应性和吸引力人机交互方面的潜力。

英文摘要

Metaphors play a critical role in shaping user experiences with Voice User Interfaces (VUIs), yet existing designs often rely on static, human-centric metaphors that fail to adapt to diverse contexts and user needs. This paper introduces Metaphor-Fluid Design, a novel approach that dynamically adjusts metaphorical representations based on conversational use-contexts. We compare this approach to a Default VUI, which characterizes the present implementation of commercial VUIs commonly designed around the persona of an assistant, offering a uniform interaction style across contexts. In Study 1 (N=130), metaphors were mapped to four key use-contexts-commands, information seeking, sociality, and error recovery-along the dimensions of formality and hierarchy, revealing distinct preferences for task-specific metaphorical designs. Study 2 (N=91) evaluates a Metaphor-Fluid VUI against a Default VUI, showing that the Metaphor-Fluid VUI enhances perceived intention to adopt, enjoyment, and likability by aligning better with user expectations for different contexts. However, individual differences in metaphor preferences highlight the need for personalization. These findings challenge the one-size-fits-all paradigm of VUI design and demonstrate the potential of Metaphor-Fluid Design to create more adaptive and engaging human-AI interactions.

URL PDF HTML 收藏
2607.11855 2026-07-14 cs.RO 新提交

Robust bipedal locomotion on flowable slopes via foot-driven terrain manipulation

通过足部驱动的地形操纵在可流动斜坡上实现稳健的双足运动

Deniz Kerimoglu, Junnosuke Kamohara, Jiyeon Maeng, Ziwon Yoon, Seth Hutchinson, Ye Zhao, Daniel I. Goldman

机构 * Georgia Institute of Technology(佐治亚理工学院) Northeastern University(东北大学)

AI总结 研究双足机器人在颗粒斜坡上的运动控制问题,通过研究带防滑钉足部的地面动力学,发现中等防滑钉间距利于行走,据此设计可调整防滑钉深度的足部,应用于大小不同的双足机器人,提出以肢体为中心调节地形相互作用的新控制方法。

Comments 38 pages, 12 figures

详情
AI中文摘要

双足机器人控制具有挑战性,因其接近不稳定状态,足部与地形接触的微小变化会迅速破坏运动稳定性。在刚性地形上,可通过成熟的接触力学和控制策略缓解这种脆弱性。而在可流动表面如颗粒斜坡上,足部接触会引发大的表面变形和类似固液转变,耦合地形效应与机器人动力学,导致性能不佳或失败,部分原因是缺乏可靠的可流动地形动力学表示方法。本文通过研究带防滑钉的足部(从鞋底伸出的薄板)的地面动力学,探讨控制地形响应如何改善颗粒斜坡上的双足运动。对小型(1.4千克)机器人物理双足的系统研究表明,防滑钉间距稀疏和密集分别会导致过度的地形屈服和阻力,降低性能并导致失败。中等防滑钉间距可分布相互作用力,使基底应力维持在(或低于)屈服阈值,从而能在高达30度的颗粒斜坡上行走。基于这些原理,设计了一种能主动调整防滑钉深度并适应刚性和颗粒地形的足部。还证明了有效的足部与地形相互作用原理可应用于更大(15千克)的自主双足机器人。本研究提出了一种替代传统以身体为中心的机器人控制方法的方案,即通过以肢体为中心的方法调节地形相互作用,而非通过身体运动调节地形诱导效应。

英文摘要

Bipedal robots are challenging to control because they operate close to instability, where small variations in foot-terrain contact can rapidly destabilize locomotion. On rigid terrain, bipedal robots mitigate this fragility by using well-established contact mechanics and control strategies. On flowable surfaces such as granular slopes, foot contact can induce large surface deformations and solid-fluid-like transitions, coupling terrain effects with robot dynamics, leading to underperformance or failure. This is partly due to the lack of reliable methods to represent the dynamics of flowable terrain, making it difficult to account for terrain effects in locomotion design. Here, we investigate how controlling terrain response can improve bipedal locomotion on granular slopes by studying the terradynamics of cleated feet, thin plates emanating from the foot soles. Systematic studies of a small-scale (1.4 kg) robophysical biped reveal that cleats with sparse and dense spacing lead to excessive terrain yielding and resistance, respectively, degrading performance and leading to failure. An intermediate cleat spacing distributes interaction forces to maintain substrate stresses near (or below) the yield threshold, enabling walking on granular slopes up to 30 degrees. Guided by these principles, we design a foot that actively adjusts cleat depth and accommodates both rigid and granular terrain. We also demonstrate that the principles of effective foot-terrain interaction translate to a larger (15 kg) autonomous biped. Our study presents an alternative to conventional body-centric robot control approaches, which regulate terrain-induced effects through body motion, by instead regulating terrain interactions through limb-centric approach.

URL PDF HTML 收藏
2607.11423 2026-07-14 cs.CL 新提交

ToFu: A White-Box, Token-Efficient Agent Harness for Researchers

ToFu:面向研究人员的白盒、令牌高效代理工具包

Junhao Ruan, Yuan Ge, Bei Li, Yongjing Yin, Yuchun Fan, Xin Chen, Jingang Wang, Chenglong Wang, Jingbo Zhu, Tong Xiao

机构 * Northeastern University(东北大学) LongCat RSI, Meituan(美团龙猫RSI) NiuTrans Research(小牛翻译研究院)

AI总结 研究提出ToFu这一代理工具包,作为研究助手,以高令牌效率等支持实际研究流程,且能本地部署;作为研究对象,提供白盒工具包供研究人员检查、修改和评估,兼具强大性能与用户体验。

详情
AI中文摘要

代理编码工具为转变研究工作流程带来新机遇。构建的代理系统性能取决于大语言模型(LLMs)及其周围的工具包,即决定代理行为的编排代码。我们提出了ToFu,一个面向研究人员的代理工具包,可读取代码库、编辑文件、运行命令并与开发工具集成。ToFu在研究中扮演双重角色。作为研究助手,与现有代理工具包相比,它以更高的令牌效率、更低的成本和多语言能力支持实际研究工作流程。其在MIT许可下发布,进一步支持对隐私敏感用户进行本地部署。作为研究对象,ToFu提供了一个白盒代理工具包,允许研究人员检查、修改和评估其编排逻辑、工具使用行为和工具包设计,同时保持强大的基准性能和应用级用户体验。

英文摘要

Agentic coding tools present new opportunities to transform research workflows. The performance of agent systems built depends on both large language models (LLMs) and the harness around LLMs, which is the orchestration code that determines an agent's behavior. We present ToFu, an agentic harness for researchers that reads your codebase, edits files, runs commands, and integrates with your development tools. ToFu plays a dual role in research. As a research assistant, it supports practical research workflows with superior token efficiency, lower cost, and multilingual capability compared with existing agentic harnesses. Its release under the MIT License further enables local deployment for privacy-sensitive users. As a research object, ToFu provides a white-box agentic harness that allows researchers to inspect, modify, and evaluate its orchestration logic, tool-use behavior, and harness design, while retaining strong benchmark performance and an application-level user experience.

URL PDF HTML 收藏