arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.20670cs.AIcs.CL

Why2Speak:弃权(不执行)策略的忠实推理

Why2Speak: Faithful Reasoning for Abstaining Action Policies

Shreya Mendi, Brinnae Bent

首次发表
浏览论文内容

中文总结 AI 辅助

该研究针对智能体需在行动与弃权间选择的问题,以多方对话干预为场景,对比多种策略,发现能力与可审计性的权衡,分析现有方法缺陷并提供监督评估的控制方法。

中文摘要 AI 辅助

许多智能体系统必须在行动与弃权(不执行)之间反复做出选择,这使得忠实推理对监督而言十分重要:只有当解释反映了产生该行动的计算过程时,解释才有用。我们通过多方对话中的干预时机来研究该问题,其中助手必须决定是发言还是保持沉默。该场景暴露出类别不平衡、行动成本不对称,以及暴露推理会改变被审计策略的可能性。我们使用Qwen3-8B模型,在有无思维链推理的情况下进行解码,比较直接决策策略、推理策略、监督微调与强化学习。我们发现存在能力-可审计性权衡:最强的直接策略实现了更高质量,但未暴露任何可检查的推理;而推理策略提供了一条轨迹,代价是性能降低,尤其是对真实干预机会的召回率。监督微调要么抑制推理,要么保留推理但不提高决策质量,而强化学习也未能改进推理策略。我们确定了此失败的一个潜在机制:当采样的 rollout 全部选择同一行动时,组相对目标无法为自信错误的提示提供学习信号。受控激活探针与行为消融表明,标准忠实性方法可能夸大了暴露的推理反映潜在决策过程的证据。基于概率的指标在自信决策下会饱和,探针易受类别不平衡与文本泄漏的影响,而推理消融可能混淆推理内容与推理模式的变化。综上,这些结果表明,暴露推理可以改变智能体的行动策略,而非仅使其可观测。我们提供了用于评估可行动或弃权(不执行)智能体的基于推理的监督的控制方法。

英文摘要

Many agentic systems must repeatedly choose between acting and abstaining, making faithful reasoning important for oversight: an explanation is useful only if it reflects the computation that produced the action. We study this problem through intervention timing in multi-party conversation, where an assistant must decide whether to speak or remain silent. This setting exposes class imbalance, asymmetric action costs, and the possibility that exposing reasoning changes the policy being audited. Using Qwen3-8B, decoded with or without chain-of-thought reasoning, we compare direct decision policies, reasoning policies, supervised fine-tuning, and reinforcement learning. We find a capability-auditability tradeoff: the strongest direct policy achieves higher quality but exposes no reasoning to inspect, while the reasoning policy provides a trace at the cost of lower performance, particularly recall of true intervention opportunities. Supervised fine-tuning either suppresses reasoning or preserves it without improving decision quality, while reinforcement learning also fails to improve the reasoning policy. We identify one mechanism underlying this failure: group relative objectives provide no learning signal on confidently wrong prompts when sampled rollouts all select the same action. Controlled activation probes and behavioral ablations show that standard faithfulness methods can overstate evidence that exposed reasoning reflects the underlying decision process. Probability-based metrics saturate under confident decisions, probes are vulnerable to class imbalance and textual leakage, and reasoning ablations can confound reasoning content with changes in inference mode. Together, these results show that exposing reasoning can change an agent's action policy rather than simply make it observable. We provide controls for evaluating reasoning-based oversight of agents that can act or abstain.

↑