arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.01050cs.AIcs.CLcs.SE

不要提供无法完成的任务:面向大规模LLM技能选择的确定性可执行性门控机制

Don't Offer What Can't Be Done: Deterministic Executability Gating for LLM Skill Selection at Scale

Ortal Ashkenazi, Vitalii Kloz, Mykhailo Ulianchenko

首次发表
浏览论文内容

中文总结 AI 辅助

该研究针对大规模LLM技能选择中不可执行技能的干扰问题,提出三阶段选择流水线,经生产分析和反事实测试,可大幅减少技能描述上下文并阻止不可执行技能影响模型选择。

中文摘要 AI 辅助

生产环境中的大语言模型(LLM)智能体从大型技能库中选择技能时面临一个仅靠语义相关性无法解决的限制:某项技能可能与用户的主题匹配,但在当前账户状态下却无法执行。本文提出了一种已部署的三阶段选择流水线,用于Wix公司的客户服务助手Helpmate。第一阶段,面向召回的语义匹配器在不参考账户状态的情况下识别与10个技能领域族相关的消息。第二阶段,确定性可执行性门控机制会移除那些存在内部硬停止条件的候选技能。由于该门控机制与技能会评估相同的退出谓词,只要保持谓词一致性且两者都观测到最新的权威状态,在相同账户状态下,每个被阻止的候选技能都将无法完成任务。最后,LLM会决定是否调用剩余候选技能中的一个。在对267600次对话中的756600条用户消息进行的发布后生产分析中,语义匹配保留了174927条消息(占比23.1%)。在该匹配流中,门控机制移除了1749270个技能-消息对中的1039462个(占比59.4%),节省了2.288亿个技能描述token,占语义匹配后技能描述总token量的59.1%。相比将全部10个技能暴露给每条消息,语义匹配与可执行性门控机制共同将技能描述上下文减少了90.5%。为测试这种剪枝是否仅影响上下文大小而非模型行为,我们用全部10个技能暴露的情况复现了包含风险的1000次对话队列。模型在78次对话中选择了生产环境阻止的技能(占比7.8%)。这一反事实结果表明,确定性门控机制可防止不可执行的候选技能影响模型选择,但不涉及下游工具执行或客户结果的影响。

英文摘要

Production LLM agents that select from large skill libraries face a limitation that semantic relevance alone cannot resolve: a skill may match a user's topic yet be impossible to execute in the current account state. We present a deployed three-stage selection pipeline for Helpmate, Wix's customer-care assistant. First, a recall-oriented semantic matcher identifies messages related to a ten-skill domain family without consulting account state. Second, a deterministic executability gate removes candidates whose internal hard-stop conditions hold. Because the gate and the skill evaluate the same exit predicates, every blocked candidate would be unable to complete under the same account state, provided predicate parity is preserved and both checks observe fresh authoritative state. Finally, the LLM decides whether to invoke one of the remaining candidates. In a post-launch production analysis of 756.6K user messages across 267.6K conversations, semantic matching retained 174,927 messages (23.1%). Within this matched stream, the gate removed 1,039,462 of 1,749,270 skill-message pairs (59.4%), saving 228.8 million skill-description tokens -- 59.1% of the post-semantic skill-description footprint. Together, semantic matching and executability gating reduced skill-description context by 90.5% relative to exposing all ten skills to every message. To test whether this pruning affects model behavior rather than context size alone, we replayed a risk-enriched cohort of 1,000 conversations with all ten skills exposed. The model selected a production-blocked skill in 78 conversations (7.8%). This counterfactual result shows that deterministic gating prevents non-executable candidates from influencing model selection, while not claiming downstream tool execution or customer-outcome effects.

发表机构

  • Wix Israel(Wix以色列)
  • Wix Ukraine(Wix乌克兰)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑