arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

可缓存决策模型能否遵循规则?

Can a Cacheable Decision Model Follow Rules?

Dushyant Rajput, Nirdesh Chauhan, Siddharth Kosaraju

arXiv 2609.37832首次发表:更新:

发表机构

AltSlate Labs LLP(AltSlate Labs LLP)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

该研究考察可缓存决策模型Certo在分离状态与候选编码后能否保持规则敏感性,发现其性能下降,但可通过反事实监督在合成任务上恢复,而联合评分器在真实规则上仍具优势。

AI 中文摘要

Certo 是一个小型非生成式决策模型(Qwen3-4B):它不生成答案,而是对候选动作的文本进行评分并返回概率。精确的设计将状态、规则和每个候选一起读取(联合评分器),因此成本随候选菜单规模增长。独立编码允许每个候选被编码一次并在不同状态间复用(在77个候选时约便宜5倍),但将状态与候选分离。我们探究这种转变在多大程度上保留规则敏感性,以及能否通过训练恢复。在Certo上的四项实验:(1)所测试的向可缓存评分的转换失去了规则敏感性(recall@1 从1.00降至0.24),而联合评分器保持1.00,且短名单+重排的补救措施失败;(2)针对性的反事实监督在留出的合成规则任务(释义、反事实、组合;跨种子可复现)上恢复了强性能,尽管我们未分离出预测是否依赖于所提供的规则;(3)在真实规则上,附加收益未得到证实——在修复截断混淆后,联合评分器在短层级上显著胜出(0.861对0.500),在困难层级上方向性胜出(0.655对0.483,n=29);(4)匹配的跨领域真实散文混合并未帮助,反而降低了合同准确率(-9.3、-16.2个百分点)。可缓存编码器可以在其训练分布上变得规则敏感,但向未见来源的真实规则的迁移未得到证实;联合评分器以牺牲缓存为代价保持优势。

英文摘要

Certo is a small non-generative decision model (Qwen3-4B): it scores candidate actions from their text and returns a probability, instead of generating an answer. The accurate design reads the state, the rules, and each candidate together (a joint scorer), so cost grows with the menu. Independent encoding lets each candidate be encoded once and reused across states (about 5x cheaper at 77 candidates), but separates state from candidate. We ask how much rule-sensitivity survives that move, and whether it can be trained back. Four experiments on Certo: (1) the tested conversion to cacheable scoring loses rule-sensitivity (recall@1 1.00 -> 0.24) while the joint scorer holds 1.00, and a shortlist+rerank rescue fails; (2) targeted counterfactual supervision restores strong performance on held-out synthetic rule tasks (paraphrase, counterfactual, composition; reproducible across seeds), though we do not isolate whether predictions depend on the supplied rule; (3) on real rules the added benefit is not established -- after fixing a truncation confound, the joint scorer wins significantly on the short tier (0.861 vs 0.500) and directionally on the hard tier (0.655 vs 0.483, n=29); (4) a matched cross-domain real-prose mixture did not help and reduced contract accuracy (-9.3, -16.2 points). A cacheable encoder can be made rule-sensitive on its training distribution, but transfer to unseen-source real rules is not established; the joint scorer keeps an edge at the cost of caching.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑