发表机构
Radiant Institute for Manifold Studies(流形研究光辉研究所)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本文探讨AI代理如何判断应遵循的规则,测试向量几何方法未获可靠结果,提出构建后果图以连接行动者、权威与政策条件,帮助代理区分行动、停止或审查。
AI 中文摘要
人工智能代理如何确定应遵循哪些规则?一条规则允许某个行动,另一条规则则施加条件、例外或冲突义务。确定性系统在规则已被明确指定时可以解决这些关系。当这些关系仍隐含在语言中时,代理可能遵循一条规则而忽略另一条本应阻止它的规则。拒绝所有未解决的行动可以避免这种风险,但也会阻止允许的行动。我们希望代理能够做出区分并仍然行动。我们最初的假设是,几何测量可以为判断提供基础。我们将行动和政策表示为向量,然后测试它们的几何结构是否能够识别主导政策并解释行动与这些政策的关系。在四项研究中,所测试的方法未能建立可靠的行动前判断。在最后的合成研究中,一个词汇路由器恢复了每一个主导和阻止政策,同时将中位政策检查减少了97.7%。然而,组合流水线仍然升级了所有2304个测试行动,包括那些本应允许的行动。将每个政策提供给相同的下游机制并未改变任何决策。找到政策并未解决解释它们的问题。这一结果促使我们修正假设:AI代理中的判断需要构建一个后果图。这样的图将把行动者和权威连接到政策条件、例外以及行动将产生的变化。后续研究将探讨使这些关系明确化是否有助于代理区分何时行动、停止或寻求审查。
英文摘要
How can an AI agent determine what rules to follow? One rule permits an action. Another imposes a condition, exception, or conflicting obligation. Deterministic systems can resolve those relationships when they have been specified. When they remain implicit in language, an agent can follow one rule while missing another that should stop it. Refusing every unresolved action avoids that risk, but also blocks permissible actions. We wanted the agent to make the distinction and still act. Our initial hypothesis was that geometric measurements could supply a basis for judgment. We represented actions and policies as vectors, then tested whether their geometry could identify governing policies and interpret the action's relation to them. Across four studies, the tested approaches did not establish reliable pre-action judgment. In the final synthetic study, a lexical router recovered every governing and blocking policy while reducing median policy checks by 97.7%. The composed pipeline nevertheless escalated all 2,304 test actions, including those it should have allowed. Supplying every policy to the same downstream mechanism changed no decision. Finding the policies had not solved the problem of interpreting them. This result led us to revise our hypothesis: judgment in AI agents requires developing a consequence graph. Such a graph would connect the actor and authority to policy conditions, exceptions, and the changes an action would produce. Follow-on studies will ask whether making those relationships explicit helps the agent distinguish when to act, stop, or seek review.
Comments16 pages, 6 figures. Four bounded studies of semantic measurements for pre-action judgment; consequence-graph hypothesis remains untested