发表机构
Monash University(莫纳什大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本文提出一种通过递归轨迹抽象层次来学习和测试语言智能体行为假设的方法,以支持跨任务和模型的科学化研究。
AI 中文摘要
语言智能体的科学研究需要能够跨任务和模型支持假设的行为变量。我们将这一研究问题表述为学习和测试轨迹抽象层次结构。一个具体的递归程序首先测量角色和阶段索引事件,提出时间约束关系,并测试它们在各种条件下的稳定性。然后,它从选定的关系构建情节级主题变量,并重复对这些变量的分析。显式测量函数将每个抽象级别连接到原始轨迹。观察和随机化协议实验评估所得假设,而干预实现之间的比较确定抽象应保留、细化还是限制。我们推导了接受约简的有限深度界限,识别了固定抽象上的协议效应,并表征了实现分歧和抽象误差的组成。一个有限样本测试使投影干预一致性可操作,构造的示例说明了主题构建和抽象细化。该公式将这种实验方法与语义分类学、定性理论归纳和行为模型恢复区分开来。它提出了一个用于发现可泛化行为假设的研究程序,其中文献相对新颖性与模型相对惊喜度分别评估。
英文摘要
Scientific studies of language agents need behavioral variables that support hypotheses across tasks and models. We formulate this research problem as learning and testing a hierarchy of trajectory abstractions. A concrete recursive procedure first measures role- and phase-indexed events, proposes temporally constrained relations, and tests their stability across conditions. It then constructs episode-level motif variables from selected relations and repeats the analysis on those variables. Explicit measurement functions connect every abstraction level to the original trajectories. Observations and randomized protocol experiments assess the resulting hypotheses, while comparisons between intervention realizations determine whether an abstraction should be retained, refined, or restricted. We derive a finite-depth bound for accepted reductions, identify protocol effects on fixed abstractions, and characterize realization disagreement and composition of abstraction error. A finite-sample test makes projected intervention consistency operational, and constructed examples illustrate motif construction and abstraction refinement. The formulation distinguishes this experimental approach from semantic taxonomies, qualitative theory induction, and behavior-model recovery. It specifies a proposed research procedure for discovering generalizable behavioral hypotheses, with literature-relative novelty assessed separately from model-relative surprise.
Comments(Work in Progress) 13 pages, 2 figures