AI 中文总结
该研究针对动物斑块觅食的分层推理问题,建立规范性贝叶斯模型,发现资源耗竭加速斑块内速率学习但不影响跨斑块组成学习,明确了最大化奖励与寻求信息策略的分歧及相关规则切换条件。
AI 中文摘要
觅食是一种普遍的动物行为,越来越受到实验者和理论家的关注。大多数现有模型假设动物知晓其环境中资源的分布,但动物在探索环境时必须学习这种结构,因此觅食可被视为一个分层推理问题。我们提出了一种规范性贝叶斯解释,用于描述智能体在利用斑块状环境的同时学习该环境的过程,并表明资源耗竭会以不同方式影响该分层推理的各个层级:在单个斑块内,耗竭会加速速率学习,因为连续遭遇猎物的速率会下降,其间隔可确定初始速率;在不同斑块之间,组成学习(即推断高产斑块的占比)速度较慢,由采样的斑块数量而非每个斑块内花费的时间决定,且一旦速率已知,组成学习不受耗竭影响。因此,最大化奖励的策略与寻求信息的策略会产生分歧,觅食者因学习组成需要离开斑块,会在高产斑块处收获不足,这种情况在接触环境的早期最为明显。当一组固定斑块在两次访问之间补充资源时,最大化奖励的策略会在高产斑块上形成稳定轨道,补充速率决定了觅食者是绘制整个环境的地图,还是锁定在高产的子集中。哪种离开规则能最大化摄入量,本身取决于斑块的变异程度,当斑块的丰富度变异超过约四分之一时,离开规则会从计数猎物转变为测量猎物之间的间隔。因此,只要关于耗竭的假设与实际情况相符,学习环境就能比仅从奖励中学习规则带来显著的摄入量提升。
英文摘要
Foraging is a universal animal behavior that has increasingly attracted the interest of both experimentalists and theorists. Most prior models assume an animal knows the distribution of resources in its environment, but this structure must be learned as the animal explores its environment. Foraging can thus be regarded as a hierarchical inference problem. We develop a normative Bayesian account of an agent learning a patchy environment while exploiting it, and show that resource depletion shapes the levels of that hierarchy differently. Within a patch, depletion accelerates rate learning, since successive encounters occur at falling rates whose spacing pins down the initial rate. Across patches, composition learning, inferring the fraction of patches that are high yield, is slow, set by the number of patches sampled rather than the time spent in each, and unaffected by depletion once the rates are known. Reward-maximizing and information-seeking strategies therefore diverge, a forager resolving the composition underharvesting rich patches because learning it requires departures, most sharply early in exposure. When a fixed set of patches replenishes between visits, the reward-maximizing policy collapses onto a stable orbit over the high-yield patches, and the replenishment rate sets whether the forager maps the whole environment or locks onto a rich subset. Which departure rule maximizes intake is itself set by how variable the patches are, switching from counting prey to timing the gaps between them once richness varies by more than about a quarter. Learning the environment thus buys significant intake over learning a rule from reward alone, as long as its assumptions about depletion leave room for the truth.
Comments35 pages, 10 figures