发表机构
Saarland University; German Research Center for Artificial Intelligence (DFKI)(萨尔兰大学; 德国人工智能研究中心(DFKI))
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对单一躯体智能体无法学习磨损的问题,提出负荷门控痛觉通道与记忆机制,证明其将分配转向未感最佳有偿工作,并在模拟中延长工作寿命4.4年且无损失,机制图双向验证价值。
AI 中文摘要
部署在单一躯体中的智能体无法了解该躯体的磨损速度,因为每一次会揭示其耐磨性的试验都会磨损它所保护的躯体。我们研究这种“纪元一”设置,其中固定权重策略的参数在躯体被抽取之前就已设定,且在生命周期中从不更新。智能体携带一个由负荷门控的痛觉通道和一个保留所感之感的记忆。我们证明,所感成本将分配转移到尚未被感受的最佳“有偿”工作,而非最温和的工作;没有记忆保留的智能体永远不会看到所感成本约束生效;该通道仅在威胁是个体不可预测、易于避免且忽视代价高昂之处才产生作用。我们按躯体进行测量,将有通道和记忆的智能体与没有它们的同一智能体个体进行对比,其中两者都不携带跨生命周期学习的调度方案。在2000个模拟的铺地板工膝盖上,以已发表的损耗率为锚定,感受、保留和替代将工作寿命从55.2岁延长至59.6岁,并将职业产出从33.7提升至36.1。69.3%的躯体获得收益,且没有躯体受损。一个能感受但当日不保留任何信息的躯体获得+4.4年中的一年,而保留则承担其余部分。一个经群体训练的智能体从同一通道获得+0.65年的收益,但产出为-0.54。差异在于物种先验已经提供的东西,而单一躯体则没有。两者通过一个恒等式相关联,消融均值报告了每躯体价值的(1-χ)倍,其中χ是盲目调度已捕获的份额,因此我们同时报告两者。在机制图预测价值的地方,护理机器人的认证服务寿命增加六倍,而现场锚定的机队报废0.15台机器而非0.55台。在它预测无价值的地方,漫游车相对于盲目谨慎几乎无增益,因此该图在两个方向上都成立。
英文摘要
An agent deployed in a single body cannot learn how fast that body wears, because every trial that would reveal its wear resistance wears the body it would protect. We study this \emph{epoch-one} setting, in which the parameters of a fixed-weight policy are set before the body is drawn and never updated in life. The agent carries a load-gated nociceptive channel and a memory that retains what was felt. We prove that felt cost moves the allocation to the best-\emph{paid} work not yet felt rather than the gentlest, that an agent without retention never sees the felt-cost constraint bind, and that the channel pays only where the threat is individually unpredictable, cheap to avoid and expensive to ignore. We measure per body, setting the agent with channel and memory against the same individual without them, where neither carries a schedule learned across lives. On $2{,}000$ simulated floor-layer knees, with wear anchored to published loss rates, feeling, retaining and substituting extends the working life from age $55.2$ to $59.6$ and raises career output from $33.7$ to $36.1$. $69.3\%$ of bodies gain and \textbf{none lose}. A body that feels but retains nothing past the day gains one of the $+4.4$ years, and retention carries the rest. A population-trained agent gains $+0.65$ years from the same channel at $-0.54$ output. The difference is what a species prior already supplies, and a single body has none. The two are related by an identity, the ablation mean reporting $(1-χ)$ of the per-body value with $χ$ the share a blind schedule already captures, so we report both. Where the regime map predicts value, a care robot sextuples its certified service life and a field-anchored fleet writes off $0.15$ of its machines instead of $0.55$. Where it predicts none, a rover gains little over blind caution, so the map holds in both directions.