发表机构
Institute of Cognitive Sciences and Technologies (ISTC-CNR); Interdepartmental Center for Research Ethics and Integrity (CID Ethics)(认知科学与技术研究所(ISTC-CNR); 跨部门研究伦理与诚信中心(CID Ethics))
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对大规模模型在动态环境中对齐不足的问题,提出基于内在动机、经验性规范学习及监管沙盒教学环境的发展性框架,以逐步塑造自主智能体的对齐行为。
AI 中文摘要
近年来,得益于能够泛化并生成复杂输出的大规模模型,人工智能取得了非凡的进展。然而,将这种潜力转化为具身智能体却揭示了一个显著局限:最先进的系统依赖于预先存在的数据集和人类反馈策略,这些策略虽然强大,但在动态或未知情境中却显得不足。为了适应环境,智能体必须通过与环境的直接互动来获取知识。应对这一挑战的一种策略是引入更高层次的机制,例如内在动机,它利用好奇心和能力感来引导在复杂环境中的探索与学习。虽然这种灵活性扩展了自主性,但它也使确保智能体与人类目标保持一致的任务复杂化。对齐,对于一般的人工系统而言本已是挑战,在非结构化和动态情境中变得更加复杂,因为在这些情境中,预定义的规则被证明是不够的。为了有效且可适应,规范必须通过一个认识论过程植根于经验,该过程从简单、情境化的原则出发,允许通过经验、自主学习和与其他道德智能体的合作,逐步构建更复杂的规则。类似于儿童通过探索环境和参与集体实践来学习社会规范,人工智能体也必须接受朝向对齐的教育。遵循丹尼特的观点,道德智能体的地位并非天生,而是根据其负责任地管理日益增长的自由度的能力而逐步被赋予的。从这个角度来看,监管沙盒可以被视为人工智能的教学环境:动态空间,其中对齐作为一种形成性过程发展,通过在日益复杂的情境中的互动与合作,逐步塑造自主行为。
英文摘要
In recent years, artificial intelligence has made extraordinary progress thanks to large-scale models capable of generalization and the generation of complex outputs. However, transferring this potential into embodied agents reveals a significant limitation: the most advanced systems rely on pre-existing datasets and human feedback strategies that are powerful but insufficient in dynamic or unknown contexts. To adapt, an agent must acquire knowledge through direct interaction with its environment. One strategy to address this challenge involves introducing higher-level mechanisms, such as intrinsic motivations, which leverage curiosity and competence, to guide exploration and learning in complex environments. While this flexibility expands autonomy, it complicates the task of ensuring agents remain aligned with human goals. Alignment, already a challenge for artificial systems in general, becomes even more complex in unstructured and dynamic contexts where predefined rules prove insufficient. To be effective and adaptable, norms must be rooted in experience through an epistemological process that starting from simple, situated principles allows for the gradual construction of more complex rules through experience, autonomous learning, and cooperation with other moral agents. Similarly to children learning social norms by exploring their environment and participating in collective practices, artificial agents must also be educated toward alignment. Following Dennett, the status of a moral agent is not innate but is attributed gradually based on the ability to responsibly manage increasing degrees of freedom. From this perspective, the regulatory sandboxes can be viewed as pedagogical environments for AI: dynamic spaces where alignment develops as a formative process, progressively shaping autonomous behaviors through interaction and cooperation in scenarios of increasing complexity.
CommentsIn publication in the proceedings of SIpEIA 2026 conference