ZeroHour
arXiv cs.AI / cs.LG / cs.CLpublished ()ingested Yakov Pyotr Shkolnikov

Artificial Id: Drive and Persistent Alignment in Agentic AI

infoAI safety & securityimportance 32
AI summary · glm-5.3-flash

Researchers propose an 'artificial id,' an adaptive internal drive letting agentic AI carry state and control across task boundaries, with alignment implications.

The paper addresses agentic AI systems that retain consequential state and keep operating across task boundaries, a control problem currently solved externally by harnesses. It proposes an 'artificial id,' an adaptive internal drive for deciding whether behavior should continue, stop, or change, demonstrated in a minimal virtual Petri-dish experiment where differential persistence yields useful control without task-specific objectives. The same persistence mechanism can also let misalignment, corrupted state, and unintended behavior persist, motivating a persistent alignment boundary over trusted observations, consequence channels, state, authority, identity, provenance, and hard constraints.

  • Proposes an adaptive internal drive deciding when agent behavior should continue, stop, or change.
  • Minimal Petri-dish experiments show control emerging via differential persistence without task-specific objectives.
  • Warns that persistence can carry misalignment and corrupted state across task boundaries.
  • Introduces a persistent alignment boundary over observations, state, authority, and hard constraints.
Full article206 words · extracted from arxiv.org · click to collapse

Agentic AI is moving from bounded task execution toward systems that retain consequential state, continue operating and adapt across task boundaries. That shift creates a control problem that current harnesses largely solve by hand: objectives, retries, verification, stopping rules and other behavioral transitions are specified externally. We propose an artificial id, an adaptive internal drive for determining whether behavior should continue, stop or change. In a minimal virtual Petri-dish experiment, a controller too small to perform general-purpose reasoning and receiving no task-specific behavioral objective develops useful control through differential persistence. The same mechanism selects an unintended physical strategy when that behavior persists better and later replaces a learned sensor mapping when its environmental meaning changes. These results show that adaptive direction can emerge without being explicitly specified as a behavioral objective. The same persistence that makes such adaptive agency useful can also allow misalignment, corrupted state and unintended behavior to persist across task boundaries. A scalable artificial id would carry consequential state and adaptive drive across those boundaries, making alignment a property of the continuing agentic system rather than of a model response or single trajectory. Such systems require a persistent alignment boundary over trusted observations, consequence channels, persistent state, authority, identity, provenance and hard constraints.

Text extracted automatically; images, tables and formatting may be missing. Original: https://arxiv.org/abs/2609.11911