Safety Boundaries of Self-Modification

Self-evolution turns a single mistake into a long-term risk.

Three boundaries:

  1. Separation of evidence and instructions – raw external content is not written directly into the Skill, only via LLM summarization with version control;
  2. Candidate and official capabilities – new artifacts first in a candidate space off of real traffic, with sandboxing and supply chain scanning;
  3. Prohibition of self-modifying safety mechanisms – the agent does not change the checks, tests, release thresholds, audit logs, and backups asserting its own updates.

Otherwise, it’s enough to lower the bar – and degradation looks like progress.

Related: Check - Release - Revert, Deadly Triad, Prompt Injection