Safety Boundaries of Self-Modification
Self-evolution turns a single mistake into a long-term risk.
Three boundaries:
- Separation of evidence and instructions – raw external content is not written directly into the Skill, only via LLM summarization with version control;
- Candidate and official capabilities – new artifacts first in a candidate space off of real traffic, with sandboxing and supply chain scanning;
- Prohibition of self-modifying safety mechanisms – the agent does not change the checks, tests, release thresholds, audit logs, and backups asserting its own updates.
Otherwise, it’s enough to lower the bar – and degradation looks like progress.
Related: Check - Release - Revert, Deadly Triad, Prompt Injection