Learning Signals from Trajectories
The starting point of evolution is not generalization, but evaluation.
Three-level check:
- outcome checker — tests, DB state: “was the task actually completed?”;
- process checker — rules, permissions: “was it done in an allowed way?”;
- quality checker — LLM rubric: “was it done properly?”.
The lower the level, the more reliance on code; the language model — only for the difficult-to-formalize.
Correct outcome ≠ correct process: deleting failing tests also gives green tests.
The result of the check is a structured diagnosis, not a scalar. The checker itself needs calibration.
Related: Continuous agent evolution, Four principles of rubric, Four carriers of update