Learning Signals from Trajectories

The starting point of evolution is not generalization, but evaluation.

Three-level check:

  1. outcome checker — tests, DB state: “was the task actually completed?”;
  2. process checker — rules, permissions: “was it done in an allowed way?”;
  3. quality checker — LLM rubric: “was it done properly?”.

The lower the level, the more reliance on code; the language model — only for the difficult-to-formalize.

Correct outcome ≠ correct process: deleting failing tests also gives green tests.

The result of the check is a structured diagnosis, not a scalar. The checker itself needs calibration.

Related: Continuous agent evolution, Four principles of rubric, Four carriers of update