LLM-as-a-Judge

LLM-as-a-Judge – a model judges according to an expert rubric: a balance between the scale of automation and the quality of human judgment for open-ended tasks.

Limitations: bias towards length, instability of repeated evaluations. The solution – multi-source heterogeneous judgment: several LLMs from different families (biases are orthogonal, Goodhart’s Law).

Judge calibration is mandatory:

  • a golden dataset with human annotation;
  • a consent threshold (kappa > 0.7);
  • recalibration upon update.

Positional bias is cured by evaluation in both orders.

Related: Four principles of the rubric, Agent evaluation metrics, Three levels of the evaluation system