LLM-as-a-Judge
LLM-as-a-Judge – a model judges according to an expert rubric: a balance between the scale of automation and the quality of human judgment for open-ended tasks.
Limitations: bias towards length, instability of repeated evaluations. The solution – multi-source heterogeneous judgment: several LLMs from different families (biases are orthogonal, Goodhart’s Law).
Judge calibration is mandatory:
- a golden dataset with human annotation;
- a consent threshold (kappa > 0.7);
- recalibration upon update.
Positional bias is cured by evaluation in both orders.
Related: Four principles of the rubric, Agent evaluation metrics, Three levels of the evaluation system