Statistical Significance of Evaluation

Observed difference may be sampling noise. Standard error of binomial distribution √(p(1-p)/n): with 100 examples and 70% success, noise band ±9 points — a 3% difference is not significant.

Correct approach for a single set of tasks is paired analysis (McNemar’s test: look only at examples where results diverged) — more sensitive than subtracting independent proportions.

Multiple comparisons trap: with 6 hypotheses, probability of false positive conclusion ~26% — Bonferroni correction or independent re-check is needed.

Multiple runs with different seeds are mandatory.

Related: Evaluation Metrics of Agents, Three Levels of Evaluation System