Statistical Significance of Evaluation
Observed difference may be sampling noise. Standard error of binomial distribution √(p(1-p)/n): with 100 examples and 70% success, noise band ±9 points — a 3% difference is not significant.
Correct approach for a single set of tasks is paired analysis (McNemar’s test: look only at examples where results diverged) — more sensitive than subtracting independent proportions.
Multiple comparisons trap: with 6 hypotheses, probability of false positive conclusion ~26% — Bonferroni correction or independent re-check is needed.
Multiple runs with different seeds are mandatory.
Related: Evaluation Metrics of Agents, Three Levels of Evaluation System