From Benchmark to Improvements
Methodology:
observation → hypothesis → experiment → solution → iteration
Reading the report: task table + capability tag matrix reveal structural weaknesses (failures in transcription, mathcounting, complexui — different symptoms of different defects).
Hypotheses of three levels:
- superficial — prompt;
- intermediate — input pipeline, reasoning;
- deep — model change, UI element tree.
Solution — not “accept everything effective”, but cost-benefit analysis: global reasoning gives +3% success at the cost of triple the latency for 8% of tasks — reject.
First check the evaluation system itself, then touch the agent.
Related: Three levels of evaluation system, Statistical significance of evaluation, Agent observability