From Benchmark to Improvements

Methodology: observation → hypothesis → experiment → solution → iteration. Reading the report: a table by tasks + a capability labels matrix reveals structural weaknesses (failures in transcription, mathcounting, complexui — different symptoms of different defects). Hypotheses on three levels: superficial (prompt), medium (input pipeline, reasoning), deep (model change, UI element tree). The solution is not to “accept everything effective”, but to analyze costs and benefits: global reasoning gives +3% success at the cost of triple the latency for 8% of tasks — reject. First, check the evaluation system itself, then touch the agent.

Related: [Three levels of evaluation system], [Statistical significance of evaluation], [Agent observability]