Agent Observability
Observability – the ability to infer the internal state of a system from external signals (logs, metrics, traces).
Trace – a single run of a task; each LLM call, tool call, search is a span with input, output, time, tokens, errors; parent-child relationships form a tree. Standards: OpenTelemetry + OpenInference.
Most valuable application – feeding data back into evaluation assets: failed production cases → de-personalization → new evaluation dataset examples & regression tests.
Observability is about “seeing”, evaluation is about enshrining that in verifiable standards.
Related: Benchmark to Improvement, Three Layers of Evaluation System