Agent Observability

Observability – the ability to infer the internal state of a system from external signals (logs, metrics, traces).

Trace – a single run of a task; each LLM call, tool call, search is a span with input, output, time, tokens, errors; parent-child relationships form a tree. Standards: OpenTelemetry + OpenInference.

Most valuable application – feeding data back into evaluation assets: failed production cases → de-personalization → new evaluation dataset examples & regression tests.

Observability is about “seeing”, evaluation is about enshrining that in verifiable standards.

Related: Benchmark to Improvement, Three Layers of Evaluation System