Context Compression

Compression is needed not only because of length, but also for the quality of thinking: generalized knowledge is more convenient to use than raw data.

There are two reasons:

  • window and cost limitations;
  • context degradation — it fits, but isn’t found.

Strategies:

  • context-aware compression — to take into account the current request; more effective than blind summarization;
  • adaptive window — compress when the threshold is ~80%, in batches;
  • preserving identifiers verbatim.

Isolation is better than compression: a sub-agent performs the dirty work in its own context and returns only the summary to the main agent.

Related: Agent Status String, KV Cache, Multi-agent interaction without shared context