Dynamic Tool Loading

Dynamic loading should not break KV Cache: the full schema of the new tool is appended to the context (trajectory), the static prefix remains stable.

Appending does not invalidate the cache: causal attention means that the K and V of already cached tokens do not change. The schema is fixed at the point of first appearance and then cached as normal history (not re-appended to the end every round).

API level support:

  • OpenAI — toolsearch + deferloading;
  • Anthropic — tool_reference;
  • Codex CLI — BM25 search.

The “define tools only at the beginning” rule is no longer a hard rule.

Related: Proactive Tool Discovery, KV Cache, MCP Protocol