Dynamic Tool Loading
Dynamic loading should not break KV Cache: the full schema of the new tool is appended to the context (trajectory), the static prefix remains stable.
Appending does not invalidate the cache: causal attention means that the K and V of already cached tokens do not change. The schema is fixed at the point of first appearance and then cached as normal history (not re-appended to the end every round).
API level support:
- OpenAI — toolsearch + deferloading;
- Anthropic — tool_reference;
- Codex CLI — BM25 search.
The “define tools only at the beginning” rule is no longer a hard rule.
Related: Proactive Tool Discovery, KV Cache, MCP Protocol