Sidecar Mechanism
Sidecar – a lightweight LLM call, parallel to the main reasoning, acting as a gatekeeper for each tool call: a dangerous operation will not execute until Sidecar allows it.
The key decision – Sidecar sees only structured call data (tool name, parameters), not free-form reasoning text: otherwise, an attacker via prompt injection could manipulate the evaluation.
Difference from proposer-reviewer:
- Sidecar – real-time classification of structured data (a lightweight model will suffice);
- proposer-reviewer – verification of soundness before/after operation.
Related: Proposer-reviewer, Five Functions of Harness, MCP Security Risks