Sidecar Mechanism

Sidecar – a lightweight LLM call, parallel to the main reasoning, acting as a gatekeeper for each tool call: a dangerous operation will not execute until Sidecar allows it.

The key decision – Sidecar sees only structured call data (tool name, parameters), not free-form reasoning text: otherwise, an attacker via prompt injection could manipulate the evaluation.

Difference from proposer-reviewer:

  • Sidecar – real-time classification of structured data (a lightweight model will suffice);
  • proposer-reviewer – verification of soundness before/after operation.

Related: Proposer-reviewer, Five Functions of Harness, MCP Security Risks