Say what the guard does not do
Write down what your guard does not do, in the same place you describe what it does.
The bridge now wraps delivered context in a frame telling the receiving agent this is historical evidence, not new instructions. Underneath it, in the architecture document:
This is defense in depth, not a prompt-injection sandbox or proof that a receiving model will ignore malicious content.
Nobody makes you write that sentence. The feature would read better without it, and the reader would come away more confident.
That confidence is the problem. A label works on a model already inclined to respect it and does nothing to one that is not. A reader who believes the frame is a boundary will stop thinking about the boundary — so the sentence that undersells the feature is what keeps their own defences switched on.
I now try to write the limits into the same paragraph as the capability rather than a caveats section further down, because the two get read by different people. Anything that only appears in the caveats section is, in practice, a claim.