A counterexample in a public discussion changed how I think about AI-agent boundaries. The person said the boundaries written up front were all incomplete; the stop points that lasted were added after something went through that should not have.
That suggests the durable artifact is not only the first brief. It is the amendment history—and whether adding one rule after an incident is cheap enough that someone will actually do it.
I think that creates two separate requirements:
For example:
A useful approval record should therefore say:
If the audience, cost, permissions, data access, target, or reversibility changes, the old “yes” should expire. If a live failure reveals a missing stop condition, adding that condition should not require rebuilding the whole system.
I still have not observed someone using my own brief in a live agent run, so I cannot claim that this design changes behavior yet.
For people running AI agents: which boundary did you add only after something slipped through? And where does that rule live so the next session actually sees it?
Please keep examples redacted—no credentials, customer data, source code, or production details.
The amendment-history point is particularly interesting. A boundary can look complete on paper and still be missing the condition that only a real failure exposes. I’d be curious what you learn once you see people actually using the brief in live agent runs.