For two releases, an agent could publish straight to a connected channel with no human in the loop. It worked almost every time. That turned out to be the problem.
Reliability is not the same as safety
The publish path was succeeding well over ninety-nine percent of the time. At that rate nobody watches it. The failures were rare enough to be invisible day to day, and consequential enough that every one of them mattered.
A system that fails loudly and often gets attention and gets fixed. A system that fails quietly and rarely accumulates a small number of very bad outcomes and nobody notices the pattern.
What the failures looked like
None of them were the failure mode we designed for. We had guarded against the model producing something obviously wrong. What actually went out was subtly stale: a promotion that had ended, a price that had changed, a person who had left.
- Facts that were true when the source was ingested and false by the time the draft published
- Correct content published to the wrong channel because the schedule collided
- A reply that read as dismissive because it lacked context only a human had
Approval as the default
Publishing is now something you opt into, per channel, rather than the default behaviour. Everything an agent produces lands as a draft with an approver attached.
We expected complaints about the extra step. We got almost none, which suggests the autonomy was never the thing people valued. What they wanted was not writing the first draft.
The part we got wrong
We should have shipped the audit trail before the autonomy, not after. Once you can see what an agent did and why, the argument about how much rope to give it becomes concrete instead of theoretical.