Every team we onboard describes its voice in adjectives — warm, confident, not corporate. Adjectives do not survive contact with a language model. We moved to a written spec with banned constructions, worked examples and a diff against your last fifty published pieces, and off-brand drafts dropped by two thirds.
For two releases our agents could post directly to a connected channel. It worked almost every time, which turned out to be the problem: the failures were rare enough that nobody was watching for them. Approval is now the default and publishing is the exception you opt into, per channel.
We looked at every draft edited before publishing and asked what people actually changed. Almost none of it was facts. It was hedging, throat-clearing openers, and a specific kind of false enthusiasm that models reach for when they have nothing to say. Here is the list, and what we did to the prompt layer.
Most attribution tools are very good at recording what happened and very bad at telling you what caused it. Adding more tracking makes the first part better and the second part worse. We think the honest version reports a range and its assumptions, so this is what our attribution view shows and what it refuses to claim.
We shipped six modules and spent a year learning that none of them matter much on their own. What makes a draft usable is that it was grounded in something true about your business. The generation is commodity; the grounding is not. That reframing changed what we build and the order we build it in.
A general assistant that can do anything is impossible to review, because you never know what it was supposed to do. Naming an agent Brand Editor or Pipeline Watcher sounds like a marketing decision. It is really a scoping decision: a title implies a remit, a remit implies a capability list, and a capability list is something you can audit.
Generating a week of posts takes seconds. Deciding when each one goes out, on which channel, without colliding with a campaign or a holiday or another team's launch, is the part that actually takes judgement. We rebuilt the calendar around that problem and stopped treating scheduling as a field on a draft.
A score between nought and a hundred is worthless if nobody can say why. We rebuilt scoring so every point is attributable to a signal you can see and switch off, and so a score that moves tells you which signal moved it. Accuracy went down slightly. Use went up a lot.
Every channel has its own idea of what a comment is, how fast it expects a reply, and whether edits are allowed. Flattening them into one queue is easy. Keeping the queue honest about which threads are actually urgent, when each platform reports timestamps differently, took three attempts.
First it was strings. Then it was templates with slots. Then it was a small compiler that resolves voice, knowledge scope and channel constraints into one prompt and can explain what it produced. Only the third one let us debug a bad draft without guessing. The two failures are the interesting part.
RIAVRichard and Avi·
We use cookies
We use strictly necessary cookies to run the site, and optional cookies to understand usage and measure marketing. You can change your choice at any time. Cookie policy