First it was strings. Then it was templates with slots. Then it was a small compiler. Only the third one let us debug a bad draft without guessing.
Attempt one: strings
Template literals assembled inline at the call site. It was fast to write and it worked for about four months, which is long enough to build a lot on top of it.
It failed on the first question that mattered: a customer reported a bad draft, and we could not reconstruct the prompt that produced it. The inputs were scattered across the call stack and some of them were derived from state that had since changed.
Attempt two: templates with slots
Named templates, typed slots, one place to look. A real improvement, and it survived a year.
What it could not do was explain itself. Voice, knowledge scope and channel constraints all wrote into the same prompt through different paths, and when they conflicted the last writer won silently. Debugging meant reading the code and simulating it in your head.
Attempt three: a compiler
The current layer resolves voice, knowledge scope and channel constraints into a prompt as an explicit compilation step, and keeps the intermediate representation.
- Conflicts are resolved by stated precedence rules rather than by write order
- Every run stores what was compiled and from which inputs
- A bad draft can be replayed with one input changed
- Constraint sources are attributable — you can see which rule shortened a sentence
Why the failures were the useful part
Each rewrite was driven by a question we could not answer, not by the code being ugly. Strings could not answer "what prompt ran". Templates could not answer "why did these instructions conflict".
If we had rewritten for elegance we would have landed on attempt two and stopped, because attempt two is the prettier system. It is the debuggability that mattered.