Your Brand Has No Training Data
· Updated
- A brand system is two halves. The rules everyone writes, and the examples almost nobody keeps. The second half is the one a model actually learns from.
- Rules tell a model what to do. Examples show it what good looks like when the rule runs out, which is where most of the drift happens.
- Keep the off brand specimens too. A model learns a boundary from both sides of it, and the rejected version with the reason is worth as much as the approved one.
- You already generate the corpus every week and throw it away. Every approval, every kill, every this not that in a review is training data, and it evaporates with the meeting.
A while back I wrote that your brand system is the only thing keeping AI honest, and I still believe every word of it. A model does not drift because it is careless. It mostly drifts because it hits a gap where nothing was written down and fills it with whatever is likely. A tight system of rules, tokens, and hard nos closes those gaps and makes the thousand small calls for you. That argument stands. It is also only half of the system.
The other half is the one almost nobody keeps. Not a handful of voice examples parked in the guidelines, but every call you have already made about what clears the bar and what does not.
What is the part of a brand system everyone forgets to keep?
The examples. Rules are the part everyone writes, because rules feel like the system: the tokens, the spacing scale, the hard nos. But a rule only covers the situation you anticipated, and the situations you did not anticipate are where models tend to drift. Your standard is also the body of work you have already judged, cleared and killed, that shows what the rules mean when they run out.
I keep coming back to how a machine learns versus how a person learns. You can hand a new hire every rule in the book, and the first brief still comes back compliant and off the mark, because rules describe the center of the target and the work happens at the edges. What fixes them is seeing your actual work. This one shipped, that one got killed, here is why. A model is the same, only it never sits in the room. It learns your brand from specimens, or it guesses at the edges.
Why does a model need examples if it already has the rules?
Because a rule tells it what to do and an example shows it what good looks like when the rule stops being enough. The moment a situation is a little off from what you wrote, the model has to infer, and that is where the drift starts. A paired example, this version yes and that version no, gives it something to reason from instead of a blank.
This is the same reason I wrote in that earlier piece that you have to write the reason down next to the rule. Examples are that idea taken all the way. A rule with its reasoning is one data point. A pile of specimens, each labeled with why it went the way it did, is a training set. It is the difference between telling someone your taste and showing them a season of your calls. One is an assertion. The other is evidence a machine can pattern match against.
Why does almost no company have a brand corpus?
Because the examples are expensive to keep and nobody is assigned to keep them. Rules get written once and live in a document. Examples get generated every week, in every review, and thrown away the moment the decision is made. The approved layout ships, the rejected ones vanish, and the note explaining the call evaporates with the meeting.
You are producing the most valuable training data you own and deleting it in real time. I have built and led creative teams for most of my career, and the specimens I would most want a model to learn from are not in any guideline. They are in the history of what we approved and what we sent back, and the reasons attached to each. That is the corpus. In the rooms I have worked in, almost no one has it, because saving it is unglamorous work that no role owns. The reasons matter as much as the pictures. Keep the off brand ones too, because a model learns a boundary from both sides of it, and a rejected version with a good reason teaches as much as a clean example alone.
What does keeping a brand corpus actually look like?
It looks like refusing to throw away the decisions you already make. You do not have to build a dataset from scratch. The raw material is already moving past you, in every review and every approval, and the only thing missing is someone catching it on the way by, with the reason still attached.
A few concrete moves:
- Save the approved work and the killed work in the same place, not just what shipped.
- Write one line on each about why it went the way it did, while the reason is still fresh.
- Treat every review as a labeling session, because that is what it is, and record the calls.
- Pair the specimens. This, not that, is worth more than either one alone.
Do that for a season and you have something no rulebook gives you: your judgment, shown rather than described, in a form a model can learn from.
That earlier piece argued that a written system is what keeps AI honest, and it is. The rules are the map. The examples are the terrain, and a model that has only ever seen the map will get lost the first time the ground disagrees with it. Good creative is judgment made visible. Your rules make a fraction of it visible. Your specimens make the rest.
Frequently asked
What is brand training data?
It is the corpus of your own work labeled by whether it hit the standard, with the reasoning attached. Not the guidelines, the specimens. The approved layout and the rejected one side by side, the note that says why one cleared the bar and the other did not. Showing a model examples is one of the surest ways to steer it, and this is the example set almost no company keeps, even though they produce it constantly.
Do rules or examples matter more for keeping AI on brand?
Both, and they do different jobs. Rules cover the cases you anticipated. Examples cover the ones you did not, which is where most of the drift lives. A rule tells the model the spacing token. A paired example shows it what to do when the situation is a little off from the rule, and that judgment is the part you cannot fully write down. Keep both or the system has a blind spot exactly where it is tested.
How do I start building a brand corpus?
Stop throwing away the decisions you already make. Save the approved work and the killed work in one place, and write one line on each about why it went the way it did. A review meeting is a labeling session you are not recording. Do that for a season and you have a specimen set with reasoning attached, which is worth more to a model than any deck of rules on its own.