Your Brand Has No Training Data
- A brand system is two halves. The rules everyone writes, and the examples almost nobody keeps. The second half is the one a model actually learns from.
- Rules tell a model what to do. Examples show it what good looks like when the rule runs out, which is exactly where drift happens.
- Keep the off brand specimens too. A model learns a boundary from both sides of it, and the rejected version with the reason is worth as much as the approved one.
- You already generate the corpus every week and throw it away. Every approval, every kill, every this not that in a review is training data nobody is saving.
A while back I wrote that your brand system is the only thing keeping AI honest, and I still believe every word of it. A model does not drift because it is careless. It drifts because there was no brand written down to hold it, so it fills the gap with whatever is likeliest. A tight system of rules, tokens, and hard nos closes those gaps and makes the thousand small calls for you. That argument stands. This piece picks up exactly where that one stopped.
Because there is a second half to a brand system, and it is the half almost nobody keeps.
What is the part of a brand system everyone forgets?
The examples. Rules are the part everyone writes, because rules feel like the system: the tokens, the spacing scale, the voice do and do not, the hard nos. But a rule only covers the situation you anticipated. The examples cover the ones you did not, and those are the situations where a model actually drifts. Your standard is not just the rules you wrote. It is the body of decisions you have already made, on standard and off, that show what the rules mean when they run out.
I keep coming back to how a machine learns versus how a person learns. You can hand a new hire every rule in the book, and they will still produce something confident and wrong on the first brief, because rules describe the center of the target and the work happens at the edges. What fixes them is seeing your actual work. This one shipped, that one got killed, here is why. A model is the same, only it never sits in the room. It learns your brand from specimens or it guesses.
Why does a model need examples if it already has the rules?
Because a rule tells it what to do and an example shows it what good looks like when the rule stops being enough. Rules handle the anticipated case. The moment a situation is a little off from what you wrote, the model has to infer, and inference from rules alone is close to a coin flip. A paired example, this version yes and that version no, gives it something to reason from instead of a blank.
This is the same reason I wrote in the last piece that you have to write the reason down next to the rule. Examples are that idea taken all the way. A rule with its reasoning is one data point. A pile of specimens, each labeled with why it went the way it did, is a training set. It is the difference between telling someone your taste and showing them a season of your calls. One is an assertion. The other is evidence a machine can pattern match against.
Why does almost no company have this corpus?
Because the examples are expensive to keep and nobody is assigned to keep them. Rules get written once and live in a document. Examples get generated every single week, in every review, and then thrown away the moment the decision is made. The approved layout ships and the three rejected ones vanish. The note that explained why one earned it evaporates with the meeting. You are producing the most valuable training data you own and deleting it in real time.
I lead brand and creative for a company building workforce housing across the Sun Belt, and the specimens I would most want a model to learn from are not in any guideline. They are in the history of what we approved and what we sent back, and the reasons attached to each. That is the corpus. Almost no one has it, not because it is hard to understand but because saving it is unglamorous work that no role owns. The reasons matter as much as the pictures. Keep the off brand ones too, because a model learns a boundary from both sides of it, and a rejected version with a good reason teaches more than a clean example alone.
So what does keeping the corpus actually look like?
It looks like refusing to throw away the decisions you already make. You do not have to build a dataset from scratch. You have to stop discarding one you generate constantly. A few concrete moves:
- Save the approved work and the killed work in the same place, not just what shipped.
- Write one line on each about why it went the way it did, while the reason is still fresh.
- Treat every review as a labeling session, because that is what it is, and record the calls.
- Pair the specimens. This, not that, is worth more than either one alone.
Do that for a season and you have something no rulebook gives you: your judgment, shown rather than described, in a form a model can learn from.
The last piece argued that a written system is what keeps AI honest, and it is. This is the part that was left implicit. The rules are the map. The examples are the terrain, and a model that has only ever seen the map will get lost the first time the ground disagrees with it. Good creative is judgment made visible. Your rules make a fraction of it visible. Your specimens make the rest.
Frequently asked
What is brand training data?
It is the corpus of your own work labeled by whether it hit the standard, with the reasoning attached. Not the guidelines, the specimens. The approved layout and the rejected one side by side, the note that says why one earned the logo and the other did not. A model reasons from examples as much as from rules, and this is the example set almost no company keeps, even though they produce it constantly.
Do rules or examples matter more for keeping AI on brand?
Both, and they do different jobs. Rules cover the cases you anticipated. Examples cover the ones you did not, which is where drift actually lives. A rule tells the model the spacing token. A paired example shows it what to do when the situation is a little off from the rule, and that judgment is the part you cannot fully write down. Keep both or the system has a blind spot exactly where it is tested.
How do I start building a brand corpus?
Stop throwing away the decisions you already make. Save the approved work and the killed work in one place, and write one line on each about why it went the way it did. A review meeting is a labeling session you are not recording. Do that for a season and you have a specimen set with reasoning attached, which is worth more to a model than any deck of rules on its own.