How to Score AI Visual Consistency Before a Set Ships
If consistency cannot be measured, it gets renegotiated on every set. A short rubric settles it in ten minutes and shows exactly what to fix.
21 de setembro de 2026

The hardest part of consistency is that everyone means something slightly different by it. One reviewer is looking at palette, another at light, a third at whether the mood feels right — and the set gets reworked because nobody agreed on what was wrong. A short rubric fixes that. It gives the set a number and, more usefully, a direction for the fix.
Four criteria, and why not more
Score four things. Every additional criterion dilutes attention and reduces agreement between reviewers.
- Palette — are all colours drawn from the approved set, and is the accent within its agreed area?
- Light — does the direction, softness and temperature match the reference?
- Treatment — do surfaces read as specified: matte, grainy, dry, clean?
- Framing — does the subject sit where the composition rule says it sits, with the specified negative space?
Notably absent: beauty, creativity, taste. Those are legitimate conversations, but they belong to approval, not to consistency. Mixing them into the score is how a rubric stops being used.
The 0-2 scale
Keep the scale small enough that two reviewers reach the same answer without negotiation.
- 2 — matches. No difference worth noting.
- 1 — acceptable difference. A visible variation that does not break the set. Write down what it is.
- 0 — breaks the set. The image would look wrong beside the others.
Any score of 1 requires a written note. Without the note, the same difference gets re-argued next month. "Slightly cooler light than reference, acceptable for the outdoor scene" is enough.
Scores of 0 are not negotiable averages. One zero means the image either gets corrected or leaves the set, regardless of how well it scores elsewhere.
How to run the session
Scoring works only if it is fast and blind enough to be honest.
- Fix the reference first. Every criterion is scored against the approved reference, not against the previous image in the set.
- Review in random order if possible. Sequential review biases you toward the last image you saw.
- Judge at thumbnail size first, then at 100%. Thumbnails expose palette and density; full resolution exposes treatment.
- Score every criterion for every image before discussing any of them. Discussing mid-pass anchors the group on the first disagreement.
- Total the columns, not the rows. The useful output is which criterion fails across the set, not which image scores lowest.
Twenty images take about ten minutes. That is faster than the discussion it replaces.
Read the failures by column
Column totals tell you what to fix:
- Palette fails most → the accent rule is unclear, or the palette in the brief lists hues without materials. Add material words: "warm oat paper" rather than "beige".
- Light fails most → the light line is too loose. Specify direction and softness explicitly, and consider whether the prompt is fighting it with an incompatible mood word.
- Treatment fails most → the model or the prompt is adding gloss. Name the surface and the finish in the brief, and remove competing style words.
- Framing fails most → the composition rule is not measurable. Give it numbers: "subject right of centre, 40% empty space, nothing within 8% of the edge".
Fixing the column is how the same defect stops recurring. Fixing individual images keeps the next batch at the same score.
Decide with the score, not around it
Write the decision rule before you score, so the outcome does not depend on how tired everyone is:
- All 2s → publish.
- 1s with notes → publish, and record the notes.
- Any 0 → correct that image or drop it from the set.
- A column average below 1.5 → do not fix images; revise the brief and regenerate the batch.
That last rule saves the most time. Once a whole column fails, individual corrections are rework. The brief is what needs attention.
This is the same principle as fixing batch-level defects in the AI image QA pass: defects shared across many images come from the instructions, not from bad luck.
Align reviewers before relying on one
For the first two or three sets, use two reviewers and compare scores per criterion. Disagreements are informative:
- Disagreement on palette usually means one reviewer is not using the reference.
- Disagreement on treatment usually means the brief's surface words are ambiguous.
- Disagreement on framing usually means the composition rule lacks numbers.
Once reviewers agree within one point on the reference set, a single consistent reviewer is enough for routine work. Keep the rubric and the reference fixed; if either changes, re-align.
Track the trend, not just the set
Log the column totals per campaign. Two patterns are worth watching:
- Improving scores over time means briefs are getting sharper. This is the expected direction after two or three iterations.
- Stable scores with rising revision counts means people are fixing images instead of briefs. The rubric is being used to justify rework rather than prevent it.
Keep the log with the campaign folder, alongside the style pack that produced the set. When a new team member joins, the log shows them what the brand tolerates and what it does not — faster than any style guide.
Score one set this week
Take the most recent batch you approved, score it against the four criteria even if it is already published, and look at the column totals. Most teams discover that one criterion accounts for the majority of their inconsistency, and that fixing it in the brief changes the next batch more than any amount of extra review.
Then generate the next set with the corrected brief in the AI image generator and score it again. Two rounds are usually enough to make consistency a measurement rather than a matter of opinion.
Perguntas frequentes
What should a consistency score measure?+
Four things: palette, light, treatment and framing. Those carry brand recognition and can be judged in a finished image. Mood and style words are too subjective to score usefully.
What scale should I use?+
Zero to two per criterion. Two means it matches the reference exactly, one means it is acceptable with a noted difference, zero means it breaks the set. Three or more levels invites arguments about degrees that do not change the decision.
How do I handle a set where most images score well but one does not?+
Do not average. One zero breaks a set, because the eye finds the misfit immediately. Either fix the image or remove it; a high average with one outlier still reads as inconsistent.
Should the same person always score?+
Use at least two reviewers for the first sets, then one consistent reviewer afterwards. Multiple reviewers align your rubric; a single consistent reviewer keeps it comparable over time.
Can I use a score to compare across months?+
Yes, if the reference and the rubric stay fixed. Track which criterion fails most often and fix it in the brief rather than in individual images.
Crie com kublaro
Descreva qualquer coisa e gere imagens incríveis em segundos; depois dê movimento com os melhores modelos de vídeo com IA.