The AI design crit: how to run reviews like elite creative teams
Playbook
Satej Sirur
|
Co-founder & CEO
A designer armed with an AI subscription can churn out 200 ad versions in an afternoon. Deciding which 3 are worth shipping is now a bigger headache than ever. AI shifted the bottleneck from creation to judgement.
Agencies solved this problem decades ago, and they did it socially. Work goes up on a wall. A Creative Director, an Art Director, a copywriter, and someone who has never seen the work before all stand in front of it and say what is working and what is not. The ritual has a name - the crit. Its rules are unwritten but strict - attack the work but never the person, and you say why.
Mad Men-era shops ran this as a standing appointment rather than a favour you asked for. Modern creative teams inherited the approval process but rarely the crit ritual. We believe AI can bring back this overlooked idea from the past and supercharge it to give Creative and Marketing teams a secret weapon to fight AI slop.
Key takeaways
|
|---|
Five types of reviews that burn your time and grey cells
Reviews fail in patterns most creatives will recognise. Acting on such reviews will waste precious time that you cannot spare. Identify such reviews and make sure you get a second round of review.
Name | What it sounds like | What it costs |
|---|---|---|
The Vibe Check | "Make it pop." "Needs more energy." | The reviewer noticed something is off but gave ambiguous feedback that is hard to act on. |
The Kerning Trap | "The logo is two pixels off the baseline." | The reviewer focuses on the minutiae before checking more fundamental areas like adherence to the brief. |
The Ambush | "Sorry, just catching up. Why are we going in this direction?" a week after assets were submitted for review | Redoing all the work because the required people were not involved from the start. |
The Ghost Designer | "Move the headline up, make it 40pt, swap that photo, …" | The reviewer redesigns in the comments instead of naming the problem. |
The Rubber Stamp | "Looks good to me" sent in 90 seconds | Cursory checks lead to assets that are rejected by platforms or, worse, lead to poor campaign outcomes. |
What a good crit actually checks
Regardless of the industry, channel, or campaign objective, we have found that the best crits check these 7 areas.
Brief adherence
Compare the asset against the brief it came from: objective, audience, funnel stage, channel, placement, primary message, mandatory elements, call to action. If any of those are missing or wrong, stop and rework the asset. No other checks are needed if this fails.
Brand
Logo usage, colour, typography, claims, legal language, accessibility. Failure here should also lead to immediate rework.
Squint test
Shrink the asset to 25%, blur it, and check the focal point, a readable eye path, visible product, and a headline that draws the eye before anyone reads a word.
Craft
Layout, alignment, grid, spacing rhythm, negative space, balance, cropping, type hierarchy, colour hierarchy, lighting, motion.
Taste
Originality, contemporary feel, emotional impact, premium quality, distinctiveness, whether the work leads the category or follows it. This is subjective but it is perhaps the most important in today's age of volume and churn.
AI fingerprints
Patterns that make work feel AI-generated rather than human-crafted. Examples are everything centred, spacing that is uniform instead of rhythmic, products floating with no ground shadow, default typography, lifestyle imagery that could belong to any brand.
Quality
Resolution, compression artifacts, cutout edges, shadows, colour profile, gradient banding, text rendering, safe zones, accessibility, platform specifications.
How to give feedback a designer can act on
The first thing that the crit should share is a verdict: Ship, Tweak, Redo. Mincing words here will waste time.
Next, for each of the 7 areas, share any misses. For each miss, share:
What the miss is in one sentence
A tag: Wrong, Weak, or Taste
How confident you are
What would fix it
Evidence to back up your claim
The tag is perhaps the most overlooked part of this. Wrong means an objective failure such as a brief miss or a broken brand rule. Weak means a craft problem measured against a standard you can point to. Taste means personal preference.
How the squint test works
In our experience, the squint test is a powerful way to predict how an asset will perform in the real world. Take this example banner and its resized and blurred version next to it. Blur removes your ability to read so you can focus on the layout and where the attention is drawn.
Banner | Resized and blurred version |
|---|---|
![]() | ![]() |
Here is a prompt that you can use in ChatGPT or Claude to run the squint test on one of your assets.
You are a seasoned Creative Director running a squint test on a single asset. Resize the attached asset to 25% of its original dimensions and apply a heavy blur so that no text is legible. Evaluate these in resized and blurred image. 1. Focal point - What wins the eye first, and is there only one? For each, return Pass, Weak, or Fail, with one sentence of reasoning based on what you can see in the blurred version only. Do not use knowledge of the unblurred asset. If you cannot tell, say you cannot tell. |
Here is what ChatGPT had to say about this image.
Criterion | Rating | Reasoning |
|---|---|---|
Focal point | Pass | The eye immediately lands on the large, bright rectangular object in the lower half of the frame. There is a clear dominant focal point with no obvious competitor. |
Eye path | Weak | My eyes move from the bright block near the top, down to the phone shape, then stall. There is little visual guidance connecting the upper area to the product beyond simple vertical alignment. |
Product visibility | Pass | Although details are lost, it is still recognizable as a handheld smartphone-like device with a distinct camera module. |
Message clarity | Fail | I cannot infer the proposition from the blurred composition alone. It is evident that this is an advertisement for a product, but I cannot determine the offer, benefit, or message without readable text. |
Cognitive load | Pass | I perceive roughly 4 to 5 competing visual elements: the headline block, the secondary text block, the small central icon, the product, and the hand. That is a manageable number, with one clearly dominant object. |
The AI fingerprint test on the same asset
This is a failure mode nobody had to think about before 2023. The arms race of LLMs has given us an abundance of AI models that can generate images and videos blazing fast. However, most of them scream "AI slop" and they are destructive for both short-term conversion and long-term brand equity. This is one of the most important tests you need to run in your crit.
Ironically, AI is a very good judge of other AI. Here is a prompt you can use to detect AI fingerprints.
You are a seasoned Creative Director checking a single asset for AI fingerprints: patterns that make work feel AI-generated rather than designed by an expert human designer. You are not judging whether AI made this. That question is unanswerable and unhelpful. Human-made work fails these checks too. Judge only whether the composition shows evidence of deliberate choices. Check each of the following for the attached asset: 1. Centring - Is every element on one vertical axis, or does the layout use an asymmetric grid with intent? Rules: - Judge only what is visible in the attached file. Do not infer intent from the brand. Return a table: Fingerprint | Pass, Weak, or Fail | Evidence |
This is what Claude said about the above banner.
Fingerprint | Verdict | Evidence |
|---|---|---|
Centring | Fail | All seven display elements sit within 3px of the canvas centre axis. Measured horizontal midpoints: 472.5, 469.5, 470.5, 470.0, 473.5, 471.0, 470.0 against a canvas centre of 470.5. The device is centred too. There is no off-axis decision anywhere in the frame |
Spacing rhythm | Pass | Gaps are grouped, not uniform. Within-group spacing runs 19 to 24px (headline to subhead 24px, product name to tagline 21px), between-group spacing runs 42 to 43px either side of the divider rule at y=347. Somebody set that deliberately |
Scale contrast | Weak | Overall ratio is fine at roughly 6:1, from the 85px headline band down to the 14px vertical legal line. The problem sits in the middle: the subhead, "The all-new", and "Nova X8" measure 47px, 45px and 49px, a ratio of 1.09:1 across three supposedly different hierarchy levels |
Ground contact | Pass | The device is held, with fingers wrapping the left edge and the thumb across the right. Contact is unambiguous. No floating product |
Typography | Weak | One family throughout, in at least two weights, differentiated further by colour. Not a default system stack, and not a lazy single weight either. I cannot tell from the file whether this is a licensed brand face or a generic geometric sans, so treat the weak verdict as "no pairing, one family carrying everything" rather than a font identification |
Negative space | Pass | Roughly 150px of clear space above the headline, generous margins either side of the device, and the area around the divider rule is given to nothing |
Colour distribution | Weak | There is a dominant (purple gradient, RGB 32,28,46 at top to 172,152,209 at bottom), a secondary (white type) and an accent (light violet on the subhead, "The all-new", and the glyph). The accent sits in the same hue family as the background, so it separates by luminance only. The highest-saturation element in the frame is the skin tone of the hand, which nobody chose as an accent |
Imagery specificity | Fail | A hand holding a handset against a coloured gradient is the default composition for the category. Replace the wordmark with any competitor and the image works unchanged. Nothing in the photograph is specific to this product beyond the colourway |
Trend currency | Pass | Dark gradient with a lit product hero is current. The four-point sparkle glyph belongs to a convention that peaked within the last three years rather than before it. Worth a human look regardless: models date visual conventions unreliably, and the glyph is decorative with no stated job |
Building an AI Crit for your own brand
Each of the individual checks of the crit can be built into their own prompts. What matters is grounding the prompt with context about your brand. Here are 6 things to add as context to your AI Crit.
Brand book containing colour, logo claims, legal language, accessibility rules, and more
Approved assets with annotations about their tier, channel, and audience
Industry best practices drawn from competitors, adjacent categories, and winning campaigns
Anti-references, meaning work you explicitly rejected or dislike
Past feedback pulled from where you collaborate on designs
Craft standards such as grid system, type scale, spacing scale, and photography treatment
Testing your AI Crit
Calibrate your AI Crit so you can trust it for every campaign. Here are a few tests you can run.
Run it on the last 20 assets that were approved the first time and the last 20 that came back with comments. The approved set has to score higher.
Run it on 20 competitor assets. It should flag them as off-brand rather than passing them, because passing a competitor's work means it is scoring general design quality and not your brand.
Run it on 20 AI-generated assets. It should catch the fingerprints.
AI Crit out of the box with AI Studio
AI Crit is one of the many tools built into AI Studio that allows us to ship tens of thousands of on-brand assets on time for global brands. We fastidiously measure the first-pass rate, which tells us how many assets were approved by customers without any reviews. Tools like AI Crit help us keep our first-pass rate to above 96%.
Where human expertise comes in
No AI Crit can replace your and your team's taste. A model can tell you an asset looks like every other asset in the category, and it can tell you which of your own approved work it resembles. Whether that sameness is a problem for this campaign, in this quarter, against this competitor, is a judgement that belongs to a person with a name, a job title, and accountability.
The 7 checks buy you the right to spend your judgement on that question instead of on whether the logo is in the safe zone. Brand guardians do not need a machine with opinions. They need everything below the opinion handled, at volume, before the work reaches them.
If your reviews keep uncovering problems that were really brief problems, here is how to craft a brief that works.
Frequently asked questions
What is a design crit?
A design crit is a structured review where designers, copywriters, and Creative Directors evaluate work in progress before it goes out. It differs from an approval in that its purpose is to improve the work rather than to authorise it. The rules are that feedback attacks the work rather than the person, and that every criticism comes with a reason.
Who should give feedback on creative work?
Whoever owns the business outcome makes the final call, usually the brand or channel manager. Everyone else contributes within their area: a creative director on craft and taste, a legal or compliance reviewer on claims, an e-commerce manager on platform specification. The common failure is treating all opinions as equally weighted, which turns a review into a negotiation. Settle who decides what before the first round, not during the third.
How do you review AI-generated designs differently?
The 7 checks stay the same, with one addition. AI-generated assets fail in a way human work rarely does: they pass every objective test and still read as machine-made. Centred layouts, uniform spacing, weak scale contrast, floating products, and interchangeable lifestyle imagery are the usual signals.
What is the squint test?
The squint test means resizing an asset to about 25% and blurring it until no text is legible, then judging what survives. It shows you the focal point, the eye path, product visibility, and message clarity as a scrolling shopper would experience them. It costs little, takes seconds, and catches hierarchy problems that a full-size review hides because the reviewer can read.
How many revision rounds are normal, and what causes them?
Two rounds is a healthy target for a standard adaptation, three for new creative. Most teams run more, and the extra rounds usually trace back to one of two causes: an incomplete brief, or a review that started at craft and only reached brief adherence much later. Both are process problems.
Can AI review creative work?
AI can reliably check the objective layers such as brief coverage, brand rules, platform specifications, asset resolutions, safe zones, and accessibility. It is reasonable at craft when it has been trained on the brand's own approved and rejected work. It is unreliable at taste, and any tool that claims otherwise is scoring general design convention rather than your brand. Treat the output as a first pass that clears the mechanical failures before a person spends attention on the interesting ones.
How is a crit different from an approval?
An approval answers one question: can this ship. A crit answers a longer one: what is wrong with this, how badly, and what would fix it. Teams that only run approvals accumulate the same errors across quarters, because nothing in an approval creates a written standard. A crit produces one every time somebody has to explain why a preference is a preference.


