AI can generate the options. It can’t tell you which one is wrong.

Anthropic launched Claude Design in mid-2026 with an explicit pitch: a tool for founders, PMs, and marketers — people who, as the product page puts it, don't live in Figma. It generates prototypes, decks, and marketing assets fast enough to be genuinely useful, and the output is polished enough to share.

We've had this land in our own projects already. PMs on two of our current engagements used it over a weekend to mock up feature ideas and came back on Monday with something that looked ready to present. First time it happened, it was a slightly strange moment.

But here's what we've actually observed across all of those situations.

The gap between plausible and right

AI-generated design output looks plausible on the surface. That's genuinely impressive, and it's a real shift from where these tools were even twelve months ago. The problem is that "plausible" and "right" are very different things in product design, and the gap between them isn't visible until you know what to look for.

The moment you run an AI-generated design through a real process — checking it against actual user flows, existing design system components, edge cases, mobile breakpoints, technical constraints — it tends to fall apart in specific, predictable ways. The AI optimised for what looked correct based on common patterns. It didn't have the context to know that your users are mostly working on small screens, that your backend can't support that interaction model, or that three versions of that component already exist and are maintained in your design system.

The tool generated five options. None of them were wrong in an obvious way. But one of them was right for your specific situation, and determining which one requires knowledge the AI doesn't have.

What this actually changes for design teams

The role that survives this shift isn't the one that produces design output. It's the one that evaluates it.

We've started describing our own work on some projects as more editorial than generative — less time producing initial concepts from scratch, more time making informed judgments about what AI produces, knowing which direction to push it in, and being the person who can look at what came out and say: this is good enough, or this isn't, and here's specifically why.

That shift requires exactly the kind of experience that's hard to shortcut: having shipped enough products to know how things fail, having worked with enough user bases to spot assumptions that won't hold, having run enough design reviews to recognise when something looks finished but isn't.

What this means for companies hiring design help

The question we're hearing more often from clients isn't "do we still need designers?" — it's "how do we know if the design we're getting is actually any good?"

AI can produce a convincing answer to a design brief. What it can't do is tell you whether it's the right answer for your specific product, your specific users, and your specific moment in the market. That evaluation is what senior design experience is actually for, and it becomes more valuable as the volume of plausible-looking output increases.

In short: more output, same need for judgment. If anything, the ability to assess quality matters more when there's more to assess.

Gytis Markevicius
July 26, 2026
5 min read

If you're building complex AI apps and the design isn't where it should be, a 20-minute conversation is a good place to start.

Every engagement includes Head of Design oversight —
a senior designer thinking about the product strategically, catching issues before they reach you, and making sure the work holds up under scrutiny.

A team that's quick to start, communicates clearly, and doesn't need managing.
No long hiring process, no onboarding overhead, no hand-holding.
Focus and dedication because we only work with a small number of clients at once.

Let's talk about your product and your business goals.

Request your FREE trial

✦ Start with a conversation ✦

Book a call to start your FREE 3-day trial

20 minutes. No pitch. We'll tell you honestly if we're the right fit.