different species of crabvietnamese mud crab
2
1 Comment

How Do You Evaluate a Designer When AI Can Generate a Good Interface?

For years, evaluating a product designer seemed relatively straightforward.

You could look at their portfolio, assess the quality of their screens, ask about their process, and discuss how they worked with product managers and developers.

The visual output was not the only thing that mattered, but it was usually the most visible proof of competence. Then AI changed the equation.

Today, a person with limited design experience can describe a product, generate a polished dashboard, add a design system, create responsive variants, and even produce a working prototype in a surprisingly short amount of time.

The result may look professional. It may follow common UX patterns. It may be more visually consistent than the work of many junior designers from only a few years ago.

So what exactly should we evaluate now?

If a good-looking interface is becoming easier to generate, the ability to produce one can no longer be the primary measure of design quality.

A polished interface is no longer strong evidence

This does not mean visual design has stopped mattering.

A badly designed interface is still a problem. Hierarchy, typography, spacing, accessibility and consistency still require attention.

But a polished screen tells us less than it used to.

It does not tell us whether the designer understood the problem.

It does not tell us whether the product serves the right user.

It does not tell us whether the proposed flow works in real situations.

It does not tell us what happens when data is missing, the AI is wrong, the user changes their mind or the system fails.

And it definitely does not tell us whether the interface should exist in the first place.

AI can generate a convincing answer before anyone has properly defined the question.

That is precisely why evaluating designers through the final screen alone is becoming increasingly dangerous.

Evaluate the decisions behind the interface

The most important design work is moving away from drawing individual screens and towards defining product behaviour.

When reviewing a designer’s work, I would pay less attention to how quickly they arrived at a polished result and more attention to the decisions they made along the way.

Why does this feature exist?

What user problem does it solve?

What assumptions are being made?

Which parts of the proposed solution were generated by AI, and which parts were consciously accepted, changed or rejected?

What alternatives were considered?

What evidence influenced the final decision?

A strong designer should be able to explain not only what they created, but why the product behaves in this particular way.

AI can suggest patterns.

The designer is still responsible for deciding whether those patterns make sense in the specific context.

Look for the ability to recognise false confidence

AI-generated interfaces often look more complete than they really are.

A prototype may contain charts without meaningful data.

A form may appear simple because difficult cases were ignored.

A generated workflow may work perfectly for the happy path but collapse as soon as the user makes a mistake.

The interface can create the impression that the product logic has already been solved.

Often, it has not.

One of the most valuable design skills will therefore be the ability to recognise when something only looks convincing.

Can the designer identify missing states?

Do they question generated assumptions?

Do they notice when a flow has no clear recovery mechanism?

Do they understand where the data comes from and whether the system is technically capable of producing the promised result?

Do they know the difference between a prototype, a demo and a production-ready product?

The strongest designer may not be the person who generates the most impressive first version.

It may be the person who finds everything that is dangerously incomplete about it.

Evaluate problem framing, not only problem solving

AI is very good at producing solutions.

It is much less reliable at deciding which problem deserves to be solved.

This makes problem framing a much more important part of evaluating designers.

Give two designers the same vague request:

We need an AI assistant for freelancers.

One may immediately generate a dashboard with a chat panel, project cards, notifications and productivity analytics.

Another may begin by asking:

What kind of freelancers?

What is currently taking too much time?

Is the real problem planning, communication, prioritisation, invoicing or finding work?

What information can the system access?

How much control does the user expect?

What happens when the assistant makes the wrong recommendation?

The second designer may initially produce less visible output.

But they are reducing the risk of building the wrong product.

In an AI-assisted environment, generating more screens is easy. Narrowing the problem intelligently is hard.

Test whether they can design systems, not isolated screens

AI products rarely behave like traditional, predictable software.

The same input may produce different outputs.

The system may be uncertain.

Recommendations may change when new information appears.

The interface may need to adapt to context, user experience, permissions and confidence levels.

That means designers need to think in rules, states and boundaries.

Instead of asking only:

“What should this screen look like?”

They need to ask:

What is the system allowed to do?

When should it ask for confirmation?

How should uncertainty be communicated?

What can the user undo?

Which decisions require human approval?

How does the interface change for a new user, an expert user or a user with limited attention?

How does the product behave when an agent fails?

Evaluating a designer should increasingly involve examining whether they can define these behaviours.

A static portfolio may not reveal that.

A conversation about edge cases often will.

Judge how they use AI, not whether they use it

Using AI should not automatically count as either a strength or a weakness.

The question is how the designer uses it.

Do they use AI to explore more alternatives?

Do they use it to accelerate repetitive production work?

Can they compare outputs critically?

Can they maintain consistency across generated components?

Do they understand when to leave the generated solution behind and design something specifically for the product?

Do they verify the output rather than treating it as a finished answer?

A designer who refuses to use AI may become unnecessarily slow.

A designer who accepts everything AI produces may become unnecessarily dangerous.

The valuable skill is orchestration: knowing what to generate, what to verify, what to change and what to reject.

Review the questions they ask

One practical way to evaluate designers is to pay close attention to their questions.

Weak questions tend to focus on appearance:

Should the button be blue?

Can we make this cleaner?

Should we use cards or a table?

Strong questions expose product risk:

What happens if the user does not trust the recommendation?

Which decision is reversible?

How will we know whether this feature is useful?

What information is the system missing?

Who is responsible when the AI takes the wrong action?

What is the simplest version that tests our main assumption?

The quality of a designer’s questions often reveals more than the quality of their first solution.

Change the design exercise

Traditional design tasks often ask candidates to redesign a screen or create a new feature.

That format becomes less useful when a candidate can generate multiple polished concepts almost immediately.

A better exercise may begin with an imperfect AI-generated product.

Ask the designer to critique it.

Ask them to identify unsupported assumptions, missing states, ethical risks, accessibility issues and technical uncertainties.

Then ask them to improve the product logic, not only the visual layer.

You can also change the input halfway through:

The data is incomplete.

The user rejected the recommendation.

The system has only 60% confidence.

A different user has different permissions.

The action cannot be undone.

Now what happens?

This tests whether the designer can respond to the complexity that polished mock-ups often hide.

The portfolio may need to change too

Design portfolios are traditionally organised around final screens.

The future portfolio may need to show much more of the reasoning underneath them.

What was generated?

What was manually designed?

What assumptions were tested?

What changed after user feedback?

Which AI suggestions were rejected?

How were product rules defined?

What were the failure scenarios?

How did the designer work with engineering, data and business constraints?

The interface should still be visible.

But it should no longer be presented as the main evidence of quality.

It is the result of the designer’s decisions, not the decisions themselves.

The new standard

When almost anyone can generate something that looks like a well-designed product, the role of the designer does not disappear.

But the standard changes.

A good designer is no longer simply someone who can produce a good interface.

A good designer is someone who can recognise whether the interface solves a real problem, define how the product should behave, challenge generated assumptions, understand system constraints and protect the user from convincing but incorrect solutions.

AI is making interface production cheaper.

It may make design judgement more valuable.

The question for hiring managers is therefore no longer:

Can this person create a polished interface?

It is:

Can this person make good product decisions when producing a polished interface is the easiest part of the job?

How would you evaluate a designer today? And what part of the traditional portfolio or recruitment process would you remove first?

on July 27, 2026
  1. 1

    Coming from a Scrum and Kanban background rather than design, I see a very similar shift in how we should evaluate people across product teams. When producing a convincing output becomes easier, the real value moves towards the reasoning behind it: how someone frames the problem, questions assumptions, notices risks and responds when reality does not match the happy path.

    I especially like the idea of giving a candidate an imperfect AI-generated product instead of asking them to create another polished concept from scratch. Seeing what they challenge, what they investigate and which questions they ask would probably reveal much more than the final interface. AI can make almost everyone look productive for a moment, but it cannot hide weak judgement for very long.