Generate, Evaluate, Refine: An AI-Enabled Workflow for UX Design

AI & Product Design

By Thomas DiNataleJune 11, 2026

There’s no shortage of AI workflow advice for designers these days, and most of it stays abstract. Just open LinkedIn. What follows is a workflow that emerged from a real project. Three moves that made it possible to go from a blank Figma screen to wireframes ready for concept testing in a single afternoon. The workflow is generate, evaluate, refine. Here’s how it actually worked.


We’ve spent several years designing software that helps teachers run their classrooms. When the navigation started getting in the way, it was time to rethink it. The navigation, a side panel with a series of open and close menus, was limiting. It couldn’t accommodate new features gracefully, and the experience had grown more complex than it needed to be. We knew we needed to start the next school year with a redesigned navigation that could handle growth and feel simpler at the same time.

We also knew going into the project we wanted a place to surface timely information to teachers, but it was unclear where that would live or what it would be. The navigation work had to come first, because understanding how teachers move through the product was a prerequisite for knowing what a time-based user experience should surface.

Our timeline was tight and the scope of the project required a lot of deep thinking to quickly deliver wireframes ready for concept testing. What we didn’t know going in was that the details of those wireframes would compress into a single day, and that AI-enabled workflows would play a central role in making that possible.

The research behind the sprint

To kick the project off, we ran a moderated open card sorting exercise with academic staff and product team members. The goal was to understand how teachers think about their daily work so we could design a more intuitive navigation.

Kicking the project off with a 15 minute open card sort exercise.

The results of the exercise, card groupings and category labels, were very similar across team members. This provided us with a clearer picture of how teachers organize their day: what they reach for first, what they return to repeatedly, and what they rarely need at all. From those patterns, a new navigation architecture emerged. One organized around how teachers spend their time, not around how the software happened to be structured.

The card sorting phase matters to the rest of the story for one reason: we already knew this product well. Years of working on its features meant we understood how teachers used it, and where it was failing them, before the card sort began. What the exercise produced was a clearer understanding of how teachers think about their work.

Surfacing the right information at the right time

The first design review covered global navigation. We designed a sectioned, icon-anchored structure that reduces visible options at any one time. The time-based user experience had no detail yet, just an open question about what information would genuinely serve teachers throughout their day.

Dashboard concept navigation diagram.

During that review we proposed two time-based concepts to test against. The first was a dashboard that contained information relevant to all teachers regardless of grade level or subject. The second was a shortcut links feature that would let teachers save direct links to the screens they use most. This second concept was a deliberate risk-reduction move. Rather than trying to anticipate every possible widget need before the school year began, we gave teachers a way to surface what mattered to them directly. If we got a widget wrong, the shortcuts would cover the gap.

By the end of the review we had alignment on both concepts and a starting point for the dashboard widgets: a dozen that were clearly required for all teachers. The list covered the moments that shape every teacher’s day: current class in session, daily attendance, unread inbox messages, daily student goal completion, daily grading, parent communication, upcoming events, and academic reporting. Plus widgets that are only needed at specific times during the year, like re-enrollment and academic progress testing.

That list left the room with us on a Thursday morning. The concept testing wireframes needed to be ready the next day.

A dashboard isn’t a screen. It’s a snapshot in time.

Good enough for concept testing means something specific. The content had to be accurate and grounded in the teachers’ actual experience, real widget names, real data states, real moments in the school day.

Every widget has to tell a true story at 9 AM, at noon, and at 3 PM. The attendance widget showing “Submitted at 9:02 AM” only works because it’s after 9. Before 9 it needs to be an alert. The current block widget showing “18 minutes left” is only accurate during an active class. The daily grading widget showing one task complete and one still pending is specific to that moment in that day, after the morning submission, before the 3:30 PM deadline.

Getting that logic right across fifteen widgets, each with their own states, their own rhythms, their own relationship to the clock and the calendar, is the work that concept testing will either validate or surface problems with. And it’s invisible when done well, which means it rarely gets acknowledged as the hard problem it actually is. This is what an afternoon produced. Here’s how it happened.

Final wireframe for concept testing from the generate, evaluate, refine workflow.

Generate, evaluate, refine: A UX workflow

With a dozen widgets to design and an afternoon to do it, the process couldn’t afford to be slow. What emerged was a way of working with AI, in this case Claude, that had three distinct moves.

The first was to generate. Use Claude as a design partner to rapidly explore the content and structure of each widget before committing anything to a Figma file. Let it be unconstrained. That’s the point. I made the explicit decision to not overcontextualize at the beginning. 

The second move was to evaluate. Apply product and design knowledge to discard ideas that don’t belong and identify the ones worth pursuing. This is where design judgement does its most important work. Not in idea generation, but in knowing which ideas fit this product, this user experience and which ones don’t. 

The third was to refine. Take what survived the evaluate phase, design the UI, and bring screenshots back to Claude with specific questions. This is the move that is easy to skip, and the one that turned out to be surprisingly valuable. 

Each move has a distinct role. Here’s how they worked in practice.

Generate

Designer is the director, AI is the producer

The first move was to open a conversation with Claude and start exploring each widget without constraints. No product context loaded upfront, no existing UI patterns referenced or Figma MCP connected to constrain the output, no boundaries set on what it could suggest. The goal was volume and variety; as many interpretations of each widget’s content and structure as possible before making any design decisions.

For each widget I asked specific questions that gave Claude a starting point. I started with the current class widget and asked: Provide options for a current class widget that communicate the time remaining in the class, the time elapsed, when the class ends and use color to provide quick visual reference. What are the different ways you can lay out content for this widget?

Claude generated current class widgets.

The deliberate decision not to overcontextualize at the start mattered here. When AI knows too much about your constraints upfront, it self-edits. It produces safer, narrower output within the lines you’ve drawn. Deciding to leave those constraints out meant the generate phase returned a wider range of options. Most of these options would be set aside, but they surfaced quickly. A paper and pencil sketch session that would have been an hour or more was done in 5 minutes.

The generate phase is intentionally high-volume, low-stakes thinking. Like a traditional design charette, it generates a lot of options quickly, most of which you’ll discard. The constraint now isn’t the tool’s ability to generate; it’s the designer’s ability to evaluate what comes back. That’s the next move. 

Evaluate

Designer only, AI is absent

The generate phase returns options. The evaluate phase is where the designer decides which ones to put aside and which ones are worth pursuing in the refine phase.

The evaluate phase requires the kind of design knowledge that is hard to feed into an AI. Yes, you could load constraints within Claude upfront to create options more tailored to the product, but that misses the point. The knowledge required to evaluate well doesn’t live in a prompt. It lives in the designer’s understanding of the product’s interaction philosophy, the way experience principles come alive across a thoughtful user experience, the UX systems and design principles that have been established over time. These aren’t design system components or UI pattern rules. They’re the deeper logic of how a product works, and they’re the hardest thing to transfer to an AI tool.

Two examples from the dashboard project  illustrate how evaluation works and why it requires a designer who knows the product well enough to catch both.

Adding what doesn’t belong

The first type of misalignment is the easiest to spot: AI adds something that doesn’t belong. If you’ve ever used AI to work through a problem, you’ve experienced this in some form. In our example, Claude suggested a streak counter for the attendance widget. Streaks work. Habit tracking works and maybe showing an attendance streak would compel teachers to complete attendance by 9AM every day.

Claude generated attendance widget with streak concept.

But this product doesn’t use that pattern anywhere else, and introducing it in a single widget creates inconsistency without a product-wide decision to support it. The streak was filtered out. Not because it was a bad idea, but because it didn’t belong in this product.

Replacing what already exists

The second type of misalignment is harder to catch: AI introduces something that looks reasonable but violates an established design principle. For our sprint, we needed to design a widget that allows the teacher to view analytics about posts they make to a class feed for parents. Simple, right? Well, Claude designed a full posting interface UI from the widget. It looked polished and it was a reasonable UI pattern. 

Claude generated parent communication widget.

The problem was that it contradicted the dashboard’s core design principle: this is a surface for information retrieval, it’s a summary layer, not a control center for taking action. That’s what makes this type of misalignment challenging to catch. The widget doesn’t look wrong. The UI is fine. Instead, it’s the interaction model that is fundamentally at odds with the design principles of the dashboard.

Wrong additions feel out of place. Wrong patterns feel at home until you check them against the design principles that define what you are designing. Both require a designer. Now that we’ve generated and evaluated our widgets, it’s time to refine.

Refine

Human & AI in dialogue, working on real designs

The third move was to take what survived the evaluate phase, design the UI, and bring screenshots back to Claude with specific questions. It’s the move that’s easiest to skip because the UI Claude generates looks complete. We found this step to be the most valuable. The current class widget is the clearest example of why. 

In the generate phase, Claude produced three layout options for communicating time in a class, a horizontal progress bar, segmented blocks and an arc with inline metadata. The evaluate phase selected option A, the horizontal progress bar as the right foundation. Options B and C were set aside. 

But a progress bar alone only tells a teacher how much time is left. The refine phase is where the designer adds what AI doesn’t know. 

In our dashboard concept we decided the current class would be literacy class. This gave us additional context for the design: the book the class is reading, the pages assigned for today, the current lesson, and a direct link to the teaching notes. None of that came from Claude. It came from knowing how a teacher’s day actually works, what they need at 12:02 PM when one class ends and another is just beginning, leaving little time to navigate across the product. 

Dashboard wireframe with the current class in focus.
Current class wireframe with option A from the generate phase and designer context.

That richer design went back to Claude as a screenshot with a simple, specific question: Here is a design of the current class module with additional details, how can I refine this experience? 

The conversation shifted from UI to UX. 

Instead of more layout options, the dialogue reflected on what the design refinement was doing for the teacher: surfacing the right context at the right moment so nothing had to be searched for when classes changed. That framing narrowed the design problem productively and opened angles that hadn’t been considered yet. One of them: five minutes before the class ends the teacher is already thinking about what comes next. Should the widget surface a preview of the upcoming class before the transition happens? Or, can a teacher customize the widget to their own teaching preferences? 

Those are UX questions, not UI questions. They came from a back-and-forth about a specific design that required the designer’s knowledge. The refine move doesn’t just iterate on AI output. It adds what AI can’t supply, then uses dialogue to explore and sharpen the experience further. That sequence is what shifts the conversation from what the UI looks like to what the experience should provide.

What made it work and what the process requires

One thing worth noting before drawing any conclusions. The sequence, generate, evaluate, refine, isn’t always linear. The examples in this article intentionally move cleanly from one phase to the next but in practice the loop is more fluid. You might be deep in the refine phase, bring a screenshot back to Claude, and find that the right response is to generate new layout options with the additional context. Claude’s output then goes back through an evaluation before design refinement. The three moves are better understood as modes, than steps, and knowing which mode the work needs at any given moment is itself a design judgment call. 

What makes this workflow possible wasn’t the AI. It was the foundation that existed before the AI phase began. Years of working on this product meant we understood the interaction philosophy, the UX systems, the design principles that had evolved over time. That knowledge is what made the evaluate phase work. Without it, there’s nothing to filter against. The generate phase returns options and they all look equally plausible. 

AI amplifies what the designer brings to it. Feed it product knowledge and design judgment, and it accelerates output. Without the designer’s knowledge and judgment, the output might look considered but it fits nowhere in particular. Just some UI.

The honest version of the project’s timeline and output: a dozen designed widgets, a single afternoon, high-fidelity wireframes ready for concept testing. Those are real. But it only worked because the designer knew which ideas to keep, which to set aside, and when to push back with refined designs and specific questions. 

The pace of design has changed. The job hasn’t.

Read more

AI & Product Design

AI and Design Thinking

AI is transforming how we live, work, and communicate. With AI’s help, we can move from idea to product launch at super-speed. Never before have we been able to move

Allison SallSeptember 26, 2025