Description: Once you've picked your evaluation variant, run it well: scope it, evaluate with discipline, score and prioritize, and land on a fix list your team will act on.
For: Practitioners (product managers, designers, researchers) and Proxy Functions (sales, customer support, analytics, etc.) who will run an evaluation.
Time to apply: Same day to 5+ days, depending on the variant you picked
Pulls from these UX Sourcebooks: Advanced Usability Checklist, Broad User Research Insights
Last updated: Sep 30, 2026
You've decided to run a heuristic evaluation and picked your variant. Now the quality of what you get out of it depends on scoping it tightly, evaluating with discipline, and finishing with a prioritized list, not just a pile of findings.
Get leadership agreement on timeline and deliverable, and assemble your evaluators: internal product trio, proxy functions (sales, customer support, data analysts, and other internal subject-matter experts), and/or external participants, depending on your variant.
If proxy functions are involved, budget 20 minutes to walk them through the template and how to evaluate with it. They know customers, not heuristics, so this briefing is what makes their findings usable. Then set scope below.
One flow: Evaluate now.
Many flows: Split. Pick the top 1 to 2 flows now, and schedule the rest.
One flow: Assign for an individual to do a quick pass (30 to 60 minutes), or defer to later.
Many flows: Defer. Revisit after the critical flows are fixed.
A note on flow length: "one flow" can still be too much if it has many steps or touch-points, a 12-step checkout flow isn't the same job as a 3-step login. Treat a long single flow the same way you'd treat many flows: evaluate the highest-risk 2 to 3 steps today, and schedule the rest.
For a human-led evaluation:
Download the template.
Read the instructions and get familiar with the core heuristics.
Determine if the evaluation will focus on core heuristics only, or if it will include add-ons (content, AI performance), and/or focus on accessibility and inclusivity.
Time: 30 minutes
Output: A shared checklist of page states in the workflow. One tab per a page state in the spreadsheet.
For an AI-assisted evaluation:
Download the AI evaluation package.
Upload the files to a new Project in Claude and rename it "Heuristic Evaluation".
Time: 30 minutes
Output: A folder of all page states in workflow(s) to evaluate, labeled logically.
For a human-led evaluation:
Assign individuals or pair of people to evaluate each page state for all heuristics.
Time: 1 to several hours per flow, with 1 or more scheduled sessions pending scope.
Output: Applicable heuristic issue(s) documented per page state per core heuristic.
For an AI-assisted evaluation:
Drop all images for each workflow separately into Heuristic Evaluation Claude Skill to generate documented issues in a spreadsheet. Upload each workflow separately or all workflows together (your preference). Human review of AI output is strongly encouraged.
Time: 30-60 minutes
Output: A single spreadsheet or set of spreadsheets of documented issues ranked by severity.
Whichever variant you're running, evaluate each screen in the taskflow in the sequence that real user actually encounters them, rather than jumping around.
Issues often compound across steps (a confusing label on screen 2 makes an ambiguous choice on screen 4 worse), and that compounding is easy to miss if you're not walking it in order or only conducting an AI-assisted evaluation.
For a human-led evaluation:
Rate each issue by severity and reach (if available).
If multiple taskflows and/or tabs in spreadsheet are used, they will need to be collated.
Merge duplicate issues documented (if applicable).
Tag the source of the issue if combining human-led and AI-assisted (if applicable).
For an AI-assisted evaluation:
Documented issues will already have severity scores. Review these, and include a reach score (if available).
If multiple taskflows and/or tabs in spreadsheet are used, they will need to be collated.
Merge duplicate issues documented (if applicable).
Tag the source of the issue if combining human-led and AI-assisted (if applicable).
How do you rate issues by severity and reach?
In the spreadsheet, for each issue documented with the heuristic violated labeled, use the following to rate severity and reach:.
Severity scale
Severity is how painful the issue is to customers who experience it.
Minor: Causes hesitation or mild irritation. User recovers without help.
Moderate: Causes occasional task failure or real delay and frustration for some users.
Critical: Leads to task failure. Severe frustration or blocks the user outright.
Reach scale
Reach is the projected or factual number of users who run into the issue.
Low: Hit by 1 evaluator, or an isolated, one-off report from proxies or customers (roughly under 20%, when a real percentage is available).
Medium: Hit by around half the evaluators or test participants, or proxies describe it as a recurring-but-not-dominant pattern (roughly 20 to 50%).
High: Hit by most or all evaluators or participants, or proxies describe it as common and frequently reported (roughly 50%+).
If you need to a qualitative count of the number of users who reported the issue in support, sales, research, or otherwise, try the Usability Problem Matrix.
For both evaluation types:
Sort issues by lowest effort (perceived or actual) and impact (estimated highest) with the team that owns the solutions where the issues exist.
Pair these issues with supportive evidence from research, analytics, sales/support volume, etc.
Discuss and align on issues to prioritize and timeline to remediate.
Moderate to Critical (task failure or real frustration)
Low reach (few users): Fix if cheap. Exception: accessibility blockers always get fixed.
High reach (many users): Prioritizing fixing these first.
Minor (hesitation, mild irritation)
Low reach (few users): Log it. Don't spend time here.
High reach (many users): Batch. Bundle into a quick-win cleanup.
High Impact
Low effort: Quick wins. Do these now.
High effort: Big bets. Roadmap them with an owner.
Low Impact
Low effort: Fill-ins. Do when convenient.
High effort: Drop. Don't spend a sprint here.
For both evaluation types:
Package top ranked issues into a deck that can be used to create shared visibility and leadership alignment. Include a view of priorities, timeline to remediate, and expected impact.
Trusting AI output unreviewed. An AI-assisted pass is a first draft, not a finding. Route it through a human review before it reaches the topline.
Evaluating alone. One evaluator misses a lot, even with AI help. Use 3 to 5 people with different backgrounds when you have the time.
Comparing notes too early in a product trio or proxy round. Independent passes are what make the merged list credible.
Stopping at the list. A findings doc nobody prioritizes is just a critique. Always finish with the prioritization step.
Reporting AI-assisted findings as if they include frequency, confirmed occurrence, or divergent use cases. A pass built on uploaded flows and screens can only flag candidate issues, not how often they happen. Say "identified as a possible issue," not "occurring for X% of users," unless the finding actually came from proxy signal, external research, or an agent with real usage data behind it.
Advanced Usability Checklist: the core heuristics plus content, AI, and performance add-ons. This is your Step 1 template.
Broad User Research Insights: use these as evaluator prompts, AI or human.
If you see… Check it against…
Long paragraphs, buried labels
Users scan, not read. Do headings carry the key word first?
Big lists, many equal options
Hick's Law / choice overload. How many options at this decision point?
Small or distant primary buttons
Fitts's Law. Is the primary action big and close?
Silent waits
Response-time thresholds. Is there feedback when it takes more than a moment?
Dense jargon
Plain language. Would this read clearly at a general reading level?
If you enjoyed this resource, consider sharing it with a friend.
Receive more resources and insights every month by signing up to receive my Unleash UX Newsletter on LinkedIn or through email.