Search this site
Embedded Files
Gerren Lamson
  • Home
  • About
  • Work
  • Writing
    • Articles
      • Driver Trees: Design Effective Alignment between Businesses and Customers
      • 4 Lessons from Leading a Vision Project for Indeed Hiring Platform
      • Assembling a Central System to Bridge Products
      • Redesigning Even Better File Organization (Part II)
      • 3 Things I Learned from Starting a Design Newsletter
      • Creating a Culture of Effective Design Feedback
      • Shaping Our Design Principles
      • Revamping Our Onboarding Process for Designers
      • How To Write Your Own UX Plan
      • Future-proofing UX in the Age of AI
    • UX Sourcebooks
      • UX Impact Data Claims
      • Broad User Research Insights
      • A11y & Inclusivity Checklist
      • Advanced Usability Checklist
      • UX Metrics 101
    • UX Field Guides
      • How to Decide If A Heuristic Evaluation Is the Right Method
      • How to Run an Effective Heuristic Evaluation
Gerren Lamson
  • Home
  • About
  • Work
  • Writing
    • Articles
      • Driver Trees: Design Effective Alignment between Businesses and Customers
      • 4 Lessons from Leading a Vision Project for Indeed Hiring Platform
      • Assembling a Central System to Bridge Products
      • Redesigning Even Better File Organization (Part II)
      • 3 Things I Learned from Starting a Design Newsletter
      • Creating a Culture of Effective Design Feedback
      • Shaping Our Design Principles
      • Revamping Our Onboarding Process for Designers
      • How To Write Your Own UX Plan
      • Future-proofing UX in the Age of AI
    • UX Sourcebooks
      • UX Impact Data Claims
      • Broad User Research Insights
      • A11y & Inclusivity Checklist
      • Advanced Usability Checklist
      • UX Metrics 101
    • UX Field Guides
      • How to Decide If A Heuristic Evaluation Is the Right Method
      • How to Run an Effective Heuristic Evaluation
  • More
    • Home
    • About
    • Work
    • Writing
      • Articles
        • Driver Trees: Design Effective Alignment between Businesses and Customers
        • 4 Lessons from Leading a Vision Project for Indeed Hiring Platform
        • Assembling a Central System to Bridge Products
        • Redesigning Even Better File Organization (Part II)
        • 3 Things I Learned from Starting a Design Newsletter
        • Creating a Culture of Effective Design Feedback
        • Shaping Our Design Principles
        • Revamping Our Onboarding Process for Designers
        • How To Write Your Own UX Plan
        • Future-proofing UX in the Age of AI
      • UX Sourcebooks
        • UX Impact Data Claims
        • Broad User Research Insights
        • A11y & Inclusivity Checklist
        • Advanced Usability Checklist
        • UX Metrics 101
      • UX Field Guides
        • How to Decide If A Heuristic Evaluation Is the Right Method
        • How to Run an Effective Heuristic Evaluation

← All UX Field Guides

 UX_FIELD_GUIDE_2 

How to Run an Effective Heuristic Evaluation

TL;DR: Once you've picked your evaluation variant, here's how to run it well: set up, evaluate, score, prioritize, and land on a ranked fix list your team can actually act on.

Overview

  • Description: Once you've picked your evaluation variant, run it well: scope it, evaluate with discipline, score and prioritize, and land on a fix list your team will act on.

  • For: Practitioners (product managers, designers, researchers) and Proxy Functions (sales, customer support, analytics, etc.) who will run an evaluation.

  • Time to apply: Same day to 5+ days, depending on the variant you picked

  • Pulls from these UX Sourcebooks: Advanced Usability Checklist, Broad User Research Insights

Last updated: Sep 30, 2026

1. The situation

You've decided to run a heuristic evaluation and picked your variant. Now the quality of what you get out of it depends on scoping it tightly, evaluating with discipline, and finishing with a prioritized list, not just a pile of findings.

2. Pre-briefing

Get leadership agreement on timeline and deliverable, and assemble your evaluators: internal product trio, proxy functions (sales, customer support, data analysts, and other internal subject-matter experts), and/or external participants, depending on your variant.

If proxy functions are involved, budget 20 minutes to walk them through the template and how to evaluate with it. They know customers, not heuristics, so this briefing is what makes their findings usable. Then set scope below.

3. Decision: What do you evaluate first?

High-traffic / business-critical

  • One flow: Evaluate now.

  • Many flows: Split. Pick the top 1 to 2 flows now, and schedule the rest.

Low-traffic / low-stakes

  • One flow: Assign for an individual to do a quick pass (30 to 60 minutes), or defer to later.

  • Many flows: Defer. Revisit after the critical flows are fixed.

A note on flow length: "one flow" can still be too much if it has many steps or touch-points, a 12-step checkout flow isn't the same job as a 3-step login. Treat a long single flow the same way you'd treat many flows: evaluate the highest-risk 2 to 3 steps today, and schedule the rest.

4. The play: five steps to conduct

Step 1: Setup

For a human-led evaluation:

  • Download the template.

  • Read the instructions and get familiar with the core heuristics.

  • Determine if the evaluation will focus on core heuristics only, or if it will include add-ons (content, AI performance), and/or focus on accessibility and inclusivity.

  • Time: 30 minutes

  • Output: A shared checklist of page states in the workflow. One tab per a page state in the spreadsheet.

For an AI-assisted evaluation:

  • Download the AI evaluation package.

  • Upload the files to a new Project in Claude and rename it "Heuristic Evaluation".

  • Time: 30 minutes

  • Output: A folder of all page states in workflow(s) to evaluate, labeled logically.

Step 2: Evaluate

For a human-led evaluation:

  • Assign individuals or pair of people to evaluate each page state for all heuristics.

  • Time: 1 to several hours per flow, with 1 or more scheduled sessions pending scope.

  • Output: Applicable heuristic issue(s) documented per page state per core heuristic.

For an AI-assisted evaluation:

  • Drop all images for each workflow separately into Heuristic Evaluation Claude Skill to generate documented issues in a spreadsheet. Upload each workflow separately or all workflows together (your preference). Human review of AI output is strongly encouraged.

  • Time: 30-60 minutes

  • Output: A single spreadsheet or set of spreadsheets of documented issues ranked by severity.

Whichever variant you're running, evaluate each screen in the taskflow in the sequence that real user actually encounters them, rather than jumping around.

Issues often compound across steps (a confusing label on screen 2 makes an ambiguous choice on screen 4 worse), and that compounding is easy to miss if you're not walking it in order or only conducting an AI-assisted evaluation.

Step 3: Consolidate & score

For a human-led evaluation:

  • Rate each issue by severity and reach (if available).

  • If multiple taskflows and/or tabs in spreadsheet are used, they will need to be collated. 

  • Merge duplicate issues documented (if applicable).

  • Tag the source of the issue if combining human-led and AI-assisted (if applicable).

For an AI-assisted evaluation:

  • Documented issues will already have severity scores. Review these, and include a reach score (if available).

  • If multiple taskflows and/or tabs in spreadsheet are used, they will need to be collated. 

  • Merge duplicate issues documented (if applicable).

  • Tag the source of the issue if combining human-led and AI-assisted (if applicable).

How do you rate issues by severity and reach?

In the spreadsheet, for each issue documented with the heuristic violated labeled, use the following to rate severity and reach:.

Severity scale
Severity is how painful the issue is to customers who experience it.

  • Minor: Causes hesitation or mild irritation. User recovers without help.

  • Moderate: Causes occasional task failure or real delay and frustration for some users.

  • Critical: Leads to task failure. Severe frustration or blocks the user outright.

Reach scale
Reach is the projected or factual number of users who run into the issue.

  • Low: Hit by 1 evaluator, or an isolated, one-off report from proxies or customers (roughly under 20%, when a real percentage is available).

  • Medium: Hit by around half the evaluators or test participants, or proxies describe it as a recurring-but-not-dominant pattern (roughly 20 to 50%).

  • High: Hit by most or all evaluators or participants, or proxies describe it as common and frequently reported (roughly 50%+).

If you need to a qualitative count of the number of users who reported the issue in support, sales, research, or otherwise, try the Usability Problem Matrix.

Step 4: Prioritize top issues

For both evaluation types:

  • Sort issues by lowest effort (perceived or actual) and impact (estimated highest) with the team that owns the solutions where the issues exist.

  • Pair these issues with supportive evidence from research, analytics, sales/support volume, etc.

  • Discuss and align on issues to prioritize and timeline to remediate.

Decision: How urgent is each issue?

Moderate to Critical (task failure or real frustration)

  • Low reach (few users): Fix if cheap. Exception: accessibility blockers always get fixed.

  • High reach (many users): Prioritizing fixing these first.

Minor (hesitation, mild irritation)

  • Low reach (few users): Log it. Don't spend time here.

  • High reach (many users): Batch. Bundle into a quick-win cleanup.

Decision: What do you fix first?

High Impact

  • Low effort: Quick wins. Do these now.

  • High effort: Big bets. Roadmap them with an owner.

Low Impact

  • Low effort: Fill-ins. Do when convenient.

  • High effort: Drop. Don't spend a sprint here.

Step 5: Write the topline

For both evaluation types:

  • Package top ranked issues into a deck that can be used to create shared visibility and leadership alignment. Include a view of priorities, timeline to remediate, and expected impact.

5. Watch outs

  • Trusting AI output unreviewed. An AI-assisted pass is a first draft, not a finding. Route it through a human review before it reaches the topline.

  • Evaluating alone. One evaluator misses a lot, even with AI help. Use 3 to 5 people with different backgrounds when you have the time.

  • Comparing notes too early in a product trio or proxy round. Independent passes are what make the merged list credible.

  • Stopping at the list. A findings doc nobody prioritizes is just a critique. Always finish with the prioritization step.

  • Reporting AI-assisted findings as if they include frequency, confirmed occurrence, or divergent use cases. A pass built on uploaded flows and screens can only flag candidate issues, not how often they happen. Say "identified as a possible issue," not "occurring for X% of users," unless the finding actually came from proxy signal, external research, or an agent with real usage data behind it.

6. Pull from the UX Sourcebooks

  • Advanced Usability Checklist: the core heuristics plus content, AI, and performance add-ons. This is your Step 1 template.

  • Broad User Research Insights: use these as evaluator prompts, AI or human.

If you see… Check it against…

  • Long paragraphs, buried labels

    • Users scan, not read. Do headings carry the key word first?

  • Big lists, many equal options

    • Hick's Law / choice overload. How many options at this decision point?

  • Small or distant primary buttons

    • Fitts's Law. Is the primary action big and close?

  • Silent waits

    • Response-time thresholds. Is there feedback when it takes more than a moment?

  • Dense jargon

    • Plain language. Would this read clearly at a general reading level?

← All UX Field Guides

If you enjoyed this resource, consider sharing it with a friend. 

Receive more resources and insights every month by signing up to receive my Unleash UX Newsletter on LinkedIn or through email.

© 2005-2026 Gerren Lamson · v2.1 · Last updated 9/28/26
Google Sites
Report abuse
Page details
Page updated
Google Sites
Report abuse