Heuristic Audit vs Usability Testing

A heuristic review and a usability test answer different questions. A heuristic review compares an interface with an explicit set of usability principles. A usability test observes relevant participants attempting defined tasks. Neither method proves that a whole service is accessible, compliant or guaranteed to convert.

Short answer

Start with the decision you need to make, not a universal order. Review an implemented interface when the question is about known design principles and consistency. Test with users when the uncertainty concerns behaviour, language, mental models or whether a task works for the intended audience. Sometimes a small review before testing removes obvious noise; sometimes prototype testing should happen before code exists.

What a heuristic review can provide

A reviewer can inspect navigation, labels, feedback, consistency, error recovery, content hierarchy and interaction patterns against a stated framework. This is useful for a live interface or sufficiently detailed prototype, particularly when a team wants a prioritised list of suspected issues before a release or research round.

The evidence is expert judgement tied to the selected heuristics and reviewed states. It does not show how representative users actually interpret the product, and different reviewers or scopes may produce different findings. Accessibility conformance also requires its own standards-based method; a heuristic review should not be relabelled as a WCAG audit.

What usability testing can provide

Moderated usability testing gives participants realistic tasks and lets the team observe what they understand, attempt and find difficult. It is useful for questions that expert inspection cannot settle, including whether terminology makes sense, whether people notice a control, and how a journey fits their expectations.

A test plan should define the research objectives, relevant user groups, recruitment criteria, tasks, prototype or service state, consent and analysis method. GOV.UK recommends clear, actionable objectives and research questions for each round. The output is evidence from that round—not a universal verdict on every user or a statistically certain conversion forecast.

  • Choose participants by the question. Recruit actual or likely users with relevant task experience and access needs.
  • Use realistic tasks. Avoid leading participants toward the interface element being evaluated.
  • Record context. State devices, service version, participant characteristics and test limitations.
  • Analyse promptly. Separate observations from interpretations and connect findings to the research objective.

Choose by product stage

  • Discovery: use user research when the primary uncertainty is who the users are, what they need or how the current task works.
  • Early concept: test sketches or prototypes if the team needs to learn whether the proposed flow and language make sense.
  • Implemented interface: a heuristic review can identify suspected pattern and consistency problems efficiently.
  • Important journey: usability testing can evaluate the task with relevant participants before or after expert review.
  • Accessibility question: combine a standards-based evaluation with involvement of disabled users where appropriate. W3C says user evaluation and WCAG conformance evaluation complement one another.

Evidence and limitations

MethodPrimary evidenceMain limitation
Heuristic reviewExpert inspection against named principlesDoes not observe representative users
Usability testingObserved task behaviour and participant feedbackFindings are bounded by participants, tasks and test context
Accessibility evaluationEvidence against an agreed WCAG scope and methodA sampled review is not a blanket legal certificate

There is no universal participant number

The number and mix of participants depend on the research question, user groups, method and risk. GOV.UK describes a typical range for a qualitative round but recommends more rounds when clearer findings are needed, and much larger samples for surveys, A/B testing and benchmarking. The plan should justify the sample instead of repeating “five users is enough” as a rule.

A practical sequence

  1. Name the decision. Write down what the team must decide after the work.
  2. List current evidence. Include analytics, support themes, prior research, standards findings and known constraints.
  3. Choose the smallest suitable method. Select review, testing, research or a combination based on the evidence gap.
  4. Fix or prototype deliberately. Record which finding each change is intended to address.
  5. Evaluate again. Retest important changes rather than treating a recommendation as proof that the problem is solved.

Method sources

Read next

Need this kind of work done?

Use the brief call to define the decision, product stage, relevant users and evidence already available. The written scope should explain why the proposed method fits.

Request an audit →

← All insights