Short answer

Incorporate LLM-driven heuristic analysis as a preliminary step in usability testing for complex engineering software to quickly identify a substantial portion of potential issues, thereby optimizing the use of human expert resources.

Field
User-Centred Design
Source
Applied Sciences (2026)
Method
Comparative Heuristic Usability Evaluation
Evidence
Moderate effect

Leveraging Large Language Models (LLMs) alongside human expert review can significantly enhance the efficiency and scope of heuristic usability evaluations for complex engineering software. This user-centred design research insight is drawn from a 2026 study published in Applied Sciences. Using Comparative heuristic usability evaluation, researchers explored how this design variable affects real-world outcomes. The key design takeaway: Incorporate LLM-driven heuristic analysis as a preliminary step in usability testing for complex engineering software to quickly identify a substantial portion of potential issues, thereby optimizing the use of human expert resources.

Study
User-Centred DesignNew This WeekModerate effect

LLM-Assisted Heuristic Evaluation Identifies 55% of Usability Issues in Automotive REM Tools

Leveraging Large Language Models (LLMs) alongside human expert review can significantly enhance the efficiency and scope of heuristic usability evaluations for complex engineering software.

Applied Sciences · 2026

01

Key Findings

  • 01Human evaluators identified usability issues within the REM tool.
  • 02The LLM identified 55% of its detected issues as valid, with 32% overlapping with human findings and 23% being novel.
  • 03LLMs show potential as complementary tools for accelerating early-stage heuristic inspections.
02

Application

Design takeaway

Incorporate LLM-driven heuristic analysis as a preliminary step in usability testing for complex engineering software to quickly identify a substantial portion of potential issues, thereby optimizing the use of human expert resources.

How to apply

When evaluating complex software, use an LLM with relevant heuristics and screenshots to generate an initial list of potential usability problems. Then, have human experts review and validate these findings, focusing their efforts on the most critical or unique issues identified by the AI.

Project actions

  • 01When evaluating a product, consider using an AI tool to help identify potential usability issues before conducting user testing.
  • 02Clearly document the prompts and parameters used when employing AI for design analysis.
03

Method & Evidence

AimTo evaluate the usability of an automotive Requirements Engineering and Management (REM) tool using both human expert heuristics and an LLM-based approach, and to compare the effectiveness of each method.
MethodComparative Heuristic Usability Evaluation
ProcedureHuman experts conducted a heuristic usability evaluation of the IBM DOORS Next Generation REM tool using Nielsen's 10 Usability Heuristics. Subsequently, an LLM (ChatGPT-5) was prompted with the same heuristics and static screenshots of the tool to perform its own evaluation. The findings from both methods were compared to identify overlapping and unique issues.
ContextAutomotive engineering software (Requirements Engineering and Management tools)

Variables

IV["Evaluation method (Human Expert vs. LLM)","Type of usability issue"]
DV["Number of identified usability issues","Percentage of overlapping issues","Percentage of novel issues identified by LLM"]
CV["Usability heuristics used (Nielsen's 10)","Target software (IBM DOORS Next Generation)","Screenshots used for LLM evaluation"]
04

Strengths & Limitations

Strengths

  • +Comparative analysis of two distinct evaluation methods.
  • +Application to a relevant and complex domain (automotive engineering software).

Limitations

AI tools may not fully grasp context or nuanced user emotions. Their findings need human validation. The quality of AI output depends heavily on the input and the AI model itself.

Reliability & validity

Reliability could be improved by using multiple LLMs or having a larger group of human experts. Validity is supported by the comparison to human findings and the focus on established heuristics, but the LLM's 'unconfirmed' findings require further validation through user testing.

Think critically

To what extent can LLMs truly understand the subjective aspects of user experience, such as delight or frustration, which are crucial for holistic design evaluation?

05

Design Principles

"Augment human expertise with AI for efficient and comprehensive design evaluation."

Requirements Engineering and Management (REM) tools are critical for safety-critical industries like automotive design. Poor usability in these tools can lead to workflow disruptions and compliance risks. This research demonstrates how AI can augment traditional usability testing, offering a faster, broader initial assessment.

06

What This Means for Your Design

Using AI like ChatGPT can help find many of the same problems in software design that human experts do, and sometimes even find new ones, making the testing process faster.

How to use in your project

  • 1.Use AI tools to generate initial hypotheses about usability problems in your design, which can then be tested through user research.
  • 2.Compare AI-generated findings with your own observations or expert reviews to identify strengths and weaknesses of the AI approach.
07

Add to My Project

08

Quick Cite

Paragraph starter

This design project explored the potential of AI in usability evaluation. By employing a Large Language Model (LLM) alongside human heuristic review, we aimed to assess the efficiency and effectiveness of AI in identifying usability issues within complex engineering software. The LLM successfully flagged a significant percentage of relevant issues, demonstrating its capacity to augment traditional design research methods and accelerate the identification of potential design flaws.

09

Source

Applied Sciences

Hybrid Usability Evaluation of an Automotive REM Tool: Human and LLM-Based Heuristic Assessment of IBM Doors Next

journal · 2026

View source

Questions About This Research

What does the research say about llm-assisted heuristic evaluation identifies 55% of usability issues in automotive rem tools?
Incorporate LLM-driven heuristic analysis as a preliminary step in usability testing for complex engineering software to quickly identify a substantial portion of potential issues, thereby optimizing the use of human expert resources. Evidence: Applied Sciences (2026).
Why does "LLM-Assisted Heuristic Evaluation Identifies 55% of Usability Issues in Automotive REM Tools" matter for design?
Requirements Engineering and Management (REM) tools are critical for safety-critical industries like automotive design. Poor usability in these tools can lead to workflow disruptions and compliance risks. This research demonstrates how AI can augment traditional usability testing, offering a faster, broader initial assessment.
How can designers apply this research?
Incorporate LLM-driven heuristic analysis as a preliminary step in usability testing for complex engineering software to quickly identify a substantial portion of potential issues, thereby optimizing the use of human expert resources.
What were the main findings?
Human evaluators identified usability issues within the REM tool.. The LLM identified 55% of its detected issues as valid, with 32% overlapping with human findings and 23% being novel.. LLMs show potential as complementary tools for accelerating early-stage heuristic inspections.
What research method was used?
Comparative Heuristic Usability Evaluation.
How strong is the evidence?
Evidence strength is rated Moderate effect, based on a 2026 journal from Applied Sciences.
What should I do differently in my next project?
When evaluating complex software, use an LLM with relevant heuristics and screenshots to generate an initial list of potential usability problems. Then, have human experts review and validate these findings, focusing their efforts on the most critical or unique issues identified by the AI.
What are the limitations?
The LLM's evaluation was based on static screenshots, lacking dynamic interaction. The LLM's unconfirmed findings require further validation. The specific LLM used may influence results.