Short answer
When using AI for design evaluation, prioritize text-based inputs for LLMs and validate AI-generated similarity assessments with human feedback, especially for critical design decisions.
- Field
- Modelling
- Source
- Journal of Mechanical Design (2025)
- Method
- Comparative analysis of embedding spaces
- Evidence
- Moderate effect
Language models can partially replicate human perception of design similarity, but their alignment is not perfect and can be influenced by the input modality. This modelling research insight is drawn from a 2025 study published in Journal of Mechanical Design. Using Comparative analysis of embedding spaces, researchers explored how this design variable affects real-world outcomes. The key design takeaway: When using AI for design evaluation, prioritize text-based inputs for LLMs and validate AI-generated similarity assessments with human feedback, especially for critical design decisions.
LLM Design Similarity Embeddings Align Moderately with Human Judgments
Language models can partially replicate human perception of design similarity, but their alignment is not perfect and can be influenced by the input modality.
Journal of Mechanical Design · 2025
Key Findings
- 01Text-based language model embeddings show moderate alignment with human judgments of design similarity.
- 02Multimodal (image and text) embeddings did not consistently improve alignment with human judgments compared to text-only embeddings.
- 03Local tripletwise similarity is a more nuanced metric for assessing alignment than raw Likert-scale scores.
Application
Design takeaway
When using AI for design evaluation, prioritize text-based inputs for LLMs and validate AI-generated similarity assessments with human feedback, especially for critical design decisions.
How to apply
In a design project, use an LLM to generate similarity scores between design concepts based on their textual descriptions. Then, conduct a small user study to compare these AI-generated scores with human ratings to identify areas of agreement and disagreement.
Project actions
- 01When comparing AI and human judgments, clearly define what 'similarity' means in your project context.
- 02Consider using qualitative feedback alongside quantitative similarity scores to understand the reasoning behind human preferences.
Method & Evidence
Variables
Strengths & Limitations
Strengths
- +Utilizes triplet comparisons for a more nuanced assessment of similarity.
- +Investigates the impact of different input modalities for AI evaluation.
Limitations
The AI models used might have biases from their training data, and the human participants might have different interpretations of design similarity.
Reliability & validity
Reliability could be assessed by repeating the human similarity judgments with a new group of participants. Validity is supported by the use of triplet comparisons, which are considered a robust method for capturing ordinal preferences, and by comparing against human ground truth.
Think critically
How might the specific training data of an LLM influence its perception of design similarity, and how could this bias be addressed in a design project?
Design Principles
"AI-driven design evaluation should be treated as a supplementary tool, requiring human oversight and validation to ensure alignment with user and stakeholder needs."
As design teams increasingly leverage AI for concept evaluation, understanding the alignment between AI and human judgment is crucial for effective integration. This research provides insights into the limitations and potential of using AI models as design evaluators, guiding practitioners on how to best utilize these tools.
What This Means for Your Design
AI can help guess if designs are similar like humans do, but it's not perfect. Text alone sometimes works better than using pictures too.
How to use in your project
- 1.Reference this study when discussing the use of AI tools for design concept evaluation or when comparing AI-generated data with user research findings.
Add to My Project
Quick Cite
Paragraph starter
This research highlights the potential and limitations of using language models for design evaluation. Findings suggest that while AI can offer moderate alignment with human judgments of design similarity, particularly through text-based embeddings, human oversight remains critical. This informs our approach by emphasizing the need to validate AI-generated insights with qualitative user feedback to ensure comprehensive design assessment.
Source
Journal of Mechanical Design
Exploring Human and Language Model Alignment in Perceived Design Similarity Using Ordinal Embeddings
journal · 2025
View sourceQuestions About This Research
- What does the research say about llm design similarity embeddings align moderately with human judgments?
- When using AI for design evaluation, prioritize text-based inputs for LLMs and validate AI-generated similarity assessments with human feedback, especially for critical design decisions. Evidence: Journal of Mechanical Design (2025).
- Why does "LLM Design Similarity Embeddings Align Moderately with Human Judgments" matter for design?
- As design teams increasingly leverage AI for concept evaluation, understanding the alignment between AI and human judgment is crucial for effective integration. This research provides insights into the limitations and potential of using AI models as design evaluators, guiding practitioners on how to best utilize these tools.
- How can designers apply this research?
- When using AI for design evaluation, prioritize text-based inputs for LLMs and validate AI-generated similarity assessments with human feedback, especially for critical design decisions.
- What were the main findings?
- Text-based language model embeddings show moderate alignment with human judgments of design similarity.. Multimodal (image and text) embeddings did not consistently improve alignment with human judgments compared to text-only embeddings.. Local tripletwise similarity is a more nuanced metric for assessing alignment than raw Likert-scale scores.
- What research method was used?
- Comparative analysis of embedding spaces.
- How strong is the evidence?
- Evidence strength is rated Moderate effect, based on a 2025 journal from Journal of Mechanical Design.
- What should I do differently in my next project?
- In a design project, use an LLM to generate similarity scores between design concepts based on their textual descriptions. Then, conduct a small user study to compare these AI-generated scores with human ratings to identify areas of agreement and disagreement.
- What are the limitations?
- The study focused on specific types of design representations (sketches and descriptions) and may not generalize to all design domains or modalities. The definition of 'similarity' can be subjective and vary across individuals.