Short answer

Incorporate LLM-based tools for evaluating the novelty and creativity of metaphorical concepts during the ideation and concept development phases of a design project.

Field
Innovation & Design
Source
Creativity Research Journal (2024)
Method
Quantitative, computational modeling
Sample
4,589 responses from 1,546 participants
Evidence
Strong effect

Large Language Models, when fine-tuned on human ratings, can accurately and reliably score the creativity of metaphors, significantly outperforming traditional semantic distance measures. This innovation & design research insight is drawn from a 2024 study published in Creativity Research Journal. Using Quantitative, computational modeling with 4,589 responses from 1,546 participants, researchers explored how this design variable affects real-world outcomes. The key design takeaway: Incorporate LLM-based tools for evaluating the novelty and creativity of metaphorical concepts during the ideation and concept development phases of a design project.

Study
Innovation & DesignRecentStrong effect

LLMs Can Automate Metaphor Creativity Scoring with High Reliability

Large Language Models, when fine-tuned on human ratings, can accurately and reliably score the creativity of metaphors, significantly outperforming traditional semantic distance measures.

Creativity Research Journal · 2024

01

Key Findings

  • 01Fine-tuned LLMs (RoBERTa and GPT-2) reliably predicted human creativity ratings for new metaphors (r = .72 and r = .70, respectively).
  • 02LLMs significantly outperformed semantic distance measures (r = .42).
  • 03The fine-tuned models demonstrated good generalization to metaphor prompts not seen during training (RoBERTa r = .68, GPT-2 r = .63).
02

Application

Design takeaway

Incorporate LLM-based tools for evaluating the novelty and creativity of metaphorical concepts during the ideation and concept development phases of a design project.

How to apply

Use fine-tuned LLMs to score the originality of user-generated ideas, marketing taglines, or conceptual product descriptions.

Project actions

  • 01Consider how you can use AI tools to help evaluate the creativity of your design concepts.
  • 02If your project involves generating new ideas or descriptions, explore using LLMs to score their originality.
03

Method & Evidence

AimCan fine-tuned Large Language Models accurately and reliably score the creativity of metaphors, and do they generalize to novel prompts?
MethodQuantitative, computational modeling
ProcedureTwo open-source LLMs (RoBERTa and GPT-2) were fine-tuned using a dataset of 4,589 metaphor responses and their corresponding human creativity ratings. The performance of these fine-tuned models was then evaluated against new human ratings and compared to a semantic distance scoring method. Their ability to generalize to metaphor prompts not included in the training data was also tested.
Sample4,589 responses from 1,546 participants
ContextCreative thinking tasks, specifically metaphor generation.

Variables

IVFine-tuned LLMs (RoBERTa, GPT-2), Semantic Distance
DVHuman creativity ratings (originality)
CVMetaphor prompts, participant responses, human rating criteria
04

Strengths & Limitations

Strengths

  • +Large dataset size.
  • +Demonstrated generalization to unseen prompts.
  • +Direct comparison with a baseline method (semantic distance).

Limitations

The LLM's performance is tied to the data it was trained on. If your design project uses very unique or niche language, the LLM might not score it accurately.

Reliability & validity

Reliability is demonstrated by the high correlation coefficients (r) between LLM scores and human ratings. Validity is supported by the LLMs' ability to generalize to new prompts and outperform a known, albeit weaker, scoring method.

Think critically

To what extent can LLMs truly capture the nuanced and subjective nature of human creativity, especially in domains beyond metaphor generation?

05

Design Principles

"Automate subjective creative evaluation where possible to increase speed and consistency."

This research introduces a scalable and efficient method for evaluating a key aspect of creative thinking. Automating this process frees up human resources and provides more consistent, reproducible results, which is invaluable for design research and development where idea generation and evaluation are critical.

06

What This Means for Your Design

Computers can now be taught to judge how creative a metaphor is, just like a person can, and they do it much faster and more consistently than before.

How to use in your project

  • 1.Reference this study when discussing methods for evaluating the creativity of your design concepts or user-generated ideas.
07

Add to My Project

08

Quick Cite

Paragraph starter

This research demonstrates that Large Language Models can be effectively fine-tuned to automate the scoring of metaphor creativity, achieving high reliability and accuracy comparable to human raters. This approach offers a significant advantage in terms of speed and consistency for evaluating creative outputs in design projects, particularly when dealing with a large volume of ideas or user feedback.

09

Source

Creativity Research Journal

Automatic Scoring of Metaphor Creativity with Large Language Models

journal · 2024

View source

Questions About This Research

What does the research say about llms can automate metaphor creativity scoring with high reliability?
Incorporate LLM-based tools for evaluating the novelty and creativity of metaphorical concepts during the ideation and concept development phases of a design project. Evidence: Creativity Research Journal (2024).
Why does "LLMs Can Automate Metaphor Creativity Scoring with High Reliability" matter for design?
This research introduces a scalable and efficient method for evaluating a key aspect of creative thinking. Automating this process frees up human resources and provides more consistent, reproducible results, which is invaluable for design research and development where idea generation and evaluation are critical.
How can designers apply this research?
Incorporate LLM-based tools for evaluating the novelty and creativity of metaphorical concepts during the ideation and concept development phases of a design project.
What were the main findings?
Fine-tuned LLMs (RoBERTa and GPT-2) reliably predicted human creativity ratings for new metaphors (r = .72 and r = .70, respectively).. LLMs significantly outperformed semantic distance measures (r = .42).. The fine-tuned models demonstrated good generalization to metaphor prompts not seen during training (RoBERTa r = .68, GPT-2 r = .63).
What research method was used?
Quantitative, computational modeling with 4,589 responses from 1,546 participants.
How strong is the evidence?
Evidence strength is rated Strong effect, based on a 2024 journal from Creativity Research Journal.
What should I do differently in my next project?
Use fine-tuned LLMs to score the originality of user-generated ideas, marketing taglines, or conceptual product descriptions.
What are the limitations?
The accuracy of LLM scoring is dependent on the quality and diversity of the training data. The models may exhibit biases present in the training corpus. Generalization to highly novel or culturally specific metaphors might still be a challenge.