Study
Classic DesignRecentStrong effect

Expert-level multimodal AI struggles with domain-specific visual reasoning

Current advanced multimodal AI models demonstrate significant limitations when tasked with understanding and reasoning across diverse, expert-level visual information from various academic disciplines.

arXiv (Cornell University) · 2023

01

Key Findings

  • 01Advanced multimodal AI models, including GPT-4V and Gemini Ultra, achieved accuracies of only 56% and 59% respectively on the MMMU benchmark.
  • 02The benchmark covers a wide range of disciplines (Art & Design, Business, Science, Health & Medicine, Humanities & Social Science, Tech & Engineering) and image types, posing significant challenges for AI perception and reasoning.
02

Application

Design takeaway

AI is not yet a substitute for human expert visual analysis and reasoning in complex, specialized domains.

How to apply

When using AI for design research or analysis involving complex visuals (e.g., interpreting technical diagrams, analyzing artistic compositions, understanding scientific charts), always cross-reference AI outputs with human expert knowledge and critical judgment.

Project actions

  • 01When evaluating AI tools for your design project, consider their ability to handle domain-specific visual information.
  • 02If your project involves complex visual data, acknowledge the current limitations of AI in interpreting it accurately.
03

Method & Evidence

AimTo what extent can current state-of-the-art multimodal AI models perform expert-level visual reasoning across a broad spectrum of academic disciplines?
MethodBenchmark evaluation
ProcedureA comprehensive benchmark (MMMU) was created, comprising over 11,500 multimodal questions from college-level exams and textbooks across six core disciplines. These questions involved diverse image types (charts, diagrams, maps, chemical structures, etc.) and required domain-specific knowledge and reasoning. Fourteen open-source models and two proprietary models (GPT-4V, Gemini) were evaluated on this benchmark.
Sample11,500+ questions
ContextArtificial Intelligence, Multimodal Understanding, Expert Reasoning, Academic Domains

Variables

IVType of AI model (e.g., GPT-4V, Gemini, open-source models)
DVAccuracy on multimodal understanding and reasoning tasks
CVDiscipline, subject matter, image type, complexity of reasoning required
04

Strengths & Limitations

Strengths

  • +Massive scale of the benchmark with over 11.5K questions.
  • +Inclusion of diverse academic disciplines and heterogeneous image types.

Limitations

The AI models tested might not represent the absolute cutting edge of AI development, and new models are constantly emerging. The specific disciplines and image types in the benchmark might not cover all areas relevant to a particular design project.

Reliability & validity

The validity of the benchmark relies on the quality and representativeness of the questions. Reliability would be assessed by the consistency of AI model performance across similar question types.

Think critically

Given the limitations of AI in expert-level visual reasoning, how can designers best leverage AI as a supplementary tool rather than a replacement for their own analytical and creative skills?

05

Design Principles

"Human expertise remains critical for nuanced visual interpretation and domain-specific reasoning."

This highlights a critical gap in the development of AI systems that can truly emulate human expertise. For designers, engineers, and researchers, it suggests that AI tools are not yet capable of nuanced, domain-specific visual interpretation required for complex problem-solving or creative ideation.

06

What This Means for Your Design

Even the smartest AI can't yet understand complex pictures and information like a human expert in subjects like art, science, or business.

How to use in your project

  • 1.Cite this research when discussing the limitations of AI tools in your design project, especially if you are using AI for visual analysis or idea generation.
07

Add to My Project

08

Quick Cite

(2023). MMMU: A Massive Multi-discipline Multimodal Understanding and Reasoning Benchmark for Expert AGI. arXiv (Cornell University). https://doi.org/10.48550/arxiv.2311.16502 Retrieved from https://designdex.org/study/f279da42-1578-4dec-8f18-0fe7748fc697/expert-level-multimodal-ai-struggles-with-domain-specific-visual-reasoning

Paragraph starter

Current research indicates that even advanced multimodal AI models exhibit significant limitations in performing expert-level visual reasoning across diverse academic disciplines, achieving accuracies below 60% on comprehensive benchmarks. This suggests that AI tools are not yet capable of the nuanced, domain-specific interpretation required for complex design challenges, necessitating continued reliance on human expertise for critical visual analysis and problem-solving.

09

Source

arXiv (Cornell University)

MMMU: A Massive Multi-discipline Multimodal Understanding and Reasoning Benchmark for Expert AGI

journal · 2023

View source

Questions about this research

What does the research say about expert-level multimodal ai struggles with domain-specific visual reasoning?
AI is not yet a substitute for human expert visual analysis and reasoning in complex, specialized domains. Evidence: arXiv (Cornell University) (2023).
Why does "Expert-level multimodal AI struggles with domain-specific visual reasoning" matter for design?
This highlights a critical gap in the development of AI systems that can truly emulate human expertise. For designers, engineers, and researchers, it suggests that AI tools are not yet capable of nuanced, domain-specific visual interpretation required for complex problem-solving or creative ideation.
How can designers apply this research?
AI is not yet a substitute for human expert visual analysis and reasoning in complex, specialized domains.
What were the main findings?
Advanced multimodal AI models, including GPT-4V and Gemini Ultra, achieved accuracies of only 56% and 59% respectively on the MMMU benchmark.. The benchmark covers a wide range of disciplines (Art & Design, Business, Science, Health & Medicine, Humanities & Social Science, Tech & Engineering) and image types, posing significant challenges for AI perception and reasoning.
What research method was used?
Benchmark evaluation with 11,500+ questions.
How strong is the evidence?
Evidence strength is rated Strong effect, based on a 2023 journal from arXiv (Cornell University).
What should I do differently in my next project?
When using AI for design research or analysis involving complex visuals (e.g., interpreting technical diagrams, analyzing artistic compositions, understanding scientific charts), always cross-reference AI outputs with human expert knowledge and critical judgment.
What are the limitations?
The benchmark may not encompass all possible forms of expert visual data or reasoning tasks. The performance of AI models is a snapshot in time and may improve with further development.
Is there evidence that domain-specific visual affects design outcomes?
Even the most advanced AI models struggle to achieve expert-level understanding and reasoning when presented with complex, domain-specific visual information from various academic fields. This highlights a critical gap in the development of AI systems that can truly emulate human expertise. For designers, engineers, an Source: arXiv (Cornell University) (2023).
Where does this visual reasoning research apply?
Artificial Intelligence, Multimodal Understanding, Expert Reasoning, Academic Domains It sits within classic design research on designdex.org.

Related research topics

domain-specific visual design research · evidence on domain-specific visual · does domain-specific visual improve design outcomes · visual reasoning studies for designers · domain-specific visual and visual reasoning findings · classic design research evidence