Short answer

When designing AI systems for medical diagnostics, rigorously assess the source and characteristics of your training data to ensure it accurately reflects the target population and clinical conditions.

Field
User-Centred Design
Source
Journal of Clinical Medicine (2023)
Method
Systematic review and meta-analysis of publicly available datasets.
Evidence
Moderate effect

Publicly available fundus image datasets exhibit significant variability in characteristics and accessibility, posing challenges for the development and reliable application of AI-driven diagnostic tools in ophthalmology. This user-centred design research insight is drawn from a 2023 study published in Journal of Clinical Medicine. Using Systematic review and meta-analysis of publicly available datasets., researchers explored how this design variable affects real-world outcomes. The key design takeaway: When designing AI systems for medical diagnostics, rigorously assess the source and characteristics of your training data to ensure it accurately reflects the target population and clinical conditions.

Study
User-Centred DesignRecentModerate effect

Fundus Image Datasets: Usability and Generalizability Challenges for AI in Ophthalmology

Publicly available fundus image datasets exhibit significant variability in characteristics and accessibility, posing challenges for the development and reliable application of AI-driven diagnostic tools in ophthalmology.

Journal of Clinical Medicine · 2023

01

Key Findings

  • 01A wide range of fundus image datasets exist, but their availability and legality vary considerably.
  • 02Significant differences in dataset characteristics (e.g., image quality, patient demographics, disease prevalence) limit the generalizability of AI models trained on them.
  • 03Barriers to access, such as complex licensing or data usage restrictions, hinder the widespread use of these datasets.
02

Application

Design takeaway

When designing AI systems for medical diagnostics, rigorously assess the source and characteristics of your training data to ensure it accurately reflects the target population and clinical conditions.

How to apply

Before commencing an AI design project involving medical imaging, conduct a thorough audit of available datasets, focusing on their origin, quality, labeling consistency, and potential biases.

Project actions

  • 01When choosing datasets for your design project, don't just pick the easiest one to access; consider its quality and how well it represents the problem you're trying to solve.
  • 02Document any limitations of your chosen datasets clearly in your project report.
03

Method & Evidence

AimTo systematically review and catalog publicly available fundus image datasets, assessing their characteristics, accessibility barriers, usability, and generalizability for AI applications in ophthalmology.
MethodSystematic review and meta-analysis of publicly available datasets.
ProcedureThe researchers conducted a comprehensive search for publicly available fundus image datasets, analyzed their characteristics (e.g., image quality, resolution, labeling), identified legal and access barriers, and evaluated their potential usability and generalizability for AI model training.
ContextOphthalmology, Medical Imaging, Artificial Intelligence, Data Science

Variables

IVCharacteristics of fundus image datasets (e.g., source, quality, labeling, size).
DVUsability and generalizability of datasets for AI applications.
04

Strengths & Limitations

Strengths

  • +Comprehensive review of a specific domain (fundus images).
  • +Addresses a critical bottleneck in AI development: data availability and quality.

Limitations

The availability and quality of datasets can change over time, and this review is a snapshot from 2023. The interpretation of 'usability' and 'generalizability' can be subjective.

Reliability & validity

The reliability of the review depends on the thoroughness of the search strategy and the consistency of the analysis criteria applied to each dataset. Validity is enhanced by the systematic approach to data extraction and synthesis.

Think critically

If a dataset is easy to access and has many images, does that automatically make it a good choice for training an AI model, or are there other, less obvious factors that are more important?

05

Design Principles

"Data representativeness and accessibility are critical for the successful development and deployment of AI-driven design solutions."

Designers and engineers developing AI solutions for medical imaging must critically evaluate the quality, accessibility, and representativeness of training data. Overlooking these factors can lead to AI models that perform poorly in real-world clinical settings, potentially impacting patient care.

06

What This Means for Your Design

When building computer programs that look at eye scans (fundus images) to find diseases, it's hard because the free picture collections are all different and sometimes tricky to get. This means the programs might not work well everywhere.

How to use in your project

  • 1.Reference this study when discussing the challenges of data acquisition and selection for your design project, particularly if your project involves AI or image analysis.
07

Add to My Project

08

Quick Cite

Paragraph starter

The development of effective AI-driven diagnostic tools, such as those for ophthalmology, is significantly challenged by the characteristics and accessibility of publicly available datasets. As demonstrated by Krzywicki et al. (2023), variations in image quality, labeling, and demographic representation across repositories can impede the generalizability of trained models, necessitating careful data selection and validation in any design project.

09

Source

Journal of Clinical Medicine

A Global Review of Publicly Available Datasets Containing Fundus Images: Characteristics, Barriers to Access, Usability, and Generalizability

journal · 2023

View source

Questions About This Research

What does the research say about fundus image datasets: usability and generalizability challenges for ai in ophthalmology?
When designing AI systems for medical diagnostics, rigorously assess the source and characteristics of your training data to ensure it accurately reflects the target population and clinical conditions. Evidence: Journal of Clinical Medicine (2023).
Why does "Fundus Image Datasets: Usability and Generalizability Challenges for AI in Ophthalmology" matter for design?
Designers and engineers developing AI solutions for medical imaging must critically evaluate the quality, accessibility, and representativeness of training data. Overlooking these factors can lead to AI models that perform poorly in real-world clinical settings, potentially impacting patient care.
How can designers apply this research?
When designing AI systems for medical diagnostics, rigorously assess the source and characteristics of your training data to ensure it accurately reflects the target population and clinical conditions.
What were the main findings?
A wide range of fundus image datasets exist, but their availability and legality vary considerably.. Significant differences in dataset characteristics (e.g., image quality, patient demographics, disease prevalence) limit the generalizability of AI models trained on them.. Barriers to access, such as complex licensing or data usage restrictions, hinder the widespread use of these datasets.
What research method was used?
Systematic review and meta-analysis of publicly available datasets..
How strong is the evidence?
Evidence strength is rated Moderate effect, based on a 2023 journal from Journal of Clinical Medicine.
What should I do differently in my next project?
Before commencing an AI design project involving medical imaging, conduct a thorough audit of available datasets, focusing on their origin, quality, labeling consistency, and potential biases.
What are the limitations?
The review is limited to publicly available datasets and may not capture all relevant proprietary or restricted datasets. The assessment of 'usability' and 'generalizability' is based on the reported characteristics of the datasets.