Short answer

When designing or implementing recommender systems, proactively assess and mitigate the potential for 'bad' recommendations, recognizing that user risk levels are under increasing scrutiny.

Field
Innovation & Markets
Source
International Journal of Information Management (2023)
Method
Meta-analysis of algorithmic audits
Sample
151 algorithmic audits from 33 studies
Evidence
Strong effect

A meta-analysis of algorithmic audits reveals that 8-10% of recommendations from major platforms are 'bad', a risk level comparable to regulated food products, necessitating a re-evaluation of oversight. This innovation & markets research insight is drawn from a 2023 study published in International Journal of Information Management. Using Meta-analysis of algorithmic audits with 151 algorithmic audits from 33 studies, researchers explored how this design variable affects real-world outcomes. The key design takeaway: When designing or implementing recommender systems, proactively assess and mitigate the potential for 'bad' recommendations, recognizing that user risk levels are under increasing scrutiny.

Study
Innovation & MarketsRecentStrong effect

Recommender Algorithms: 8-10% 'Bad' Recommendations Comparable to Food Safety Risks

A meta-analysis of algorithmic audits reveals that 8-10% of recommendations from major platforms are 'bad', a risk level comparable to regulated food products, necessitating a re-evaluation of oversight.

International Journal of Information Management · 2023

01

Key Findings

  • 01Approximately 8-10% of algorithmic recommendations are classified as 'bad'.
  • 02About 25% of recommendations actively protect users from self-induced harm ('do good').
  • 03The observed risk levels are comparable to those monitored for generic food defects by regulatory bodies like the FDA or FSIS.
  • 04Negative feedback loops ('rabbit holes') exist, but positive feedback loops ('good recommendations') are also prevalent.
02

Application

Design takeaway

When designing or implementing recommender systems, proactively assess and mitigate the potential for 'bad' recommendations, recognizing that user risk levels are under increasing scrutiny.

How to apply

When developing a new product or feature that uses algorithmic recommendations, conduct a risk assessment to quantify the potential for negative user experiences and compare it against industry benchmarks and regulatory standards.

Project actions

  • 01When evaluating a system, consider not just its intended function but also its potential negative side effects.
  • 02Use quantitative data to support claims about the effectiveness or risks of a design.
03

Method & Evidence

AimTo assess the frequency of 'good' and 'bad' recommendations generated by algorithmic systems across various platforms and to understand the implications for regulatory oversight.
MethodMeta-analysis of algorithmic audits
ProcedureThe researchers quantitatively analyzed 151 algorithmic audits from 33 studies that reported risk-utility statistics for recommender algorithms on platforms like YouTube, Google Search, Twitter, Facebook, TikTok, and Amazon.
Sample151 algorithmic audits from 33 studies
ContextDigital platforms and recommender systems

Variables

IVPlatform type, type of risk (bias, mental health, misinformation, extremism)
DVPercentage of 'bad' recommendations, percentage of 'do good' recommendations
CVType of algorithmic audit, risk-utility statistics reported
04

Strengths & Limitations

Strengths

  • +Wide-ranging evaluation across multiple platforms.
  • +Quantitative meta-analysis of existing audit data.

Limitations

The definition of 'bad' recommendations can be subjective and may vary across different user groups or contexts. The study did not delve into the severity of the harm caused by these recommendations.

Reliability & validity

The reliability of the findings is supported by the consistency of the 8-10% figure across numerous audits and platforms. Validity is enhanced by the meta-analytic approach, which synthesizes data from multiple studies, though the 'coarse-grained' nature of the analysis might limit the depth of causal claims.

Think critically

If 8-10% of recommendations are 'bad', but 25% are 'good' at protecting users, how should designers balance the pursuit of engagement with the responsibility for user well-being?

05

Design Principles

"Algorithmic recommendations should be designed with a quantifiable risk-utility balance, aiming for risk levels comparable to or better than those in established consumer product safety regulations."

This finding is critical for businesses deploying recommender systems, as it quantifies a significant level of user risk that could impact brand reputation, user trust, and potentially lead to regulatory scrutiny. Understanding these risks allows for proactive design and mitigation strategies.

06

What This Means for Your Design

Imagine a website that suggests videos to watch. This study found that about 1 in 10 of those suggestions are actually unhelpful or even harmful. This is a big deal because it's as risky as some food products that are carefully checked by the government.

How to use in your project

  • 1.Reference this study when discussing the potential negative impacts of algorithmic design choices in your design project, particularly if your project involves recommendation systems or user engagement strategies.
07

Add to My Project

08

Quick Cite

Paragraph starter

The meta-analysis by Hilbert et al. (2023) indicates that recommender algorithms on major digital platforms consistently generate 8-10% 'bad' recommendations, a risk level comparable to regulated food products. This underscores the importance of rigorous risk assessment and mitigation strategies in the design of algorithmic systems to ensure user safety and trust.

09

Source

International Journal of Information Management

8–10% of algorithmic recommendations are ‘bad’, but… an exploratory risk-utility meta-analysis and its regulatory implications

journal · 2023

View source

Questions About This Research

What does the research say about recommender algorithms: 8-10% 'bad' recommendations comparable to food safety risks?
When designing or implementing recommender systems, proactively assess and mitigate the potential for 'bad' recommendations, recognizing that user risk levels are under increasing scrutiny. Evidence: International Journal of Information Management (2023).
Why does "Recommender Algorithms: 8-10% 'Bad' Recommendations Comparable to Food Safety Risks" matter for design?
This finding is critical for businesses deploying recommender systems, as it quantifies a significant level of user risk that could impact brand reputation, user trust, and potentially lead to regulatory scrutiny. Understanding these risks allows for proactive design and mitigation strategies.
How can designers apply this research?
When designing or implementing recommender systems, proactively assess and mitigate the potential for 'bad' recommendations, recognizing that user risk levels are under increasing scrutiny.
What were the main findings?
Approximately 8-10% of algorithmic recommendations are classified as 'bad'.. About 25% of recommendations actively protect users from self-induced harm ('do good').. The observed risk levels are comparable to those monitored for generic food defects by regulatory bodies like the FDA or FSIS.. Negative feedback loops ('rabbit holes') exist, but positive feedback loops ('good recommendations') are also prevalent.
What research method was used?
Meta-analysis of algorithmic audits with 151 algorithmic audits from 33 studies.
How strong is the evidence?
Evidence strength is rated Strong effect, based on a 2023 journal from International Journal of Information Management.
What should I do differently in my next project?
When developing a new product or feature that uses algorithmic recommendations, conduct a risk assessment to quantify the potential for negative user experiences and compare it against industry benchmarks and regulatory standards.
What are the limitations?
The analysis is quantitatively coarse-grained and refrains from judging the causal consequences or severity of risks, focusing instead on the frequency of 'bad' recommendations.