Short answer

Leverage LLMs and agentic systems for efficient and scalable analysis of textual data in design research projects.

Field
Innovation & Design
Source
Applied Sciences (2025)
Method
Experimental Evaluation
Sample
200 news articles
Evidence
Strong effect

Large Language Models (LLMs), especially when deployed in multi-agent systems, can automate structured media content analysis with accuracy approaching commercial standards, significantly reducing the need for manual coding. This innovation & design research insight is drawn from a 2025 study published in Applied Sciences. Using Experimental evaluation with 200 news articles, researchers explored how this design variable affects real-world outcomes. The key design takeaway: Leverage LLMs and agentic systems for efficient and scalable analysis of textual data in design research projects.

Study
Innovation & DesignNew This WeekStrong effect

LLMs Achieve Near Commercial Accuracy in Automated Content Analysis

Large Language Models (LLMs), especially when deployed in multi-agent systems, can automate structured media content analysis with accuracy approaching commercial standards, significantly reducing the need for manual coding.

Applied Sciences · 2025

01

Key Findings

  • 01LLMs achieved weighted global F1-scores between 0.636 and 0.822 in direct prompting.
  • 02Claude-3-7-Sonnet showed the highest direct-prompt performance.
  • 03A multi-agent system improved the F1-score from 0.757 to 0.805 compared to direct prompting.
  • 04The multi-agent system demonstrated consistent gains across different task types (binary, categorical, multi-label).
02

Application

Design takeaway

Leverage LLMs and agentic systems for efficient and scalable analysis of textual data in design research projects.

How to apply

Use LLMs to analyze user reviews, survey responses, or competitor documentation to quickly identify key themes and sentiment.

Project actions

  • 01Consider using LLMs to help analyze qualitative data from user interviews or open-ended survey questions.
  • 02Explore how different prompting strategies can influence the accuracy of LLM-based analysis for your specific design context.
03

Method & Evidence

AimCan large language models and multi-agent systems effectively automate structured media content analysis based on codebook-driven annotation?
MethodExperimental Evaluation
ProcedureA dataset of 200 news articles was manually annotated using a detailed codebook. Seven LLMs were tested using zero-shot prompting with role-based instructions and schema-constrained outputs. A multi-agent system was then developed using a specific LLM, incorporating expert role profiling, shared memory, and coordinated planning, and its performance was compared to direct prompting.
Sample200 news articles
ContextNews Content Analysis

Variables

IV["Type of LLM (low- to high-capacity)","Direct prompting vs. multi-agent system","Prompting strategy (zero-shot, role-based instructions, schema-constrained outputs)"]
DV["Weighted global F1-score","Accuracy across binary, categorical, and multi-label tasks"]
CV["Dataset size (200 articles)","Codebook complexity (26 questions, 122 codes)","News articles on U.S. tariff policies"]
04

Strengths & Limitations

Strengths

  • +Rigorous ground truth established through manual annotation.
  • +Evaluation of multiple LLMs across different capacity tiers.
  • +Development and testing of a novel multi-agent system architecture.

Limitations

LLMs might not fully grasp nuanced cultural contexts or highly specialized jargon without specific fine-tuning. The 'ground truth' data used for training or evaluation is critical for accuracy.

Reliability & validity

The study establishes reliability through a manually annotated ground truth and assesses validity by comparing LLM performance against this benchmark across various task types. The use of F1-scores provides a robust measure of both precision and recall.

Think critically

How might the 'ground truth' data used in this study have influenced the LLMs' performance, and what are the implications for applying these models to datasets with different characteristics or biases?

05

Design Principles

"Automate repetitive analytical tasks using AI to free up human cognitive resources for higher-level design thinking and strategy."

This advancement offers a scalable and cost-effective solution for analyzing large volumes of media content, which is crucial for understanding public opinion, market trends, and the impact of communication strategies. Designers and researchers can leverage these tools to gain faster, more comprehensive insights, informing design decisions and product development.

06

What This Means for Your Design

Computers that understand language (like ChatGPT) can now read lots of news articles and sort them into categories almost as well as people can, and using a team of these computer 'brains' can make them even better.

How to use in your project

  • 1.Reference this study when discussing the use of AI tools for data analysis in your design project, particularly for qualitative or textual data.
07

Add to My Project

08

Quick Cite

Paragraph starter

This research demonstrates the efficacy of Large Language Models (LLMs) in automating structured content analysis, achieving accuracy levels comparable to commercial standards. The study highlights that advanced LLMs, particularly when deployed within multi-agent systems, can significantly reduce the manual effort required for tasks such as coding and categorizing textual data, offering a scalable and cost-effective approach for design research.

09

Source

Applied Sciences

Beyond Manual Media Coding: Evaluating Large Language Models and Agents for News Content Analysis

journal · 2025

View source

Questions About This Research

What does the research say about llms achieve near commercial accuracy in automated content analysis?
Leverage LLMs and agentic systems for efficient and scalable analysis of textual data in design research projects. Evidence: Applied Sciences (2025).
Why does "LLMs Achieve Near Commercial Accuracy in Automated Content Analysis" matter for design?
This advancement offers a scalable and cost-effective solution for analyzing large volumes of media content, which is crucial for understanding public opinion, market trends, and the impact of communication strategies. Designers and researchers can leverage these tools to gain faster, more comprehensive insights, informing design decisions and product development.
How can designers apply this research?
Leverage LLMs and agentic systems for efficient and scalable analysis of textual data in design research projects.
What were the main findings?
LLMs achieved weighted global F1-scores between 0.636 and 0.822 in direct prompting.. Claude-3-7-Sonnet showed the highest direct-prompt performance.. A multi-agent system improved the F1-score from 0.757 to 0.805 compared to direct prompting.. The multi-agent system demonstrated consistent gains across different task types (binary, categorical, multi-label).
What research method was used?
Experimental Evaluation with 200 news articles.
How strong is the evidence?
Evidence strength is rated Strong effect, based on a 2025 journal from Applied Sciences.
What should I do differently in my next project?
Use LLMs to analyze user reviews, survey responses, or competitor documentation to quickly identify key themes and sentiment.
What are the limitations?
The performance of LLMs can be sensitive to prompt engineering and the quality of the ground truth data. The cost-effectiveness may vary depending on the specific LLM and usage volume.