Short answer

Designers of educational tools should explore integrating AI models to assist in the assessment of written content, focusing on areas where AI demonstrates reliable agreement with human expertise.

Field
Modelling
Source
Journal of Educational Technology and Online Learning (2023)
Method
Comparative analysis
Sample
43 participants
Evidence
Moderate effect

Generative AI models like ChatGPT-3.5 can achieve a moderate level of agreement with human raters when assessing second-language academic writing, suggesting potential for AI to assist in this process. This modelling research insight is drawn from a 2023 study published in Journal of Educational Technology and Online Learning. Using Comparative analysis with 43 participants, researchers explored how this design variable affects real-world outcomes. The key design takeaway: Designers of educational tools should explore integrating AI models to assist in the assessment of written content, focusing on areas where AI demonstrates reliable agreement with human expertise.

Study
ModellingRecentModerate effect

AI-Assisted Academic Writing Assessment: ChatGPT-3.5 Shows Moderate Agreement with Human Raters

Generative AI models like ChatGPT-3.5 can achieve a moderate level of agreement with human raters when assessing second-language academic writing, suggesting potential for AI to assist in this process.

Journal of Educational Technology and Online Learning · 2023

01

Key Findings

  • 01Human raters demonstrated statistically significant low to high positive correlation in their scoring.
  • 02ChatGPT-3.5 showed a slight to fair but significant level of agreement with two of the five human raters.
02

Application

Design takeaway

Designers of educational tools should explore integrating AI models to assist in the assessment of written content, focusing on areas where AI demonstrates reliable agreement with human expertise.

How to apply

In a design project involving educational software, consider implementing an AI-powered feedback mechanism that provides initial scoring and identifies common errors, which can then be reviewed by an instructor.

Project actions

  • 01When evaluating written outputs, consider using AI tools to assist in initial scoring or identification of common errors.
  • 02Compare AI-generated scores with human expert scores to understand the AI's strengths and weaknesses in your specific context.
03

Method & Evidence

AimTo investigate the extent to which generative AI (ChatGPT-3.5) can replicate the scoring of human raters for first-year college students' second-language academic writing.
MethodComparative analysis
ProcedureFirst-year college students completed a paragraph writing task. The same writing criteria were used to evaluate the submissions by five human raters and ChatGPT-3.5. The scores from human raters were analyzed for inter-rater reliability, and the scores from ChatGPT-3.5 were compared against those of the human raters.
Sample43 participants
ContextAcademic writing assessment, educational technology

Variables

IVScoring method (AI vs. Human Rater)
DVScore assigned to written work
CVWriting criteria, writing task, student participant pool
04

Strengths & Limitations

Strengths

  • +Direct comparison of AI and human performance on a specific task.
  • +Analysis of inter-rater reliability among human evaluators.

Limitations

The AI's performance might be limited by the specific training data it used and the complexity of the writing task. The study only used one AI model, so results might differ with other AI systems.

Reliability & validity

Inter-rater reliability among human raters was assessed, and the agreement between AI and human raters was statistically analyzed to determine validity.

Think critically

To what extent can AI truly capture the nuances of academic writing, such as creativity, critical thinking, and cultural context, that human raters might perceive?

05

Design Principles

"Leverage AI for scalable, consistent preliminary evaluation, reserving human expertise for complex judgment and nuanced feedback."

The integration of AI into assessment workflows can streamline the evaluation of written work, potentially reducing the workload on educators and providing more consistent feedback. This opens avenues for developing intelligent tutoring systems and automated feedback tools that support learners.

06

What This Means for Your Design

Computers using AI can score student writing somewhat like a teacher can, but it's best to have a teacher check the AI's work too.

How to use in your project

  • 1.This study can be referenced when discussing the use of AI in assessment, particularly for evaluating written outputs in a design project.
07

Add to My Project

08

Quick Cite

Paragraph starter

Research indicates that generative AI models, such as ChatGPT-3.5, can achieve a moderate level of agreement with human raters when assessing second-language academic writing (Geçkin et al., 2023). This suggests that AI has the potential to serve as a supplementary tool in educational assessment, offering efficiency while maintaining a degree of reliability when used in conjunction with human expertise.

09

Source

Journal of Educational Technology and Online Learning

Assessing second-language academic writing: AI vs. Human raters

journal · 2023

View source

Questions About This Research

What does the research say about ai-assisted academic writing assessment: chatgpt-3.5 shows moderate agreement with human raters?
Designers of educational tools should explore integrating AI models to assist in the assessment of written content, focusing on areas where AI demonstrates reliable agreement with human expertise. Evidence: Journal of Educational Technology and Online Learning (2023).
Why does "AI-Assisted Academic Writing Assessment: ChatGPT-3.5 Shows Moderate Agreement with Human Raters" matter for design?
The integration of AI into assessment workflows can streamline the evaluation of written work, potentially reducing the workload on educators and providing more consistent feedback. This opens avenues for developing intelligent tutoring systems and automated feedback tools that support learners.
How can designers apply this research?
Designers of educational tools should explore integrating AI models to assist in the assessment of written content, focusing on areas where AI demonstrates reliable agreement with human expertise.
What were the main findings?
Human raters demonstrated statistically significant low to high positive correlation in their scoring.. ChatGPT-3.5 showed a slight to fair but significant level of agreement with two of the five human raters.
What research method was used?
Comparative analysis with 43 participants.
How strong is the evidence?
Evidence strength is rated Moderate effect, based on a 2023 journal from Journal of Educational Technology and Online Learning.
What should I do differently in my next project?
In a design project involving educational software, consider implementing an AI-powered feedback mechanism that provides initial scoring and identifies common errors, which can then be reviewed by an instructor.
What are the limitations?
The study focused on a single writing task and a specific AI model; generalizability to other writing genres or AI versions may vary. The agreement level was moderate, not high, indicating AI cannot fully replace human judgment.