Short answer
Designers of educational tools should explore integrating AI models to assist in the assessment of written content, focusing on areas where AI demonstrates reliable agreement with human expertise.
- Field
- Modelling
- Source
- Journal of Educational Technology and Online Learning (2023)
- Method
- Comparative analysis
- Sample
- 43 participants
- Evidence
- Moderate effect
Generative AI models like ChatGPT-3.5 can achieve a moderate level of agreement with human raters when assessing second-language academic writing, suggesting potential for AI to assist in this process. This modelling research insight is drawn from a 2023 study published in Journal of Educational Technology and Online Learning. Using Comparative analysis with 43 participants, researchers explored how this design variable affects real-world outcomes. The key design takeaway: Designers of educational tools should explore integrating AI models to assist in the assessment of written content, focusing on areas where AI demonstrates reliable agreement with human expertise.
AI-Assisted Academic Writing Assessment: ChatGPT-3.5 Shows Moderate Agreement with Human Raters
Generative AI models like ChatGPT-3.5 can achieve a moderate level of agreement with human raters when assessing second-language academic writing, suggesting potential for AI to assist in this process.
Journal of Educational Technology and Online Learning · 2023
Key Findings
- 01Human raters demonstrated statistically significant low to high positive correlation in their scoring.
- 02ChatGPT-3.5 showed a slight to fair but significant level of agreement with two of the five human raters.
Application
Design takeaway
Designers of educational tools should explore integrating AI models to assist in the assessment of written content, focusing on areas where AI demonstrates reliable agreement with human expertise.
How to apply
In a design project involving educational software, consider implementing an AI-powered feedback mechanism that provides initial scoring and identifies common errors, which can then be reviewed by an instructor.
Project actions
- 01When evaluating written outputs, consider using AI tools to assist in initial scoring or identification of common errors.
- 02Compare AI-generated scores with human expert scores to understand the AI's strengths and weaknesses in your specific context.
Method & Evidence
Variables
Strengths & Limitations
Strengths
- +Direct comparison of AI and human performance on a specific task.
- +Analysis of inter-rater reliability among human evaluators.
Limitations
The AI's performance might be limited by the specific training data it used and the complexity of the writing task. The study only used one AI model, so results might differ with other AI systems.
Reliability & validity
Inter-rater reliability among human raters was assessed, and the agreement between AI and human raters was statistically analyzed to determine validity.
Think critically
To what extent can AI truly capture the nuances of academic writing, such as creativity, critical thinking, and cultural context, that human raters might perceive?
Design Principles
"Leverage AI for scalable, consistent preliminary evaluation, reserving human expertise for complex judgment and nuanced feedback."
The integration of AI into assessment workflows can streamline the evaluation of written work, potentially reducing the workload on educators and providing more consistent feedback. This opens avenues for developing intelligent tutoring systems and automated feedback tools that support learners.
What This Means for Your Design
Computers using AI can score student writing somewhat like a teacher can, but it's best to have a teacher check the AI's work too.
How to use in your project
- 1.This study can be referenced when discussing the use of AI in assessment, particularly for evaluating written outputs in a design project.
Add to My Project
Quick Cite
Paragraph starter
Research indicates that generative AI models, such as ChatGPT-3.5, can achieve a moderate level of agreement with human raters when assessing second-language academic writing (Geçkin et al., 2023). This suggests that AI has the potential to serve as a supplementary tool in educational assessment, offering efficiency while maintaining a degree of reliability when used in conjunction with human expertise.
Source
Journal of Educational Technology and Online Learning
Assessing second-language academic writing: AI vs. Human raters
journal · 2023
View sourceQuestions About This Research
- What does the research say about ai-assisted academic writing assessment: chatgpt-3.5 shows moderate agreement with human raters?
- Designers of educational tools should explore integrating AI models to assist in the assessment of written content, focusing on areas where AI demonstrates reliable agreement with human expertise. Evidence: Journal of Educational Technology and Online Learning (2023).
- Why does "AI-Assisted Academic Writing Assessment: ChatGPT-3.5 Shows Moderate Agreement with Human Raters" matter for design?
- The integration of AI into assessment workflows can streamline the evaluation of written work, potentially reducing the workload on educators and providing more consistent feedback. This opens avenues for developing intelligent tutoring systems and automated feedback tools that support learners.
- How can designers apply this research?
- Designers of educational tools should explore integrating AI models to assist in the assessment of written content, focusing on areas where AI demonstrates reliable agreement with human expertise.
- What were the main findings?
- Human raters demonstrated statistically significant low to high positive correlation in their scoring.. ChatGPT-3.5 showed a slight to fair but significant level of agreement with two of the five human raters.
- What research method was used?
- Comparative analysis with 43 participants.
- How strong is the evidence?
- Evidence strength is rated Moderate effect, based on a 2023 journal from Journal of Educational Technology and Online Learning.
- What should I do differently in my next project?
- In a design project involving educational software, consider implementing an AI-powered feedback mechanism that provides initial scoring and identifies common errors, which can then be reviewed by an instructor.
- What are the limitations?
- The study focused on a single writing task and a specific AI model; generalizability to other writing genres or AI versions may vary. The agreement level was moderate, not high, indicating AI cannot fully replace human judgment.