Short answer
Do not rely solely on vendor claims of AI performance; conduct independent, rigorous testing to validate accuracy and identify potential failure points before integrating AI tools into critical workflows.
- Field
- Innovation & Design
- Source
- Journal of Empirical Legal Studies (2025)
- Method
- Empirical evaluation with a preregistered methodology.
- Evidence
- Strong effect
Leading AI-powered legal research tools, despite claims of being 'hallucination-free,' continue to generate inaccurate information between 17% and 33% of the time. This innovation & design research insight is drawn from a 2025 study published in Journal of Empirical Legal Studies. Using Empirical evaluation with a preregistered methodology., researchers explored how this design variable affects real-world outcomes. The key design takeaway: Do not rely solely on vendor claims of AI performance; conduct independent, rigorous testing to validate accuracy and identify potential failure points before integrating AI tools into critical workflows.
AI Legal Research Tools Hallucinate 17-33% of the Time, Despite Vendor Claims
Leading AI-powered legal research tools, despite claims of being 'hallucination-free,' continue to generate inaccurate information between 17% and 33% of the time.
Journal of Empirical Legal Studies · 2025
Key Findings
- 01AI legal research tools from LexisNexis (Lexis+ AI) and Thomson Reuters (Westlaw AI-Assisted Research and Ask Practical Law AI) hallucinate between 17% and 33% of the time.
- 02Claims of 'hallucination-free' or 'eliminating' hallucinations by vendors are overstated.
- 03Substantial differences exist between systems in terms of responsiveness and accuracy.
Application
Design takeaway
Do not rely solely on vendor claims of AI performance; conduct independent, rigorous testing to validate accuracy and identify potential failure points before integrating AI tools into critical workflows.
How to apply
When evaluating or developing AI tools for professional use, establish clear metrics for accuracy and hallucination, and design testing protocols that mimic real-world usage scenarios to uncover potential flaws.
Project actions
- 01When using AI for research, always cross-reference its output with original sources.
- 02Consider how to design systems that flag potential inaccuracies for user review.
Method & Evidence
Variables
Strengths & Limitations
Strengths
- +First preregistered empirical evaluation of proprietary RAG-based legal AI tools.
- +Development of a comprehensive dataset and typology for identifying AI hallucinations.
Limitations
The specific AI tools tested might not represent all AI technologies, and the legal domain has unique requirements for accuracy.
Reliability & validity
The study's preregistered methodology and comprehensive dataset aim to enhance the reliability and validity of its findings regarding AI hallucination rates.
Think critically
Given the significant hallucination rates, what ethical responsibilities do designers and developers of AI tools have to inform users about these limitations?
Design Principles
"Prioritize empirical validation and transparency in AI system development and deployment."
This research highlights a critical gap between the marketing of advanced AI tools and their actual performance in high-stakes professional environments. Designers and engineers developing AI solutions must prioritize rigorous, independent evaluation over unsubstantiated claims to ensure user trust and mitigate risks.
What This Means for Your Design
Even the best AI tools for lawyers sometimes make things up, about 1 in 5 to 1 in 3 times, so you can't trust them completely without checking.
How to use in your project
- 1.Reference this study when discussing the limitations of AI tools or the importance of user verification in your design project.
Add to My Project
Quick Cite
Paragraph starter
This research highlights the critical issue of AI 'hallucinations' in professional tools, demonstrating that even advanced AI legal research platforms exhibit significant error rates (17-33%). This underscores the necessity for designers to implement robust validation processes and for users to maintain critical oversight, as AI outputs cannot be blindly trusted in high-stakes applications.
Source
Journal of Empirical Legal Studies
Hallucination‐Free? Assessing the Reliability of Leading <scp>AI</scp> Legal Research Tools
journal · 2025
View sourceQuestions About This Research
- What does the research say about ai legal research tools hallucinate 17-33% of the time, despite vendor claims?
- Do not rely solely on vendor claims of AI performance; conduct independent, rigorous testing to validate accuracy and identify potential failure points before integrating AI tools into critical workflows. Evidence: Journal of Empirical Legal Studies (2025).
- Why does "AI Legal Research Tools Hallucinate 17-33% of the Time, Despite Vendor Claims" matter for design?
- This research highlights a critical gap between the marketing of advanced AI tools and their actual performance in high-stakes professional environments. Designers and engineers developing AI solutions must prioritize rigorous, independent evaluation over unsubstantiated claims to ensure user trust and mitigate risks.
- How can designers apply this research?
- Do not rely solely on vendor claims of AI performance; conduct independent, rigorous testing to validate accuracy and identify potential failure points before integrating AI tools into critical workflows.
- What were the main findings?
- AI legal research tools from LexisNexis (Lexis+ AI) and Thomson Reuters (Westlaw AI-Assisted Research and Ask Practical Law AI) hallucinate between 17% and 33% of the time.. Claims of 'hallucination-free' or 'eliminating' hallucinations by vendors are overstated.. Substantial differences exist between systems in terms of responsiveness and accuracy.
- What research method was used?
- Empirical evaluation with a preregistered methodology..
- How strong is the evidence?
- Evidence strength is rated Strong effect, based on a 2025 journal from Journal of Empirical Legal Studies.
- What should I do differently in my next project?
- When evaluating or developing AI tools for professional use, establish clear metrics for accuracy and hallucination, and design testing protocols that mimic real-world usage scenarios to uncover potential flaws.
- What are the limitations?
- The study focused on specific AI legal research tools and may not generalize to all AI applications or legal domains. The 'closed nature' of proprietary systems limits full transparency into their underlying mechanisms.