Short answer

Designers should focus on developing AI agents that can handle the inherent complexity and variability of real-world online interactions, prioritizing user needs for reliable and comprehensive digital assistance.

Field
User-Centred Design
Source
arXiv preprint (2026)
Method
Empirical evaluation using a benchmark framework.
Sample
7 frontier AI models were evaluated.
Evidence
Strong effect

Current AI agents demonstrate significant limitations in autonomously completing common, multi-step online tasks across diverse platforms, indicating a gap between AI capabilities and user needs for practical assistance. This user-centred design research insight is drawn from a 2026 study published in arXiv preprint. Using Empirical evaluation using a benchmark framework. with 7 frontier AI models were evaluated., researchers explored how this design variable affects real-world outcomes. The key design takeaway: Designers should focus on developing AI agents that can handle the inherent complexity and variability of real-world online interactions, prioritizing user needs for reliable and comprehensive digital assistance.

Study
User-Centred DesignNew This WeekStrong effect

AI Agents Struggle with Real-World Online Tasks, Highlighting Need for User-Centred Design

Current AI agents demonstrate significant limitations in autonomously completing common, multi-step online tasks across diverse platforms, indicating a gap between AI capabilities and user needs for practical assistance.

arXiv preprint · 2026

01

Key Findings

  • 01AI agents can only complete a small fraction of everyday online tasks.
  • 02Current AI models struggle with tasks requiring information extraction from user documents, multi-step navigation across diverse platforms, and extensive form filling.
02

Application

Design takeaway

Designers should focus on developing AI agents that can handle the inherent complexity and variability of real-world online interactions, prioritizing user needs for reliable and comprehensive digital assistance.

How to apply

When designing AI-powered tools or assistants, rigorously test their performance on a wide range of realistic, multi-step tasks across different platforms, mimicking actual user behaviour and environmental complexities.

Project actions

  • 01When designing an AI assistant, think about all the different websites and steps a user might go through.
  • 02Test your AI on real websites, not just pretend ones, to see if it really works.
03

Method & Evidence

AimCan AI agents reliably complete a broad spectrum of everyday online tasks that require navigating multiple platforms and complex workflows?
MethodEmpirical evaluation using a benchmark framework.
ProcedureAn evaluation framework named ClawBench was developed, comprising 153 everyday online tasks across 144 live platforms. AI agents were tasked with completing these tasks, with only the final submission requests being intercepted to prevent real-world side effects. Performance was measured by the success rate of task completion.
Sample7 frontier AI models were evaluated.
ContextOnline task completion, AI agent capabilities, digital assistants.

Variables

IVType of AI agent, complexity of online task.
DVTask completion rate, accuracy of form filling, success in multi-step workflows.
CVLive production websites, specific task definitions, interception layer mechanism.
04

Strengths & Limitations

Strengths

  • +Utilizes live, production websites, reflecting real-world complexity.
  • +Covers a broad range of everyday tasks across multiple categories and platforms.

Limitations

The evaluation only blocked final submissions, so actual financial or personal data risks were avoided, but this might not capture all real-world consequences of AI errors. The definition of 'everyday tasks' can be subjective.

Reliability & validity

The use of live production websites enhances ecological validity. Reliability would depend on the consistency of AI agent performance across multiple trials for the same task. The benchmark's breadth contributes to construct validity.

Think critically

Given the current limitations, what are the most critical design considerations for developing AI agents that can effectively and safely assist users with complex online tasks?

05

Design Principles

"AI systems designed for user assistance must be evaluated and developed within the context of real-world, dynamic environments, not just isolated simulations."

This research underscores that while AI can handle isolated functions, its ability to act as a general-purpose assistant is still nascent. Designers and developers must focus on bridging this gap by creating AI systems that are more robust, adaptable, and aligned with the complexities of human workflows and user expectations for seamless digital interactions.

06

What This Means for Your Design

AI assistants aren't very good yet at doing everyday online jobs for you, like buying things or booking appointments, because the internet is complicated and changes a lot.

How to use in your project

  • 1.Use this research to justify the need for robust testing of AI agents in your design project, especially if you are developing or integrating AI functionalities.
  • 2.Cite this study when discussing the current limitations of AI in practical applications and the importance of user-centred evaluation.
07

Add to My Project

08

Quick Cite

Paragraph starter

Research indicates that current AI agents exhibit significant limitations in autonomously completing a wide array of everyday online tasks, often failing in scenarios requiring multi-platform navigation and complex data input. This highlights a critical gap between AI capabilities and the demands of real-world user workflows, underscoring the necessity for design approaches that prioritize robustness and user-centred evaluation in dynamic digital environments.

09

Source

arXiv preprint

ClawBench: Can AI Agents Complete Everyday Online Tasks?

journal · 2026

View source

Questions About This Research

What does the research say about ai agents struggle with real-world online tasks, highlighting need for user-centred design?
Designers should focus on developing AI agents that can handle the inherent complexity and variability of real-world online interactions, prioritizing user needs for reliable and comprehensive digital assistance. Evidence: arXiv preprint (2026).
Why does "AI Agents Struggle with Real-World Online Tasks, Highlighting Need for User-Centred Design" matter for design?
This research underscores that while AI can handle isolated functions, its ability to act as a general-purpose assistant is still nascent. Designers and developers must focus on bridging this gap by creating AI systems that are more robust, adaptable, and aligned with the complexities of human workflows and user expectations for seamless digital interactions.
How can designers apply this research?
Designers should focus on developing AI agents that can handle the inherent complexity and variability of real-world online interactions, prioritizing user needs for reliable and comprehensive digital assistance.
What were the main findings?
AI agents can only complete a small fraction of everyday online tasks.. Current AI models struggle with tasks requiring information extraction from user documents, multi-step navigation across diverse platforms, and extensive form filling.
What research method was used?
Empirical evaluation using a benchmark framework. with 7 frontier AI models were evaluated..
How strong is the evidence?
Evidence strength is rated Strong effect, based on a 2026 journal from arXiv preprint.
What should I do differently in my next project?
When designing AI-powered tools or assistants, rigorously test their performance on a wide range of realistic, multi-step tasks across different platforms, mimicking actual user behaviour and environmental complexities.
What are the limitations?
The evaluation framework intercepts final submission requests, which might not fully capture all potential failure points in a live transaction. The benchmark focuses on 'simple' tasks, and the complexity of 'everyday' tasks can vary significantly.