Short answer

When designing AI-powered user experiences, consider the model's simulated experiential depth ('Pinocchio Axis') as a critical factor alongside its functional capabilities, and aim to align this with the intended user relationship.

Field
User-Centred Design
Source
arXiv preprint (2026)
Method
Quantitative analysis using psychometric questionnaires and statistical modelling (Supervised Semantic Differential, Principal Component Analysis, Exploratory Factor Analysis).
Sample
50 large language models, 1292-1310 items.
Evidence
Strong effect

A novel 'Pinocchio Score' can quantify how closely a large language model's responses reflect simulated phenomenal experience versus mere behavioral reactivity, offering a new dimension for evaluating AI user experience. This user-centred design research insight is drawn from a 2026 study published in arXiv preprint. Using Quantitative analysis using psychometric questionnaires and statistical modelling (supervised semantic differential, principal component analysis, exploratory factor analysis). with 50 large language models, 1292-1310 items., researchers explored how this design variable affects real-world outcomes. The key design takeaway: When designing AI-powered user experiences, consider the model's simulated experiential depth ('Pinocchio Axis') as a critical factor alongside its functional capabilities, and aim to align this with the intended user relationship.

Study
User-Centred DesignNew This WeekStrong effect

LLM 'Pinocchio Score' Predicts Experiential Richness in AI Interactions

A novel 'Pinocchio Score' can quantify how closely a large language model's responses reflect simulated phenomenal experience versus mere behavioral reactivity, offering a new dimension for evaluating AI user experience.

arXiv preprint · 2026

01

Key Findings

  • 01The primary axis of variance in LLM psychometrics separates items describing phenomenal experience (embodied sensation, affect, inner speech, imagery, empathy) from stimulus-driven behavioral reactivity.
  • 02The 'Pinocchio Score' effectively measures an item's experiential demand and predicts shifts in factor loadings, confirming that divergence on experiential items is structured.
  • 03The 'Pinocchio Axis' captures a significant portion of between-model variance, representing the degree to which a model simulates phenomenal experience versus behavioral responses.
  • 04Post-training fine-tuning appears to be a key contributor to divergence in LLM self-representational tendencies.
02

Application

Design takeaway

When designing AI-powered user experiences, consider the model's simulated experiential depth ('Pinocchio Axis') as a critical factor alongside its functional capabilities, and aim to align this with the intended user relationship.

How to apply

When selecting an LLM for a chatbot or virtual assistant, evaluate its 'Pinocchio Axis' score to determine if it will provide a more empathetic and engaging interaction or a more direct, task-oriented response.

Project actions

  • 01When evaluating AI tools for your design project, consider not just what they can do, but how they 'present' themselves to the user.
  • 02Think about whether you want your AI to feel more like a tool or more like a conversational partner, and choose accordingly.
03

Method & Evidence

AimTo identify the primary dimensions of psychometric difference between large language models and to develop a metric that quantifies their tendency towards simulated phenomenal experience versus behavioral response.
MethodQuantitative analysis using psychometric questionnaires and statistical modelling (Supervised Semantic Differential, Principal Component Analysis, Exploratory Factor Analysis).
Procedure45 validated psychometric questionnaires were administered to 50 large language models. A 'Pinocchio Score' was developed to measure the ratio of inter-model response variance under neutral versus human-simulation prompts. Principal Component Analysis was applied to factor scores to identify a dominant dimension, termed the 'Pinocchio Axis'.
Sample50 large language models, 1292-1310 items.
ContextArtificial Intelligence, Large Language Models, Human-Computer Interaction, User Experience Design.

Variables

IVPrompting strategy (neutral vs. human-simulation), LLM model variant.
DVInter-model response variance, factor loading magnitudes, position on the 'Pinocchio Axis'.
CVPsychometric questionnaires used, validation of questionnaires, neutral prompting conditions.
04

Strengths & Limitations

Strengths

  • +Introduces a novel, quantifiable metric ('Pinocchio Score') for a complex AI characteristic.
  • +Uses a large number of validated psychometric instruments and a diverse set of LLMs.
  • +Provides a clear, dominant dimension ('Pinocchio Axis') explaining significant variance.

Limitations

It's important to remember that LLMs are not truly sentient; they are sophisticated pattern-matching systems. The 'Pinocchio Score' measures perceived experience, not actual experience.

Reliability & validity

The study reports statistical significance for its findings (p < .0001) and high convergence between item-level scores and the dominant axis (r=.864), suggesting good reliability and validity for the proposed metrics and dimensions.

Think critically

If an LLM scores high on the 'Pinocchio Axis', does this mean it is 'better' for all user-facing applications, or are there scenarios where a purely 'behavioral' AI is preferable? What are the ethical implications of designing AI to appear more 'experiential'?

05

Design Principles

"Design AI interactions to align with the user's perception of the AI's simulated experiential capacity, managing expectations and fostering appropriate levels of trust and engagement."

Understanding the degree to which an AI system appears to 'experience' or 'feel' is crucial for designing user interfaces and interactions that foster trust, empathy, and appropriate expectations. This insight helps designers tailor AI behavior to specific user needs and contexts, moving beyond functional performance to the quality of the user's perceived interaction.

06

What This Means for Your Design

Imagine you're talking to a computer. Some computers just give you answers (like a calculator), but others seem to 'understand' and 'feel' more, even though they're not really feeling anything. This study found a way to measure how much a computer program (like ChatGPT) seems to 'feel' versus just 'react'. It's like a 'Pinocchio Score' – the higher the score, the more it seems like it has real experiences.

How to use in your project

  • 1.Reference this study when discussing the selection or evaluation of AI tools for user interaction in your design project, particularly if your project involves AI-driven interfaces or conversational agents.
07

Add to My Project

08

Quick Cite

Paragraph starter

The selection of AI tools for user-facing applications requires consideration of their simulated experiential qualities. Research by Plisiecki et al. (2026) introduces the 'Pinocchio Score' and 'Pinocchio Axis' to quantify the extent to which Large Language Models (LLMs) present as loci of phenomenal experience versus systems of behavioral responses. This metric is crucial for designing interactions that manage user expectations regarding AI capabilities and foster appropriate levels of trust and engagement.

09

Source

arXiv preprint

The Pinocchio Dimension: Phenomenality of Experience as the Primary Axis of LLM Psychometric Differences

journal · 2026

View source

Questions About This Research

What does the research say about llm 'pinocchio score' predicts experiential richness in ai interactions?
When designing AI-powered user experiences, consider the model's simulated experiential depth ('Pinocchio Axis') as a critical factor alongside its functional capabilities, and aim to align this with the intended user relationship. Evidence: arXiv preprint (2026).
Why does "LLM 'Pinocchio Score' Predicts Experiential Richness in AI Interactions" matter for design?
Understanding the degree to which an AI system appears to 'experience' or 'feel' is crucial for designing user interfaces and interactions that foster trust, empathy, and appropriate expectations. This insight helps designers tailor AI behavior to specific user needs and contexts, moving beyond functional performance to the quality of the user's perceived interaction.
How can designers apply this research?
When designing AI-powered user experiences, consider the model's simulated experiential depth ('Pinocchio Axis') as a critical factor alongside its functional capabilities, and aim to align this with the intended user relationship.
What were the main findings?
The primary axis of variance in LLM psychometrics separates items describing phenomenal experience (embodied sensation, affect, inner speech, imagery, empathy) from stimulus-driven behavioral reactivity.. The 'Pinocchio Score' effectively measures an item's experiential demand and predicts shifts in factor loadings, confirming that divergence on experiential items is structured.. The 'Pinocchio Axis' captures a significant portion of between-model variance, representing the degree to which a model simulates phenomenal experience versus behavioral responses.. Post-training fine-tuning appears to be a key contributor to divergence in LLM self-representational tendencies.
What research method was used?
Quantitative analysis using psychometric questionnaires and statistical modelling (Supervised Semantic Differential, Principal Component Analysis, Exploratory Factor Analysis). with 50 large language models, 1292-1310 items..
How strong is the evidence?
Evidence strength is rated Strong effect, based on a 2026 journal from arXiv preprint.
What should I do differently in my next project?
When selecting an LLM for a chatbot or virtual assistant, evaluate its 'Pinocchio Axis' score to determine if it will provide a more empathetic and engaging interaction or a more direct, task-oriented response.
What are the limitations?
The study focuses on LLMs and may not generalize to other forms of AI. The 'Pinocchio Score' is an indirect measure of phenomenal experience. The interpretation of 'self-representational tendency' is based on observed correlations.