Short answer
When designing or selecting AI tools for analyzing public feedback, prioritize models that demonstrate equitable summarization across diverse demographic groups, especially socioeconomic status, and ensure the preservation of substantive arguments.
- Field
- User-Centred Design
- Source
- arXiv preprint (2026)
- Method
- Counterfactual experimental design
- Sample
- 182 public comments across 32 identity conditions, generating over 106,000 summaries.
- Evidence
- Strong effect (for socioeconomic status)
Large Language Models used by government agencies can produce biased summaries of public comments, disproportionately preserving less meaning and using simpler language for comments attributed to lower socioeconomic status individuals. This user-centred design research insight is drawn from a 2026 study published in arXiv preprint. Using Counterfactual experimental design with 182 public comments across 32 identity conditions, generating over 106,000 summaries., researchers explored how this design variable affects real-world outcomes. The key design takeaway: When designing or selecting AI tools for analyzing public feedback, prioritize models that demonstrate equitable summarization across diverse demographic groups, especially socioeconomic status, and ensure the preservation of substantive arguments.
LLM Summaries Exhibit Socioeconomic Bias, Undermining Equitable Public Input
Large Language Models used by government agencies can produce biased summaries of public comments, disproportionately preserving less meaning and using simpler language for comments attributed to lower socioeconomic status individuals.
arXiv preprint · 2026
Key Findings
- 01Comments attributed to individuals with lower socioeconomic status (e.g., 'street vendor') were summarized with less preservation of original meaning and simpler language compared to identical comments attributed to higher socioeconomic status individuals (e.g., 'financial analyst').
- 02This socioeconomic bias was consistent across different names, prompts, LLMs, and regulatory contexts.
- 03Race-based differential treatment was inconsistent and appeared linked to specific name tokens rather than racial categories.
- 04Gender-based differential treatment was absent.
- 05Writing quality of the original comment, specifically argument substance, influenced summarization outcomes more than superficial errors like spelling or grammar.
Application
Design takeaway
When designing or selecting AI tools for analyzing public feedback, prioritize models that demonstrate equitable summarization across diverse demographic groups, especially socioeconomic status, and ensure the preservation of substantive arguments.
How to apply
When developing or implementing AI tools for analyzing public comments, conduct thorough bias audits focusing on socioeconomic indicators and ensure that the AI's output accurately reflects the substance of the original input, regardless of the attributed commenter's background.
Project actions
- 01When testing AI tools, consider how different user characteristics might influence the AI's output.
- 02Focus on testing for fairness and bias in AI applications, especially those that impact public services or democratic processes.
Method & Evidence
Variables
Strengths & Limitations
Strengths
- +Employs a rigorous counterfactual design to isolate the impact of demographic attributions.
- +Tests a large number of comments and summaries across multiple LLMs, increasing generalizability.
Limitations
It's hard to test every possible demographic factor or every AI tool. The study focused on specific types of bias, and other biases might exist. Also, the study didn't look at how these biased summaries might affect real-world decisions.
Reliability & validity
The study's reliability is supported by the consistent findings across different names, prompts, and models for socioeconomic status. Validity is enhanced by the counterfactual design which isolates variables, though the specific metrics for 'meaning preservation' and 'language complexity' would need to be clearly defined and measured to ensure construct validity.
Think critically
If AI is used to summarize public input for policy decisions, how can designers ensure that the AI's output is a true and equitable representation of all voices, rather than a biased interpretation?
Design Principles
"AI systems processing user-generated content should be designed to ensure equitable representation and avoid bias based on demographic attributes, particularly those related to socioeconomic status."
This bias can lead to inequitable representation of public opinion in regulatory processes, potentially disenfranchising certain groups and undermining the democratic principle of notice-and-comment rulemaking. Designers of AI systems for public engagement must prioritize fairness and equity to ensure all voices are heard and considered.
What This Means for Your Design
Computers that read public comments for the government sometimes treat comments from poorer people differently than comments from richer people, even if the comments say the same thing. They might not capture the full meaning or use more basic words for comments from poorer people.
How to use in your project
- 1.Use this research to justify the importance of testing for bias in your own design project if it involves AI or user data.
- 2.Cite this study when discussing the potential for AI to create inequitable outcomes in your design analysis.
Add to My Project
Quick Cite
Paragraph starter
Research indicates that AI systems, such as Large Language Models used for processing public comments, can exhibit significant bias. A study by Kim et al. (2026) found that identical comments attributed to individuals of different socioeconomic statuses resulted in summaries that preserved less of the original meaning and used simpler language for those perceived as lower socioeconomic status. This highlights a critical design challenge: ensuring AI tools do not inadvertently disenfranchise certain user groups and that fairness benchmarks are integrated into AI development and procurement.
Source
arXiv preprint
All Public Voices Are Equal, But Are Some More Equal Than Others to LLMs?
journal · 2026
View sourceQuestions About This Research
- What does the research say about llm summaries exhibit socioeconomic bias, undermining equitable public input?
- When designing or selecting AI tools for analyzing public feedback, prioritize models that demonstrate equitable summarization across diverse demographic groups, especially socioeconomic status, and ensure the preservation of substantive arguments. Evidence: arXiv preprint (2026).
- Why does "LLM Summaries Exhibit Socioeconomic Bias, Undermining Equitable Public Input" matter for design?
- This bias can lead to inequitable representation of public opinion in regulatory processes, potentially disenfranchising certain groups and undermining the democratic principle of notice-and-comment rulemaking. Designers of AI systems for public engagement must prioritize fairness and equity to ensure all voices are heard and considered.
- How can designers apply this research?
- When designing or selecting AI tools for analyzing public feedback, prioritize models that demonstrate equitable summarization across diverse demographic groups, especially socioeconomic status, and ensure the preservation of substantive arguments.
- What were the main findings?
- Comments attributed to individuals with lower socioeconomic status (e.g., 'street vendor') were summarized with less preservation of original meaning and simpler language compared to identical comments attributed to higher socioeconomic status individuals (e.g., 'financial analyst').. This socioeconomic bias was consistent across different names, prompts, LLMs, and regulatory contexts.. Race-based differential treatment was inconsistent and appeared linked to specific name tokens rather than racial categories.. Gender-based differential treatment was absent.
- What research method was used?
- Counterfactual experimental design with 182 public comments across 32 identity conditions, generating over 106,000 summaries..
- How strong is the evidence?
- Evidence strength is rated Strong effect (for socioeconomic status), based on a 2026 journal from arXiv preprint.
- What should I do differently in my next project?
- When developing or implementing AI tools for analyzing public comments, conduct thorough bias audits focusing on socioeconomic indicators and ensure that the AI's output accurately reflects the substance of the original input, regardless of the attributed commenter's background.
- What are the limitations?
- The study focused on specific demographic attributions and a limited set of LLMs; broader demographic categories and a wider range of AI models might yield different results. The long-term impact of such biased summaries on policy outcomes was not directly measured.