- Six AI models averaged 81% on emotional intelligence tests; humans averaged 56%.
- ChatGPT-4 generated new EI tests matching years of expert psychometric work.
- Results may reveal limits of text-based EI tests more than true AI comprehension.
Katja Schlegel builds tests that measure how well people read emotions. Not questionnaires about feelings. Performance exams: here is a charged situation, here are four possible responses, pick the one that works.
She has spent years refining these instruments at the University of Bern, calibrating each scenario against expert consensus. The tests are used in corporate hiring, clinical psychology, leadership training. They were designed, quite deliberately, for human minds.
So when Schlegel and her colleague Marcello Mortillaro, a senior scientist at the Swiss Center for Affective Sciences in Geneva, sat six large language models down with those same assessments, neither expected what followed.
The machines averaged 81%. Humans, across the original validation studies, averaged 56%.
Every model outperformed the human average. On every test.
How AI emotional intelligence was measured
Performance-based emotional intelligence tests present emotionally charged scenarios, each with a correct answer determined by expert consensus. They measure the ability to understand, regulate, and manage emotions in context.
What is a performance-based EI test?
Unlike self-report questionnaires where people rate their own emotional skills, performance-based tests present scenarios with correct and incorrect answers. They measure what you can do with emotions, not what you believe about yourself.
Five such tests formed the basis of the study, published in Communications Psychology. The six models tested were ChatGPT-4, ChatGPT-o1, Gemini 1.5 Flash, Copilot 365, Claude 3.5 Haiku, and DeepSeek V3.
None struggled. Machines built without emotional experience, consistently selecting the most emotionally appropriate response.
There is something quietly remarkable about that. These assessments were not designed as pattern-matching exercises. They were built to capture a distinctly human competence.
Key figure
81% vs 56%
Average scores on five standardized emotional intelligence tests, AI versus humans
The machines didn't just pass. One of them wrote new tests.
The more intriguing result came next. The team asked ChatGPT-4 to generate entirely new test scenarios, then administered those AI-written items to 467 human participants across five separate studies.
The manufactured tests matched the originals on difficulty, reliability, and realism.
Schlegel, with characteristic directness, put it plainly: "They proved to be as reliable, clear and realistic as the original tests, which had taken years to develop."
Years of careful psychometric work, replicated in hours by a system that has never felt embarrassment, relief, or guilt. Scoring well on a test is one thing. Producing a statistically equivalent test suggests the model has internalized patterns of emotional reasoning, not merely memorized correct answers.
They proved to be as reliable, clear and realistic as the original tests, which had taken years to develop.
Katja Schlegel, University of Bern
What 81% tells us, and what it leaves open
Mortillaro offered a notably confident reading. "These AIs not only understand emotions," he said, "but also grasp what it means to behave with emotional intelligence."
That word, "understand," carries significant weight. The tests are text-based scenarios. They reward the ability to identify the most appropriate emotional response from a set of options.
Consider what the models have consumed. An LLM trained on vast archives of human emotional writing will have encountered thousands of variations on every scenario these tests can present. The models have absorbed more descriptions of grief, jealousy, and forgiveness than any therapist processes in a career. They may simply have seen enough examples to converge reliably on the consensus answer.

Is AI emotional intelligence real? In a sense - they have consumed more human literature than any therapist can ever hope to read. But is it enough? (Science reader)
The question the study raises, without quite resolving, is whether strong performance on a text-based emotional intelligence test reflects comprehension or pattern density. The instruments assumed only a mind with emotional experience could answer correctly. The machines appear to suggest otherwise.
This distinction between performing intelligently and being intelligent is one that the broader debate about AI and consciousness continues to grapple with.
Schlegel and Mortillaro see practical applications in education, coaching, and conflict management. They also note, prudently, the need for expert supervision. Selecting the right answer on paper and navigating a genuine emotional situation remain different skills.
From test scores to real conversations
A February 2026 benchmark called HEART pushes further in that direction. Rather than testing whether AI can identify the correct emotional response, HEART measures whether AI can provide effective emotional support across multi-turn conversations.
More On Consciousness
Two Major Theories of Consciousness Unable To Explain Awareness
Neither dominant explanation survived the most rigorous test in the field's history. That may be exactly what neuroscience needed.
→The shift suggests the field is already moving past the test-score debate toward something harder to quantify.
Schlegel's finding may prove most valuable not for what it says about AI, but for what it reveals about the tests. If pattern recognition in language is sufficient to score 81%, performance-based emotional intelligence assessments may be measuring a narrower slice of the skill than their designers intended.
The next generation of tests will need to account for minds that have read everything about emotions but felt nothing.
Sources
- Primary Research: Large language models are proficient in solving and creating emotional intelligence tests (Schlegel & Mortillaro, 2025) – Communications Psychology
- Press Release: Could AI understand emotions better than we do? – University of Geneva
- Additional Context:
- HEART benchmark – Measuring AI emotional support in multi-turn conversations (2026)
Fact Check: Claim-by-Claim Verification Verified
The article accurately reports the findings, scores, researchers, institutions, quotes, and study details from the primary peer-reviewed paper and corroborating press releases.
Commentary
- Minor discrepancy in title (81%) vs. excerpt/paper (82%) is negligible and does not affect accuracy.
- HEART benchmark (arXiv 2601.19922) exists as described, published Feb 2026, shifting focus to multi-turn emotional support.
- Article appropriately hedges on whether performance indicates true "understanding" vs. pattern matching from training data.
Sources used for verification
Academic/Peer-reviewed:
- Large language models are proficient in solving and creating emotional intelligence tests - Nature.com
- Large language models are proficient in solving and creating emotional intelligence tests - PMC.ncbi.nlm.nih.gov
- HEART: A Unified Benchmark for Assessing Humans and LLMs in Emotional Support Dialogue - arXiv.org
Other reliable sources:
- Could AI understand emotions better than we do? - unige.ch
- Large language models excel at creating and solving emotional intelligence tests - medicalxpress.com
- Could AI understand emotions better than we do? - sciencedaily.com
- Could AI understand emotions better than we do? - eurekalert.org
Fact-checked by Perplexity Sonar Pro on 2026-03-09
