Evaluating Large Language Models in Theory of Mind Tasks

Proceedings of the National Academy of Sciences

October2024 Vol. 121 Issue 45

View Publication

Eleven large language models (LLMs) were assessed using 40 bespoke false-belief tasks, considered a gold standard in testing theory of mind (ToM) in humans. Each task included a false-belief scenario, three closely matched true-belief control scenarios, and the reversed versions of all four. An LLM had to solve all eight scenarios to solve a single task. Older models solved no tasks; Generative Pre-trained Transformer (GPT)-3-davinci-003 (from November 2022) and ChatGPT-3.5-turbo (from March 2023) solved 20% of the tasks; ChatGPT-4 (from June 2023) solved 75% of the tasks, matching the performance of 6-y-old children observed in past studies. We explore the potential interpretation of these results, including the intriguing possibility that ToM-like ability, previously considered unique to humans, may have emerged as an unintended by-product of LLMs’ improving language skills. Regardless of how we interpret these outcomes, they signify the advent of more powerful and socially skilled AI—with profound positive and negative implications.

About Stanford GSB

About the Experience

Full-Time Degree Programs

Non-Degree & Certificate Programs

Faculty

Faculty Research

Research Hub

Research Labs

Research Initiatives

Topics

Welcome, Alumni

Admission Events & Information Sessions

Evaluating Large Language Models in Theory of Mind Tasks

Xubo Cao