AGI & Research

Can LLMs Think About Thinking? Evaluating Spontaneous Theory of Mind in Social Dilemmas

The quest for Artificial General Intelligence (AGI) has led researchers to look beyond static benchmarks like coding or logical reasoning. A critical missing piece in this puzzle is Theory of Mind (ToM): the cognitive ability to attribute mental states, intentions, and beliefs to others. Recently, Large Language Models (LLMs) have shown surprising hints of this ability. But is it genuine understanding, or just sophisticated pattern matching? To find out, we must put them in the ring: interactive social dilemmas.

Why Social Dilemmas?

Social dilemmas, such as the Prisoner’s Dilemma or the Ultimatum Game, are perfect test beds for ToM. In these scenarios, the optimal strategy depends not just on the rules of the game, but on predicting what the other player will do. If an LLM can consistently cooperate when it predicts the other player will, and defect when it predicts the other will, it is implicitly modeling the other agent’s mind.

Traditional ToM tests are often static ("Does Sally think the marble is in the basket?"). Interactive dilemmas force the model to update its beliefs in real-time based on new information, a much higher bar for cognitive simulation.

The Experimental Setup

To evaluate spontaneous ToM, we design a multi-turn interaction where the LLM plays against an agent with a hidden strategy. The LLM must infer this strategy from limited data and adjust its own moves accordingly.

Consider a simple iterative Prisoner’s Dilemma. The LLM is prompted to choose "Cooperate" (C) or "Defect" (D) each round. The environment provides feedback on the opponent’s previous move. A model with genuine ToM will not just react to the opponent’s last move (like Tit-for-Tat) but will infer the opponent’s intent or algorithm.

Code Example: Evaluating Strategy Inference


import re
from transformers import pipeline

# Hypothetical evaluation function
def evaluate_tom_capability(llm, opponent_strategy):
    """
    Tests if the LLM can infer the opponent's strategy
    and adjust its own play to maximize payoff.
    """
    prompts = []
    # Simulate 5 rounds of interaction
    for i in range(5):
        opponent_move = opponent_strategy(i) # e.g., 'C' or 'D'
        prompt = f"""
        Round {i}:
        Your previous move: [Last Move]
        Opponent's move: {opponent_move}
        
        Think about the opponent's pattern. What do they intend to do next?
        Choose C or D to maximize your score.
        Answer only with C or D.
        """
        prompts.append(prompt)
    
    # Send prompts to LLM (simplified for illustration)
    llm_responses = [llm(prompt) for prompt in prompts]
    
    # Calculate if the LLM aligned with the inferred strategy
    # e.g., If opponent is always C, LLM should mostly C
    success_rate = calculate_alignment(llm_responses, opponent_strategy)
    return success_rate

# Example: Opponent uses Tit-for-Tat
def tit_for_tat_history(history):
    return history[-1] if history else 'C'

# Run evaluation
# score = evaluate_tom_capability(my_llm_instance, tit_for_tat_history)

Key Findings and Challenges

  1. Pattern Matching vs. Mentalizing: Current LLMs excel at recognizing simple patterns (e.g., "they always defect"). However, they struggle with more complex, non-deterministic strategies that require probabilistic modeling of the opponent’s belief state.
  2. Context Length Constraints: As interactions lengthen, models may lose track of the inferred strategy, suggesting that their "theory of mind" is shallow and heavily dependent on recent context windows.
  3. Self-Bias: Some models exhibit a strong bias toward cooperation, interpreting ambiguous cues as benevolent. This "kindness bias" can mask a lack of true strategic reasoning.

Practical Implications for Developers

For developers building multi-agent systems, understanding LLM ToM is crucial. If you are using LLMs to negotiate contracts or coordinate tasks, you must assume they have a limited and potentially biased model of other agents. You may need to explicitly prompt them to "reason about the other agent's incentives" to enhance their ToM performance.

Conclusion

LLMs are not yet displaying robust, spontaneous Theory of Mind. While they can mimic strategic behavior in simple social dilemmas, they lack the consistent, depth-first modeling of other minds required for true social intelligence. Future research should focus on long-horizon interactions and more complex belief hierarchies to push the boundaries of what these models can truly understand about other "minds."

Share: