LLM-Simulated Conversations Underplay Human Miscommunication and Conflict
The paper argues that evaluating LLM-simulated human conversations should account for the fact that real human dialogue is full of misunderstandings, interruptions, and uncooperative moments—not just coherent, cooperative exchanges. The authors introduce CoCoEval, a framework that detects 10 types of inconsistent or uncollaborative behaviors turn-by-turn in simulated professional-scenario conversations. Testing GPT-4.1, GPT-5.1, and Claude Opus 4 against human transcripts, they find LLM-simulated conversations show far fewer such behaviors than real ones under standard prompting, and that prompt engineering or fine-tuning fails to reliably fix this—often overcorrecting into specific artificial behaviors instead. The authors say these gaps are invisible to conventional conversation-level Likert-scale evaluations, raising concerns about using LLMs as stand-ins for human social interaction in research.
Discussion: 2 tweets from 2 authors · @RyoKamoi, @RyoKamoi_ja