Industry Evals
Measuring across the industry for simulation, HCI, and human models.
59 evals · updated 2026-09-22
Overview
Two of the leading benchmarks people have put forward for how well a model can simulate a human being. tau-USI scores how closely a model matches what one real person did in a conversation. SOUL-Index runs models across 23 tasks and is the most comprehensive to date.
Talks like a person
tau-USI/ · published table
| System | USI | |
|---|---|---|
| Human (inter-annotator)Mar 2026 | 92.7 | |
| DeepSeek-V3.1best general modelMar 2026 | 76.0 | |
| OSim-8BOdysSimJun 2026 | 75.4 | |
| OSim-4B-MidOdysSimJun 2026 | 72.6 | |
| OSim-Inst-8BOdysSimJun 2026 | 71.4 | |
| CoSER-8BMar 2026 | 67.2 | |
| OSim-8B-MidOdysSimJun 2026 | 67.1 | |
| OSim-4BOdysSimJun 2026 | 66.8 | |
| OSim-Inst-4BOdysSimJun 2026 | 64.1 | |
| UserLM-8BMar 2026 | 62.0 | |
| HumanLike-7BMar 2026 | 59.8 | |
| HumanLM-opinionMar 2026 | 46.9 |
Behaves like a person across 23 tasks
SOUL-Index (OdysSim)/ · published table
| System | Average | |
|---|---|---|
| Claude Opus 4.7best general modelJun 2026 | 65.5 | |
| Ditto-v2-8BJun 2026 | 66.5 | |
| OSim-8B-Inst-PostJun 2026 | 66.0 | |
| OSim-8B-InstJun 2026 | 65.7 | |
| OSim-8B-Inst-Post (no distillation)Jun 2026 | 65.3 | |
| OSim-8BJun 2026 | 64.6 | |
| OSim-8B-PostJun 2026 | 63.8 | |
| OSim-4BJun 2026 | 62.6 | |
| OSim-4B-PostJun 2026 | 60.5 | |
| HumanLMJun 2026 | 48.7 | |
| OSim-8B-Inst-MidJun 2026 | 43.1 | |
| OSim-8B-MidJun 2026 | 41.1 | |
| SotopiaRL-7BJun 2026 | 39.7 |
Run by us
SOUL verifiable subsetEleven verifiable tasks from OdysSim’s SOUL-Index, rerun by Jean on the authors’ own harness and data. On the rows the paper publishes, our numbers are theirs.
| Gemini 3.1 Pro | 79.0 | Gemini 3.8 Flash | 78.0 | |
| Gemini 3.1 Pro | 77.0 | Claude Fable 5.1 | 76.0 | |
| GPT-5.5 | 98.0 | Gemini 3.1 Pro | 97.0 | |
| Gemini 3.8 Flash | 78.6 | DeepSeek V4 Pro | 78.4 | |
| GPT-6 (Astra) | 86.0 | GPT-5.5 | 85.0 | |
| Ditto-v2-8B | 96.0 | Ditto-8B | 95.0 | |
| Gemini 3.1 Pro | 84.0 | Claude Opus 5 | 78.0 | |
| GPT-6 (Astra) | 92.0 | GLM 5.3 | 90.0 | |
| GPT-6 (Astra) | 77.6 | Gemini 3.1 Pro | 74.8 | |
| GPT-6 (Astra) | 48.3 | Gemini 3.1 Pro | 47.5 | |
| GPT-6 (Astra) | 100.0 | Ditto-v2-8B | 94.0 |
GPT-6 (Astra) 5, Gemini 3.1 Pro 3, GPT-5.5 1, Gemini 3.8 Flash 1, Ditto-v2-8B 1. No system is best at this, and there is no average over the 11 tasks because only GPT-6 (Astra) has a score on every one.
A social situation, then a question about it. Can the model read the room?
31 scored · 3 unreadable · line: the paper’s own OSim-8B score, 60.0
/
- Gemini 3.1 Proreproduced +0.0 · parse 99% · 100 items · Sep 202679.0
- Gemini 3.8 Flashprompted · parse 100% · 50 items · Sep 202678.0
- Claude Opus 5prompted · parse 100% · 50 items · Sep 202678.0
- Claude Fable 5.1prompted · parse 100% · 50 items · Sep 202678.0
- GPT-6 (Astra)prompted · parse 100% · 50 items · Sep 202678.0
- Grok 4.6prompted · parse 100% · 50 items · Sep 202674.0
- GPT-5.5reproduced +2.0 · parse 100% · 100 items · Sep 202671.0
- Qwen 3.8 Flashprompted · parse 100% · 100 items · Sep 202670.0
- Claude Sonnet 5prompted · parse 100% · 50 items · Sep 202670.0
- GLM 5.3 Flash (half-item arm)prompted · parse 100% · 50 items · Sep 202668.0
- GLM 5.3prompted · parse 96% · 50 items · Sep 202668.0
- DeepSeek Flashprompted · parse 100% · 50 items · Sep 202664.0
- Kimi K2.6prompted · parse 96% · 50 items · Sep 202664.0
- Ditto-v2-8Btrained · parse 100% · 100 items · Sep 202664.0
- GLM 5.3 Flashprompted · parse 99% · 100 items · Sep 202663.0
- OSim-8Breproduced +2.0 · parse 100% · 100 items · Sep 202662.0
- Kimi K3prompted · parse 100% · 50 items · Sep 202662.0
- OSim-4Btrained · parse 100% · 100 items · Sep 202662.0
- OSim-Inst-4Btrained · parse 99% · 100 items · Sep 202662.0
- OSim-Inst-8Btrained · parse 100% · 100 items · Sep 202661.0
31 scored of 73 models in the registry ·
31 scored on this task, 3 asked and unreadable ·
We reran someone else’s benchmark on their own code and data. Where they published a number for a model, ours matches theirs (Gemini 3.1 Pro within 4.0 points on 10; GPT-5.5 within 3.0 points on 5), so the chart is measuring what it claims to.
A bar is the share the model got right, out of 100. No bar means one of two things, and the line above the chart counts them separately: unreadable is the model answering in a form the benchmark’s own answer-reader could not parse, which counts against the model; not run is us never making the call, which counts against us.
There is no overall score, and that is deliberate. Only GPT-6 (Astra) has a score on all eleven tasks; every other system loses at least one to the unreadable rule or was not run on it, so an average across them does not exist for anybody else, and inventing one would hide the thing this page is about.









