Jean Technologies Icon
Jean Technologies Text

Industry Evals

Measuring across the industry for simulation, HCI, and human models.

59 evals · updated 2026-09-22

MethodologyDataSubmit

Overview

Two of the leading benchmarks people have put forward for how well a model can simulate a human being. tau-USI scores how closely a model matches what one real person did in a conversation. SOUL-Index runs models across 23 tasks and is the most comprehensive to date.

Talks like a person

tau-USI

/ · published table

SystemUSI
Human (inter-annotator)Mar 202692.7
DeepSeek-V3.1best general modelMar 202676.0
OSim-8BOdysSimJun 202675.4
OSim-4B-MidOdysSimJun 202672.6
OSim-Inst-8BOdysSimJun 202671.4
CoSER-8BMar 202667.2
OSim-8B-MidOdysSimJun 202667.1
OSim-4BOdysSimJun 202666.8
OSim-Inst-4BOdysSimJun 202664.1
UserLM-8BMar 202662.0
HumanLike-7BMar 202659.8
HumanLM-opinionMar 202646.9

Behaves like a person across 23 tasks

SOUL-Index (OdysSim)

/ · published table

SystemAverage
Claude Opus 4.7best general modelJun 202665.5
Ditto-v2-8BJun 202666.5
OSim-8B-Inst-PostJun 202666.0
OSim-8B-InstJun 202665.7
OSim-8B-Inst-Post (no distillation)Jun 202665.3
OSim-8BJun 202664.6
OSim-8B-PostJun 202663.8
OSim-4BJun 202662.6
OSim-4B-PostJun 202660.5
HumanLMJun 202648.7
OSim-8B-Inst-MidJun 202643.1
OSim-8B-MidJun 202641.1
SotopiaRL-7BJun 202639.7

Run by us

SOUL verifiable subset

Eleven verifiable tasks from OdysSim’s SOUL-Index, rerun by Jean on the authors’ own harness and data. On the rows the paper publishes, our numbers are theirs.

Best score on each task, and who is second
Gemini 3.1 Pro79.0Gemini 3.8 Flash78.0
Gemini 3.1 Pro77.0Claude Fable 5.176.0
GPT-5.598.0Gemini 3.1 Pro97.0
Gemini 3.8 Flash78.6DeepSeek V4 Pro78.4
GPT-6 (Astra)86.0GPT-5.585.0
Ditto-v2-8B96.0Ditto-8B95.0
Gemini 3.1 Pro84.0Claude Opus 578.0
GPT-6 (Astra)92.0GLM 5.390.0
GPT-6 (Astra)77.6Gemini 3.1 Pro74.8
GPT-6 (Astra)48.3Gemini 3.1 Pro47.5
GPT-6 (Astra)100.0Ditto-v2-8B94.0

GPT-6 (Astra) 5, Gemini 3.1 Pro 3, GPT-5.5 1, Gemini 3.8 Flash 1, Ditto-v2-8B 1. No system is best at this, and there is no average over the 11 tasks because only GPT-6 (Astra) has a score on every one.

A social situation, then a question about it. Can the model read the room?

31 scored · 3 unreadable · line: the paper’s own OSim-8B score, 60.0

/

  • Gemini 3.1 Proreproduced +0.0 · parse 99% · 100 items · Sep 202679.0
  • Gemini 3.8 Flashprompted · parse 100% · 50 items · Sep 202678.0
  • Claude Opus 5prompted · parse 100% · 50 items · Sep 202678.0
  • Claude Fable 5.1prompted · parse 100% · 50 items · Sep 202678.0
  • GPT-6 (Astra)prompted · parse 100% · 50 items · Sep 202678.0
  • Grok 4.6prompted · parse 100% · 50 items · Sep 202674.0
  • GPT-5.5reproduced +2.0 · parse 100% · 100 items · Sep 202671.0
  • Qwen 3.8 Flashprompted · parse 100% · 100 items · Sep 202670.0
  • Claude Sonnet 5prompted · parse 100% · 50 items · Sep 202670.0
  • GLM 5.3 Flash (half-item arm)prompted · parse 100% · 50 items · Sep 202668.0
  • GLM 5.3prompted · parse 96% · 50 items · Sep 202668.0
  • DeepSeek Flashprompted · parse 100% · 50 items · Sep 202664.0
  • Kimi K2.6prompted · parse 96% · 50 items · Sep 202664.0
  • Ditto-v2-8Btrained · parse 100% · 100 items · Sep 202664.0
  • GLM 5.3 Flashprompted · parse 99% · 100 items · Sep 202663.0
  • OSim-8Breproduced +2.0 · parse 100% · 100 items · Sep 202662.0
  • Kimi K3prompted · parse 100% · 50 items · Sep 202662.0
  • OSim-4Btrained · parse 100% · 100 items · Sep 202662.0
  • OSim-Inst-4Btrained · parse 99% · 100 items · Sep 202662.0
  • OSim-Inst-8Btrained · parse 100% · 100 items · Sep 202661.0
0255075100

31 scored of 73 models in the registry ·

We reran someone else’s benchmark on their own code and data. Where they published a number for a model, ours matches theirs (Gemini 3.1 Pro within 4.0 points on 10; GPT-5.5 within 3.0 points on 5), so the chart is measuring what it claims to.

A bar is the share the model got right, out of 100. No bar means one of two things, and the line above the chart counts them separately: unreadable is the model answering in a form the benchmark’s own answer-reader could not parse, which counts against the model; not run is us never making the call, which counts against us.

There is no overall score, and that is deliberate. Only GPT-6 (Astra) has a score on all eleven tasks; every other system loses at least one to the unreadable rule or was not run on it, so an average across them does not exist for anybody else, and inventing one would hide the thing this page is about.