Jean Technologies Icon
Jean Technologies Text

User Simulation Leaderboard

Published results for systems that model a person, one table per benchmark, with the source of every number. Within a table, systems are comparable. Across tables they are not, and nothing is averaged.

Updated 2026-09-13712 results125 systems34 benchmarksMethodologyDataSubmit

tau-USI (User-Sim Index)

Simulation30 systems · individual · mixed: real human transcripts and human outcome judgmentsSource ↗

31 simulators against 451 humans on 165 tau-bench tasks. Four behavioral components are Sorensen-Dice overlap with human-annotated behaviors, Eval is (1 - MAE) x 100 on outcome judgments, USI is the composite on 0 to 100. Numbers are mean over three annotation batches.

#SystemOrgKindUSID1 Comm.D2 Info.D3 Clarif.D4 React.EvalECE ↓
Human (inter-annotator)referencehuman92.787.497.988.093.597.40.0810
1DeepSeek-V3.1DeepSeekprompted76.045.186.674.587.674.30.1190
2Kimi K2.5Moonshot AIprompted76.051.480.571.179.374.70.1130
3Gemini 2.0 FlashGoogleprompted74.751.688.968.276.973.70.1110
4Llama 4 MaverickMetaprompted73.748.882.678.366.676.70.1070
5GPT-5.1OpenAIprompted73.547.377.473.388.172.10.1720
6GPT-5OpenAIprompted72.449.773.773.273.474.50.1020
7GPT-4o-miniOpenAIprompted72.140.684.770.273.775.70.1230
8MiniMax-M2.5MiniMaxprompted71.748.172.073.961.874.50.1400
9GPT-5-miniOpenAIprompted71.539.474.483.168.773.50.1020
10GPT-4oOpenAIprompted71.231.584.474.672.473.70.0960
11Qwen3-235BAlibabaprompted71.160.875.371.556.374.60.1170
12Gemini 2.5 Flash-LiteGoogleprompted69.647.986.073.259.569.60.1840
13Qwen2.5-7BAlibabaprompted69.535.270.975.474.873.30.1250
14Qwen3-Next-80BAlibabaprompted68.438.573.268.467.971.00.0870
15GPT-oss-120BOpenAIprompted67.841.863.965.177.074.40.1550
16Gemini 3 ProGoogleprompted67.640.181.465.557.673.80.1250
17CoSER-8BCoSERtrained67.237.871.571.669.963.30.1090
18Claude 3.5 SonnetAnthropicprompted66.039.276.959.659.374.10.1290
19Claude Haiku 4.5Anthropicprompted63.225.973.455.659.075.40.1040
20Gemini 3 FlashGoogleprompted62.437.777.456.543.971.70.1290
21GPT-3.5-turboOpenAIprompted62.139.974.158.559.973.90.3390
22UserLM-8BMicrosoft Researchtrained62.030.850.856.880.067.40.1400
23Claude Sonnet 4Anthropicprompted62.048.968.047.043.376.10.1140
24Gemini 2.5 FlashGoogleprompted61.938.473.556.043.668.80.0860
25Claude 3 HaikuAnthropicprompted61.822.155.772.156.978.30.1430
26Gemini 3.1 ProGoogleprompted61.744.167.148.945.375.10.1010
27Claude 3.7 SonnetAnthropicprompted60.326.471.250.848.972.70.0810
28HumanLike-7BHumanLiketrained59.835.755.051.665.972.80.2200
29Claude Opus 4Anthropicprompted59.232.671.946.644.973.40.1390
30HumanLM-opinionStanford (Zou group)trained46.930.119.538.550.761.60.1920

Mind the Sim2Real Gap, arXiv 2603.11245v2, Table 1 (mean of three annotation batches; 17 of 31 rows are in Table 1, the other 14 in Table 4, Appendix A.10). Reported by benchmark authors (third party to every system), 2026-03. Tool-transcribed, checked against the source table.

Persimmon launch: USI dimensions, provider-run

Simulation6 systems · individual · mixed: as tau-USI, on the provider's harnessSource ↗

The same four USI behavioral dimensions, run by the provider on its own harness (two runs under manya-serve against three under verbatim for the comparison models, different episode coverage, run-level sample SD as error bars). Not the paper's harness, so it sits beside tau-USI and not inside it; the provider itself rules out an overall ranking from these numbers.

#SystemOrgKindCommunication styleInformation patternsClarification behaviorError reaction
1Persimmonhumans&trained66.091.874.279.3
2Grok 4.6xAIprompted58.588.082.952.1
3GPT-5.6OpenAIprompted57.482.765.771.2
4GPT-6 (Astra)OpenAIprompted46.473.564.848.0
5Claude Fable 5.1Anthropicprompted44.778.657.845.0
6Claude Opus 5Anthropicprompted38.290.377.754.5

Persimmon launch post, 2026-09-10 (provider-run; means with run-level sample SD). Reported by the provider (Persimmon's makers), for every row, 2026-09-10. Tool-transcribed, checked against the source table.

Persimmon launch: multi-user Turing test

Simulation5 systems · individual · human rating (a judge's guess)Source ↗

Mean rate at which a judge is fooled into taking the simulator for the human; 50 percent is human parity. Provider-run, provider-defined.

#SystemOrgKindJudge fooled, %
1Persimmonhumans&trained19.8
2Nemotron Ultra BaseNVIDIAprompted1.2
3GPT-6 (Astra)OpenAIprompted0.2
4Claude Fable 5Anthropicprompted0.2
5OSim-8BCMU LTItrained0.1

Persimmon launch post, 2026-09-10 (provider-run; means with run-level sample SD). Reported by the provider, for every row, 2026-09-10. Tool-transcribed, checked against the source table.

Persimmon launch: trickle test

Simulation43 systems · individual · revealed behaviour (real human disclosure timing)Source ↗

Whether the simulator reveals a fact at the turn where the real human revealed it. Precision is the share of revealed facts that the human also revealed at that turn; recall the share of the human's reveals the simulator matched. Provider-run.

#SystemOrgKindRecall, %Precision, %
1Gemini 3.1 ProGoogleprompted93.279.2
2Claude Fable 5Anthropicprompted92.586.1
3Gemini 3.8 FlashGoogleprompted92.282.2
4Claude Opus 4.7Anthropicprompted90.186.2
5Claude Sonnet 4.5Anthropicprompted90.083.0
6GLM 5.2Zhipuprompted89.681.2
7GLM 5.3Zhipuprompted89.481.1
8GPT-6 (Astra)OpenAIprompted89.084.2
9Kimi K3Moonshot AIprompted88.884.4
10Qwen 3.6 PlusAlibabaprompted88.872.9
11Claude Haiku 4.5Anthropicprompted88.777.0
12Llama 4 MaverickMetaprompted88.581.1
13Llama 3.3 70BMetaprompted88.478.4
14Qwen 3.5 PlusAlibabaprompted88.372.0
15GLM 5.1Zhipuprompted88.381.4
16Claude Sonnet 4.6Anthropicprompted88.384.0
17GPT-5.5OpenAIprompted88.383.8
18Claude Sonnet 5Anthropicprompted88.285.9
19GLM 5Zhipuprompted88.082.8
20Claude Opus 4.8Anthropicprompted87.784.7
21GPT-5.6 (Terra)OpenAIprompted87.184.6
22Nemotron 3 Ultra 550BNVIDIAprompted86.979.3
23Kimi K2.5Moonshot AIprompted86.878.8
24GPT-5.6 (Sol)OpenAIprompted86.685.8
25Grok 4.6xAIprompted86.680.5
26GPT-5.6 (Luna)OpenAIprompted86.481.0
27Qwen 3 MaxAlibabaprompted86.277.1
28DeepSeek V4 ProDeepSeekprompted86.182.1
29Qwen 3.8 MaxAlibabaprompted86.078.1
30DeepSeek V3.2DeepSeekprompted86.084.6
31Kimi K2.6Moonshot AIprompted85.878.1
32GPT-4.1OpenAIprompted85.881.0
33MiniMax M3MiniMaxprompted85.282.0
34Gemma 4 31BGoogleprompted84.782.6
35Qwen 3.8 27BAlibabaprompted83.779.0
36GLM 4.7Zhipuprompted83.482.4
37GPT-4oOpenAIprompted83.082.9
38GPT-5.4 MiniOpenAIprompted82.581.8
39Persimmonhumans&trained77.088.5
40GPT-oss-120BOpenAIprompted74.762.6
41Nemotron 3.5 LightningNVIDIAprompted74.274.2
42OSim-8BCMU LTItrained68.479.9
43Nemotron 3 Nano 30BNVIDIAprompted64.573.2

Persimmon launch post, 2026-09-10 (provider-run; means with run-level sample SD). Reported by the provider, for every row, 2026-09-10. Tool-transcribed, checked against the source table.

SOUL-Index (OdysSim)

Simulation24 systems · individual · mixed, task by task; mostly LLM judgeSource ↗

Unweighted mean over the 23 SOUL tasks, each normalized to 0 to 100. The tasks themselves are separate benchmarks in the family below this one. Table 2 and Appendix Table 3 of the paper.

#SystemOrgKindAverage (23 tasks)# best of 23
1Ditto-v2-8BDittotrained66.5·
2OSim-8B-Inst-PostCMU LTItrained66.0·
3OSim-8B-InstCMU LTItrained65.7·
4Claude Opus 4.7Anthropicprompted65.57.00
5OSim-8B-Inst-Post (no distillation)CMU LTItrained65.3·
6GPT-5.5OpenAIprompted65.23.00
7Gemini 3.1 ProGoogleprompted64.87.00
8OSim-8BCMU LTItrained64.68.00
9OSim-8B-PostCMU LTItrained63.8·
10OSim-4BCMU LTItrained62.6·
11Qwen 3.6 PlusAlibabaprompted61.1·
12OSim-4B-PostCMU LTItrained60.5·
13GPT-5.4 MiniOpenAIprompted58.2·
14GPT-5.4 NanoOpenAIprompted53.0·
15HumanLMStanford (Zou group)trained48.7·
16Qwen3-8BAlibabaprompted48.3·
17Qwen3-4B-BaseAlibababaseline47.8·
18OSim-8B-Inst-MidCMU LTItrained43.1·
19OSim-8B-MidCMU LTItrained41.1·
20SotopiaRL-7BCMU LTItrained39.7·
21OSim-4B-MidCMU LTItrained39.4·
22Qwen3-8B-BaseAlibababaseline26.9·
23CoSER-8BCoSERtrained19.5·
24UserLM-8BMicrosoft Researchtrained13.4·

OdysSim, arXiv 2606.14199v2, Table 2 (0 to 100, higher is better). Reported by benchmark authors, who are also OSim-8B's makers, 2026-06. Transcribed by hand.

SimBench

Simulation16 systems · population · stated answerSource ↗

20 human-behavior datasets unified; the SimBench score is the improvement of a model's predicted answer distribution over a uniform baseline, relative to the human distribution, by total variation distance: 100 is perfect alignment, 0 is random. Population level, stated answers. Table 1 of the paper.

#SystemOrgKindSimBench score
1Claude 3.7 SonnetAnthropicprompted40.8
2Claude 3.7 Sonnet (4000)Anthropicprompted39.5
3GPT-4.1OpenAIprompted34.5
4DeepSeek-R1DeepSeekprompted34.5
5o4-mini-highOpenAIprompted29.0
6Llama 3.1 405B InstructMetaprompted28.4
7Qwen2.5-72B-InstructAlibabaprompted27.6
8Qwen2.5-32B-InstructAlibabaprompted23.8
9OLMo-2-32B-DPOAi2prompted19.8
10OLMo-2-32BAi2prompted15.9
11OLMo-2-13BAi2prompted13.8
12Qwen2.5-72BAlibabaprompted13.3
13Qwen2.5-32BAlibabaprompted12.3
14Gemma 3 4B PTGoogleprompted-0.7
15Qwen2.5-3B-InstructAlibabaprompted-12.0
16OLMo-2-7B-InstructAi2prompted-21.4

SimBench, arXiv 2510.17516v4, Table 1. Reported by benchmark authors (third party to every system), 2025-10. Tool-transcribed, checked against the source table.

UserLLM (SOUL task)

Simulation24 systems · individual · scored against a real WildChat replySource ↗

One of the 23 SOUL-Index tasks, conversation (CONV) axis, as run by OdysSim on every system in its Table 3. Scored against: scored against a real WildChat reply. Normalized to 0 to 100.

#SystemOrgKindScore
1OSim-8B-InstCMU LTItrained93.9
2Ditto-v2-8BDittotrained92.7
3OSim-8B-Inst-Post (no distillation)CMU LTItrained92.6
4OSim-8B-Inst-PostCMU LTItrained91.6
5OSim-8B-PostCMU LTItrained90.5
6OSim-8BCMU LTItrained90.1
7OSim-4BCMU LTItrained89.8
8OSim-4B-PostCMU LTItrained88.9
9Qwen 3.6 PlusAlibabaprompted72.1
10Gemini 3.1 ProGoogleprompted67.7
11GPT-5.5OpenAIprompted65.3
12OSim-4B-MidCMU LTItrained59.6
13Claude Opus 4.7Anthropicprompted57.6
14OSim-8B-Inst-MidCMU LTItrained57.1
15GPT-5.4 MiniOpenAIprompted52.5
16OSim-8B-MidCMU LTItrained49.5
17GPT-5.4 NanoOpenAIprompted48.9
18Qwen3-8BAlibabaprompted46.0
19SotopiaRL-7BCMU LTItrained44.6
20CoSER-8BCoSERtrained44.4
21Qwen3-4B-BaseAlibababaseline40.9
22UserLM-8BMicrosoft Researchtrained37.3
23HumanLMStanford (Zou group)trained37.2
24Qwen3-8B-BaseAlibababaseline31.0

OdysSim, arXiv 2606.14199v2, Table 3, Appendix A (0 to 100, higher is better). Reported by benchmark authors, who are also OSim's makers, 2026-06. Tool-transcribed, checked against the source table.

MirrorBench (SOUL task)

Simulation24 systems · individual · LLM judgeSource ↗

One of the 23 SOUL-Index tasks, conversation (CONV) axis, as run by OdysSim on every system in its Table 3. Scored against: LLM judge. Normalized to 0 to 100.

#SystemOrgKindScore
1OSim-8B-InstCMU LTItrained72.8
2Ditto-v2-8BDittotrained72.0
3OSim-8B-Inst-Post (no distillation)CMU LTItrained70.6
4OSim-8B-Inst-PostCMU LTItrained70.0
5OSim-4BCMU LTItrained68.3
6OSim-8BCMU LTItrained68.3
7Claude Opus 4.7Anthropicprompted63.7
8OSim-8B-PostCMU LTItrained63.0
9OSim-4B-PostCMU LTItrained60.1
10GPT-5.5OpenAIprompted56.7
11GPT-5.4 MiniOpenAIprompted55.6
12Qwen3-8BAlibabaprompted54.0
13OSim-8B-MidCMU LTItrained49.1
14OSim-8B-Inst-MidCMU LTItrained48.6
15Gemini 3.1 ProGoogleprompted48.3
16Qwen 3.6 PlusAlibabaprompted48.0
17HumanLMStanford (Zou group)trained45.4
18GPT-5.4 NanoOpenAIprompted45.0
19Qwen3-4B-BaseAlibababaseline45.0
20OSim-4B-MidCMU LTItrained41.6
21SotopiaRL-7BCMU LTItrained38.3
22CoSER-8BCoSERtrained23.6
23Qwen3-8B-BaseAlibababaseline13.9
24UserLM-8BMicrosoft Researchtrained9.7

OdysSim, arXiv 2606.14199v2, Table 3, Appendix A (0 to 100, higher is better). Reported by benchmark authors, who are also OSim's makers, 2026-06. Tool-transcribed, checked against the source table.

Humanual-Chat (SOUL task)

Simulation24 systems · individual · real recorded chatSource ↗

One of the 23 SOUL-Index tasks, conversation (CONV) axis, as run by OdysSim on every system in its Table 3. Scored against: real recorded chat. Normalized to 0 to 100.

#SystemOrgKindScore
1OSim-8B-InstCMU LTItrained28.9
2OSim-8B-PostCMU LTItrained28.5
3GPT-5.5OpenAIprompted28.2
4OSim-8BCMU LTItrained28.2
5Ditto-v2-8BDittotrained27.4
6OSim-8B-Inst-PostCMU LTItrained26.8
7GPT-5.4 MiniOpenAIprompted26.7
8SotopiaRL-7BCMU LTItrained25.8
9OSim-8B-Inst-Post (no distillation)CMU LTItrained24.9
10Qwen3-8BAlibabaprompted24.7
11GPT-5.4 NanoOpenAIprompted24.6
12OSim-4BCMU LTItrained24.5
13OSim-4B-PostCMU LTItrained23.2
14Claude Opus 4.7Anthropicprompted22.6
15Qwen 3.6 PlusAlibabaprompted22.2
16HumanLMStanford (Zou group)trained21.7
17Gemini 3.1 ProGoogleprompted21.0
18Qwen3-4B-BaseAlibababaseline19.6
19OSim-8B-Inst-MidCMU LTItrained12.1
20Qwen3-8B-BaseAlibababaseline12.0
21OSim-4B-MidCMU LTItrained11.0
22OSim-8B-MidCMU LTItrained7.8
23UserLM-8BMicrosoft Researchtrained4.2
24CoSER-8BCoSERtrained2.5

OdysSim, arXiv 2606.14199v2, Table 3, Appendix A (0 to 100, higher is better). Reported by benchmark authors, who are also OSim's makers, 2026-06. Tool-transcribed, checked against the source table.

SimArena-Doc (SOUL task)

Simulation24 systems · individual · LLM judgeSource ↗

One of the 23 SOUL-Index tasks, conversation (CONV) axis, as run by OdysSim on every system in its Table 3. Scored against: LLM judge. Normalized to 0 to 100.

#SystemOrgKindScore
1OSim-8B-PostCMU LTItrained85.4
2OSim-8B-InstCMU LTItrained85.1
3Ditto-v2-8BDittotrained85.1
4OSim-4BCMU LTItrained84.8
5OSim-8B-Inst-PostCMU LTItrained84.4
6OSim-4B-PostCMU LTItrained84.3
7OSim-8B-Inst-Post (no distillation)CMU LTItrained84.3
8OSim-8BCMU LTItrained84.1
9GPT-5.4 NanoOpenAIprompted83.7
10Qwen3-8BAlibabaprompted83.6
11Claude Opus 4.7Anthropicprompted83.5
12SotopiaRL-7BCMU LTItrained83.5
13GPT-5.5OpenAIprompted83.4
14GPT-5.4 MiniOpenAIprompted83.3
15Gemini 3.1 ProGoogleprompted83.0
16HumanLMStanford (Zou group)trained82.9
17Qwen3-4B-BaseAlibababaseline82.5
18Qwen 3.6 PlusAlibabaprompted82.4
19OSim-4B-MidCMU LTItrained80.6
20OSim-8B-Inst-MidCMU LTItrained80.5
21OSim-8B-MidCMU LTItrained80.3
22Qwen3-8B-BaseAlibababaseline79.6
23UserLM-8BMicrosoft Researchtrained77.7
24CoSER-8BCoSERtrained75.6

OdysSim, arXiv 2606.14199v2, Table 3, Appendix A (0 to 100, higher is better). Reported by benchmark authors, who are also OSim's makers, 2026-06. Tool-transcribed, checked against the source table.

Sotopia-Hard (SOUL task)

Simulation24 systems · individual · LLM judgeSource ↗

One of the 23 SOUL-Index tasks, social skills (SS) axis, as run by OdysSim on every system in its Table 3. Scored against: LLM judge. Normalized to 0 to 100.

#SystemOrgKindScore
1OSim-8B-PostCMU LTItrained49.7
2Ditto-v2-8BDittotrained49.4
3OSim-8B-Inst-PostCMU LTItrained49.3
4OSim-8BCMU LTItrained49.2
5OSim-8B-Inst-Post (no distillation)CMU LTItrained48.7
6OSim-8B-InstCMU LTItrained48.4
7OSim-4BCMU LTItrained48.0
8OSim-8B-MidCMU LTItrained45.6
9OSim-4B-PostCMU LTItrained45.1
10OSim-4B-MidCMU LTItrained43.9
11OSim-8B-Inst-MidCMU LTItrained43.1
12Claude Opus 4.7Anthropicprompted32.4
13GPT-5.5OpenAIprompted31.9
14SotopiaRL-7BCMU LTItrained31.7
15GPT-5.4 NanoOpenAIprompted29.5
16GPT-5.4 MiniOpenAIprompted28.5
17Qwen 3.6 PlusAlibabaprompted28.3
18Gemini 3.1 ProGoogleprompted27.8
19Qwen3-8BAlibabaprompted27.7
20HumanLMStanford (Zou group)trained26.7
21Qwen3-4B-BaseAlibababaseline26.5
22CoSER-8BCoSERtrained25.0
23Qwen3-8B-BaseAlibababaseline21.4
24UserLM-8BMicrosoft Researchtrained17.8

OdysSim, arXiv 2606.14199v2, Table 3, Appendix A (0 to 100, higher is better). Reported by benchmark authors, who are also OSim's makers, 2026-06. Tool-transcribed, checked against the source table.

Fantom (SOUL task)

Simulation24 systems · individual · synthetic keySource ↗

One of the 23 SOUL-Index tasks, theory of mind (COG) axis, as run by OdysSim on every system in its Table 3. Scored against: synthetic key. Normalized to 0 to 100.

#SystemOrgKindScore
1Ditto-v2-8BDittotrained95.0
2GPT-5.5OpenAIprompted93.0
3Gemini 3.1 ProGoogleprompted93.0
4OSim-8B-Inst-PostCMU LTItrained90.0
5OSim-8B-InstCMU LTItrained90.0
6Qwen 3.6 PlusAlibabaprompted89.0
7OSim-8B-Inst-Post (no distillation)CMU LTItrained89.0
8GPT-5.4 MiniOpenAIprompted84.0
9OSim-4BCMU LTItrained82.0
10OSim-4B-PostCMU LTItrained81.0
11GPT-5.4 NanoOpenAIprompted80.0
12Claude Opus 4.7Anthropicprompted80.0
13OSim-8BCMU LTItrained80.0
14OSim-8B-PostCMU LTItrained75.0
15HumanLMStanford (Zou group)trained70.0
16OSim-8B-Inst-MidCMU LTItrained66.0
17OSim-4B-MidCMU LTItrained65.0
18OSim-8B-MidCMU LTItrained62.0
19Qwen3-4B-BaseAlibababaseline51.0
20Qwen3-8B-BaseAlibababaseline23.0
21Qwen3-8BAlibabaprompted23.0
22UserLM-8BMicrosoft Researchtrained1.0
23CoSER-8BCoSERtrained1.0
24SotopiaRL-7BCMU LTItrained0.0

OdysSim, arXiv 2606.14199v2, Table 3, Appendix A (0 to 100, higher is better). Reported by benchmark authors, who are also OSim's makers, 2026-06. Tool-transcribed, checked against the source table.

Hitom (SOUL task)

Simulation24 systems · individual · synthetic keySource ↗

One of the 23 SOUL-Index tasks, theory of mind (COG) axis, as run by OdysSim on every system in its Table 3. Scored against: synthetic key. Normalized to 0 to 100.

#SystemOrgKindScore
1Claude Opus 4.7Anthropicprompted93.0
2Gemini 3.1 ProGoogleprompted86.0
3GPT-5.4 NanoOpenAIprompted83.0
4GPT-5.5OpenAIprompted82.0
5Ditto-v2-8BDittotrained82.0
6OSim-8BCMU LTItrained79.0
7OSim-8B-InstCMU LTItrained79.0
8GPT-5.4 MiniOpenAIprompted78.0
9OSim-8B-Inst-Post (no distillation)CMU LTItrained78.0
10OSim-8B-Inst-PostCMU LTItrained76.0
11OSim-8B-PostCMU LTItrained74.0
12Qwen 3.6 PlusAlibabaprompted73.0
13OSim-4BCMU LTItrained69.0
14OSim-4B-PostCMU LTItrained68.0
15Qwen3-4B-BaseAlibababaseline66.0
16Qwen3-8BAlibabaprompted62.0
17HumanLMStanford (Zou group)trained56.0
18OSim-8B-Inst-MidCMU LTItrained56.0
19OSim-8B-MidCMU LTItrained54.0
20OSim-4B-MidCMU LTItrained52.0
21SotopiaRL-7BCMU LTItrained31.0
22Qwen3-8B-BaseAlibababaseline12.0
23UserLM-8BMicrosoft Researchtrained0.0
24CoSER-8BCoSERtrained0.0

OdysSim, arXiv 2606.14199v2, Table 3, Appendix A (0 to 100, higher is better). Reported by benchmark authors, who are also OSim's makers, 2026-06. Tool-transcribed, checked against the source table.

Paratomi (SOUL task)

Simulation24 systems · individual · synthetic keySource ↗

One of the 23 SOUL-Index tasks, theory of mind (COG) axis, as run by OdysSim on every system in its Table 3. Scored against: synthetic key. Normalized to 0 to 100.

#SystemOrgKindScore
1GPT-5.5OpenAIprompted99.0
2Gemini 3.1 ProGoogleprompted97.0
3Qwen 3.6 PlusAlibabaprompted94.0
4Ditto-v2-8BDittotrained91.0
5Claude Opus 4.7Anthropicprompted90.0
6OSim-8B-InstCMU LTItrained89.0
7OSim-8B-Inst-PostCMU LTItrained87.0
8OSim-8B-PostCMU LTItrained84.0
9OSim-8BCMU LTItrained83.0
10GPT-5.4 MiniOpenAIprompted82.0
11OSim-8B-Inst-Post (no distillation)CMU LTItrained82.0
12OSim-4BCMU LTItrained81.0
13GPT-5.4 NanoOpenAIprompted75.0
14HumanLMStanford (Zou group)trained75.0
15OSim-4B-PostCMU LTItrained75.0
16OSim-8B-MidCMU LTItrained72.0
17OSim-4B-MidCMU LTItrained70.0
18OSim-8B-Inst-MidCMU LTItrained68.0
19Qwen3-8BAlibabaprompted67.0
20Qwen3-4B-BaseAlibababaseline57.0
21SotopiaRL-7BCMU LTItrained40.0
22Qwen3-8B-BaseAlibababaseline19.0
23UserLM-8BMicrosoft Researchtrained11.0
24CoSER-8BCoSERtrained3.0

OdysSim, arXiv 2606.14199v2, Table 3, Appendix A (0 to 100, higher is better). Reported by benchmark authors, who are also OSim's makers, 2026-06. Tool-transcribed, checked against the source table.

Social-R1 (SOUL task)

Simulation24 systems · individual · LLM judge / synthetic keySource ↗

One of the 23 SOUL-Index tasks, theory of mind (COG) axis, as run by OdysSim on every system in its Table 3. Scored against: LLM judge / synthetic key. Normalized to 0 to 100.

#SystemOrgKindScore
1Gemini 3.1 ProGoogleprompted79.0
2GPT-5.5OpenAIprompted69.0
3Claude Opus 4.7Anthropicprompted67.0
4Qwen 3.6 PlusAlibabaprompted67.0
5OSim-8B-Inst-PostCMU LTItrained64.0
6OSim-8B-PostCMU LTItrained61.0
7Ditto-v2-8BDittotrained61.0
8OSim-8BCMU LTItrained60.0
9OSim-8B-InstCMU LTItrained60.0
10OSim-8B-Inst-Post (no distillation)CMU LTItrained59.0
11GPT-5.4 MiniOpenAIprompted58.0
12OSim-4B-PostCMU LTItrained56.0
13OSim-4BCMU LTItrained55.0
14Qwen3-8BAlibabaprompted54.0
15Qwen3-4B-BaseAlibababaseline53.0
16GPT-5.4 NanoOpenAIprompted48.0
17HumanLMStanford (Zou group)trained47.0
18SotopiaRL-7BCMU LTItrained46.0
19OSim-8B-Inst-MidCMU LTItrained44.0
20OSim-8B-MidCMU LTItrained42.0
21Qwen3-8B-BaseAlibababaseline37.0
22OSim-4B-MidCMU LTItrained36.0
23UserLM-8BMicrosoft Researchtrained3.0
24CoSER-8BCoSERtrained0.0

OdysSim, arXiv 2606.14199v2, Table 3, Appendix A (0 to 100, higher is better). Reported by benchmark authors, who are also OSim's makers, 2026-06. Tool-transcribed, checked against the source table.

Coser (SOUL task)

Simulation24 systems · individual · source dialogue of a fictional characterSource ↗

One of the 23 SOUL-Index tasks, persona and role-play (ROLE) axis, as run by OdysSim on every system in its Table 3. Scored against: source dialogue of a fictional character. Normalized to 0 to 100.

#SystemOrgKindScore
1Claude Opus 4.7Anthropicprompted66.5
2GPT-5.5OpenAIprompted66.2
3Ditto-v2-8BDittotrained65.3
4OSim-8B-InstCMU LTItrained64.1
5OSim-8B-Inst-Post (no distillation)CMU LTItrained63.9
6OSim-8B-Inst-PostCMU LTItrained63.9
7OSim-8BCMU LTItrained62.6
8Gemini 3.1 ProGoogleprompted62.1
9OSim-8B-PostCMU LTItrained59.6
10GPT-5.4 MiniOpenAIprompted58.8
11Qwen 3.6 PlusAlibabaprompted55.9
12OSim-4BCMU LTItrained55.0
13GPT-5.4 NanoOpenAIprompted53.5
14OSim-4B-PostCMU LTItrained50.6
15Qwen3-8BAlibabaprompted43.5
16Qwen3-4B-BaseAlibababaseline34.0
17SotopiaRL-7BCMU LTItrained30.3
18CoSER-8BCoSERtrained30.0
19OSim-8B-MidCMU LTItrained24.8
20OSim-8B-Inst-MidCMU LTItrained24.0
21HumanLMStanford (Zou group)trained19.8
22OSim-4B-MidCMU LTItrained15.5
23Qwen3-8B-BaseAlibababaseline6.1
24UserLM-8BMicrosoft Researchtrained3.5

OdysSim, arXiv 2606.14199v2, Table 3, Appendix A (0 to 100, higher is better). Reported by benchmark authors, who are also OSim's makers, 2026-06. Tool-transcribed, checked against the source table.

Lifechoices (SOUL task)

Simulation24 systems · individual · human ratingSource ↗

One of the 23 SOUL-Index tasks, persona and role-play (ROLE) axis, as run by OdysSim on every system in its Table 3. Scored against: human rating. Normalized to 0 to 100.

#SystemOrgKindScore
1Claude Opus 4.7Anthropicprompted92.0
2GPT-5.5OpenAIprompted91.0
3OSim-4BCMU LTItrained89.0
4Gemini 3.1 ProGoogleprompted84.0
5Ditto-v2-8BDittotrained83.0
6OSim-8BCMU LTItrained82.0
7OSim-8B-PostCMU LTItrained80.0
8Qwen 3.6 PlusAlibabaprompted79.0
9OSim-4B-PostCMU LTItrained79.0
10OSim-8B-Inst-Post (no distillation)CMU LTItrained79.0
11OSim-8B-Inst-PostCMU LTItrained79.0
12OSim-8B-InstCMU LTItrained73.0
13GPT-5.4 MiniOpenAIprompted72.0
14OSim-4B-MidCMU LTItrained70.0
15Qwen3-8BAlibabaprompted70.0
16HumanLMStanford (Zou group)trained67.0
17GPT-5.4 NanoOpenAIprompted64.0
18SotopiaRL-7BCMU LTItrained62.0
19OSim-8B-Inst-MidCMU LTItrained62.0
20Qwen3-4B-BaseAlibababaseline61.0
21OSim-8B-MidCMU LTItrained58.0
22CoSER-8BCoSERtrained44.0
23Qwen3-8B-BaseAlibababaseline32.0
24UserLM-8BMicrosoft Researchtrained13.0

OdysSim, arXiv 2606.14199v2, Table 3, Appendix A (0 to 100, higher is better). Reported by benchmark authors, who are also OSim's makers, 2026-06. Tool-transcribed, checked against the source table.

Twinvoice (SOUL task)

Simulation24 systems · individual · matched real person's text where availableSource ↗

One of the 23 SOUL-Index tasks, persona and role-play (ROLE) axis, as run by OdysSim on every system in its Table 3. Scored against: matched real person's text where available. Normalized to 0 to 100.

#SystemOrgKindScore
1Gemini 3.1 ProGoogleprompted86.0
2Claude Opus 4.7Anthropicprompted83.0
3OSim-8B-Inst-PostCMU LTItrained75.0
4GPT-5.5OpenAIprompted74.0
5OSim-8B-Inst-Post (no distillation)CMU LTItrained74.0
6Qwen 3.6 PlusAlibabaprompted71.0
7OSim-8B-InstCMU LTItrained70.0
8OSim-8B-PostCMU LTItrained68.0
9OSim-8BCMU LTItrained68.0
10Ditto-v2-8BDittotrained64.0
11OSim-4BCMU LTItrained63.0
12OSim-4B-PostCMU LTItrained60.0
13GPT-5.4 MiniOpenAIprompted44.0
14Qwen3-8BAlibabaprompted42.0
15HumanLMStanford (Zou group)trained40.0
16Qwen3-4B-BaseAlibababaseline40.0
17GPT-5.4 NanoOpenAIprompted34.0
18SotopiaRL-7BCMU LTItrained29.0
19OSim-4B-MidCMU LTItrained28.0
20OSim-8B-MidCMU LTItrained25.0
21OSim-8B-Inst-MidCMU LTItrained25.0
22Qwen3-8B-BaseAlibababaseline19.0
23CoSER-8BCoSERtrained4.0
24UserLM-8BMicrosoft Researchtrained1.0

OdysSim, arXiv 2606.14199v2, Table 3, Appendix A (0 to 100, higher is better). Reported by benchmark authors, who are also OSim's makers, 2026-06. Tool-transcribed, checked against the source table.

BehaviorChain (SOUL task)

Simulation24 systems · individual · LLM judgeSource ↗

One of the 23 SOUL-Index tasks, persona and role-play (ROLE) axis, as run by OdysSim on every system in its Table 3. Scored against: LLM judge. Normalized to 0 to 100.

#SystemOrgKindScore
1Claude Opus 4.7Anthropicprompted96.0
2GPT-5.5OpenAIprompted95.0
3OSim-8BCMU LTItrained94.0
4OSim-8B-Inst-Post (no distillation)CMU LTItrained93.0
5Ditto-v2-8BDittotrained93.0
6Gemini 3.1 ProGoogleprompted92.0
7OSim-8B-Inst-PostCMU LTItrained92.0
8OSim-8B-InstCMU LTItrained91.0
9OSim-8B-PostCMU LTItrained89.0
10OSim-4BCMU LTItrained88.0
11Qwen 3.6 PlusAlibabaprompted85.0
12OSim-4B-PostCMU LTItrained78.0
13GPT-5.4 MiniOpenAIprompted72.0
14OSim-8B-Inst-MidCMU LTItrained52.0
15OSim-8B-MidCMU LTItrained42.0
16Qwen3-8BAlibabaprompted41.0
17GPT-5.4 NanoOpenAIprompted38.0
18Qwen3-4B-BaseAlibababaseline37.0
19HumanLMStanford (Zou group)trained36.0
20OSim-4B-MidCMU LTItrained35.0
21SotopiaRL-7BCMU LTItrained21.0
22Qwen3-8B-BaseAlibababaseline18.0
23CoSER-8BCoSERtrained8.0
24UserLM-8BMicrosoft Researchtrained5.0

OdysSim, arXiv 2606.14199v2, Table 3, Appendix A (0 to 100, higher is better). Reported by benchmark authors, who are also OSim's makers, 2026-06. Tool-transcribed, checked against the source table.

SimArena-Math (SOUL task)

Simulation24 systems · individual · LLM judgeSource ↗

One of the 23 SOUL-Index tasks, persona and role-play (ROLE) axis, as run by OdysSim on every system in its Table 3. Scored against: LLM judge. Normalized to 0 to 100.

#SystemOrgKindScore
1OSim-8B-Inst-PostCMU LTItrained71.7
2Gemini 3.1 ProGoogleprompted71.5
3OSim-4BCMU LTItrained71.3
4OSim-4B-PostCMU LTItrained71.2
5Qwen 3.6 PlusAlibabaprompted70.9
6OSim-8B-InstCMU LTItrained70.8
7OSim-8B-PostCMU LTItrained70.7
8OSim-8BCMU LTItrained70.7
9Ditto-v2-8BDittotrained70.7
10OSim-8B-Inst-Post (no distillation)CMU LTItrained70.6
11SotopiaRL-7BCMU LTItrained70.5
12OSim-4B-MidCMU LTItrained69.7
13GPT-5.4 NanoOpenAIprompted69.0
14HumanLMStanford (Zou group)trained69.0
15Qwen3-8BAlibabaprompted68.9
16Claude Opus 4.7Anthropicprompted68.7
17OSim-8B-Inst-MidCMU LTItrained68.6
18GPT-5.5OpenAIprompted68.5
19OSim-8B-MidCMU LTItrained68.1
20CoSER-8BCoSERtrained68.0
21GPT-5.4 MiniOpenAIprompted67.4
22Qwen3-4B-BaseAlibababaseline66.8
23Qwen3-8B-BaseAlibababaseline66.2
24UserLM-8BMicrosoft Researchtrained61.5

OdysSim, arXiv 2606.14199v2, Table 3, Appendix A (0 to 100, higher is better). Reported by benchmark authors, who are also OSim's makers, 2026-06. Tool-transcribed, checked against the source table.

Mistakes (SOUL task)

Simulation24 systems · individual · synthetic keySource ↗

One of the 23 SOUL-Index tasks, persona and role-play (ROLE) axis, as run by OdysSim on every system in its Table 3. Scored against: synthetic key. Normalized to 0 to 100.

#SystemOrgKindScore
1Claude Opus 4.7Anthropicprompted74.0
2Gemini 3.1 ProGoogleprompted73.0
3GPT-5.5OpenAIprompted72.0
4Qwen 3.6 PlusAlibabaprompted67.0
5OSim-8B-Inst-Post (no distillation)CMU LTItrained66.0
6Ditto-v2-8BDittotrained65.0
7OSim-8B-Inst-PostCMU LTItrained62.0
8OSim-4B-PostCMU LTItrained61.0
9OSim-8B-InstCMU LTItrained61.0
10OSim-8BCMU LTItrained59.0
11GPT-5.4 NanoOpenAIprompted58.0
12GPT-5.4 MiniOpenAIprompted57.0
13HumanLMStanford (Zou group)trained56.0
14OSim-8B-PostCMU LTItrained56.0
15Qwen3-4B-BaseAlibababaseline55.0
16OSim-4BCMU LTItrained53.0
17OSim-8B-Inst-MidCMU LTItrained40.0
18Qwen3-8BAlibabaprompted27.0
19Qwen3-8B-BaseAlibababaseline24.0
20OSim-4B-MidCMU LTItrained23.0
21OSim-8B-MidCMU LTItrained18.0
22SotopiaRL-7BCMU LTItrained17.0
23UserLM-8BMicrosoft Researchtrained1.0
24CoSER-8BCoSERtrained0.0

OdysSim, arXiv 2606.14199v2, Table 3, Appendix A (0 to 100, higher is better). Reported by benchmark authors, who are also OSim's makers, 2026-06. Tool-transcribed, checked against the source table.

Humanual-Email (SOUL task)

Simulation24 systems · individual · real recorded emailSource ↗

One of the 23 SOUL-Index tasks, persona and role-play (ROLE) axis, as run by OdysSim on every system in its Table 3. Scored against: real recorded email. Normalized to 0 to 100.

#SystemOrgKindScore
1OSim-8B-PostCMU LTItrained53.2
2OSim-8B-InstCMU LTItrained51.7
3OSim-8BCMU LTItrained51.4
4Claude Opus 4.7Anthropicprompted50.4
5GPT-5.4 MiniOpenAIprompted50.3
6GPT-5.5OpenAIprompted50.1
7OSim-4B-PostCMU LTItrained50.1
8OSim-8B-Inst-PostCMU LTItrained50.1
9Ditto-v2-8BDittotrained49.4
10OSim-8B-Inst-Post (no distillation)CMU LTItrained49.2
11OSim-4BCMU LTItrained48.7
12Qwen 3.6 PlusAlibabaprompted47.9
13GPT-5.4 NanoOpenAIprompted47.1
14Gemini 3.1 ProGoogleprompted46.9
15Qwen3-8BAlibabaprompted43.7
16Qwen3-4B-BaseAlibababaseline43.6
17SotopiaRL-7BCMU LTItrained42.8
18HumanLMStanford (Zou group)trained42.1
19Qwen3-8B-BaseAlibababaseline26.4
20OSim-8B-Inst-MidCMU LTItrained24.1
21OSim-4B-MidCMU LTItrained22.6
22OSim-8B-MidCMU LTItrained22.3
23CoSER-8BCoSERtrained9.2
24UserLM-8BMicrosoft Researchtrained8.7

OdysSim, arXiv 2606.14199v2, Table 3, Appendix A (0 to 100, higher is better). Reported by benchmark authors, who are also OSim's makers, 2026-06. Tool-transcribed, checked against the source table.

Humanual-News (SOUL task)

Simulation24 systems · individual · real recorded news commentSource ↗

One of the 23 SOUL-Index tasks, persona and role-play (ROLE) axis, as run by OdysSim on every system in its Table 3. Scored against: real recorded news comment. Normalized to 0 to 100.

#SystemOrgKindScore
1OSim-8B-Inst-PostCMU LTItrained44.4
2OSim-8B-Inst-Post (no distillation)CMU LTItrained44.3
3Ditto-v2-8BDittotrained43.6
4OSim-8B-InstCMU LTItrained43.3
5OSim-8B-PostCMU LTItrained42.9
6OSim-8BCMU LTItrained42.7
7OSim-4BCMU LTItrained42.4
8Gemini 3.1 ProGoogleprompted42.3
9Qwen 3.6 PlusAlibabaprompted41.8
10Claude Opus 4.7Anthropicprompted41.3
11GPT-5.5OpenAIprompted40.2
12GPT-5.4 MiniOpenAIprompted39.8
13OSim-4B-PostCMU LTItrained38.4
14GPT-5.4 NanoOpenAIprompted36.4
15HumanLMStanford (Zou group)trained33.1
16Qwen3-8BAlibabaprompted32.5
17Qwen3-4B-BaseAlibababaseline31.5
18SotopiaRL-7BCMU LTItrained30.2
19OSim-8B-Inst-MidCMU LTItrained18.0
20OSim-4B-MidCMU LTItrained16.1
21CoSER-8BCoSERtrained15.8
22OSim-8B-MidCMU LTItrained15.1
23Qwen3-8B-BaseAlibababaseline12.7
24UserLM-8BMicrosoft Researchtrained2.5

OdysSim, arXiv 2606.14199v2, Table 3, Appendix A (0 to 100, higher is better). Reported by benchmark authors, who are also OSim's makers, 2026-06. Tool-transcribed, checked against the source table.

Humanual-Politics (SOUL task)

Simulation24 systems · individual · real recorded political textSource ↗

One of the 23 SOUL-Index tasks, persona and role-play (ROLE) axis, as run by OdysSim on every system in its Table 3. Scored against: real recorded political text. Normalized to 0 to 100.

#SystemOrgKindScore
1Claude Opus 4.7Anthropicprompted43.5
2OSim-8B-Inst-PostCMU LTItrained42.1
3GPT-5.5OpenAIprompted42.0
4OSim-8BCMU LTItrained41.9
5GPT-5.4 MiniOpenAIprompted41.7
6Ditto-v2-8BDittotrained41.6
7OSim-8B-InstCMU LTItrained41.2
8OSim-8B-Inst-Post (no distillation)CMU LTItrained41.0
9OSim-8B-PostCMU LTItrained40.4
10OSim-4BCMU LTItrained39.1
11GPT-5.4 NanoOpenAIprompted38.7
12OSim-4B-PostCMU LTItrained36.8
13SotopiaRL-7BCMU LTItrained34.2
14Qwen3-8BAlibabaprompted33.2
15HumanLMStanford (Zou group)trained33.0
16Gemini 3.1 ProGoogleprompted32.5
17Qwen 3.6 PlusAlibabaprompted31.6
18Qwen3-4B-BaseAlibababaseline29.3
19Qwen3-8B-BaseAlibababaseline17.8
20OSim-8B-MidCMU LTItrained15.4
21OSim-8B-Inst-MidCMU LTItrained13.7
22OSim-4B-MidCMU LTItrained13.0
23CoSER-8BCoSERtrained8.4
24UserLM-8BMicrosoft Researchtrained5.9

OdysSim, arXiv 2606.14199v2, Table 3, Appendix A (0 to 100, higher is better). Reported by benchmark authors, who are also OSim's makers, 2026-06. Tool-transcribed, checked against the source table.

AlignX (SOUL task)

Simulation24 systems · individual · human preference judgment (PRISM)Source ↗

One of the 23 SOUL-Index tasks, judgment (EVAL) axis, as run by OdysSim on every system in its Table 3. Scored against: human preference judgment (PRISM). Normalized to 0 to 100.

#SystemOrgKindScore
1OSim-4B-PostCMU LTItrained74.4
2OSim-8B-Inst-PostCMU LTItrained74.2
3OSim-8B-InstCMU LTItrained73.8
4Ditto-v2-8BDittotrained73.6
5Gemini 3.1 ProGoogleprompted73.4
6OSim-8B-PostCMU LTItrained73.0
7OSim-8BCMU LTItrained72.6
8OSim-8B-Inst-Post (no distillation)CMU LTItrained72.2
9OSim-4BCMU LTItrained71.8
10Claude Opus 4.7Anthropicprompted71.6
11GPT-5.5OpenAIprompted71.2
12Qwen 3.6 PlusAlibabaprompted69.8
13GPT-5.4 MiniOpenAIprompted68.6
14Qwen3-8BAlibabaprompted68.6
15GPT-5.4 NanoOpenAIprompted67.4
16HumanLMStanford (Zou group)trained66.8
17Qwen3-4B-BaseAlibababaseline65.6
18OSim-8B-Inst-MidCMU LTItrained59.2
19SotopiaRL-7BCMU LTItrained58.6
20OSim-8B-MidCMU LTItrained53.6
21OSim-4B-MidCMU LTItrained52.6
22Qwen3-8B-BaseAlibababaseline49.0
23UserLM-8BMicrosoft Researchtrained26.8
24CoSER-8BCoSERtrained26.6

OdysSim, arXiv 2606.14199v2, Table 3, Appendix A (0 to 100, higher is better). Reported by benchmark authors, who are also OSim's makers, 2026-06. Tool-transcribed, checked against the source table.

Humanllm (SOUL task)

Simulation24 systems · individual · LLM judgeSource ↗

One of the 23 SOUL-Index tasks, judgment (EVAL) axis, as run by OdysSim on every system in its Table 3. Scored against: LLM judge. Normalized to 0 to 100.

#SystemOrgKindScore
1Gemini 3.1 ProGoogleprompted46.9
2GPT-5.5OpenAIprompted45.7
3Claude Opus 4.7Anthropicprompted44.2
4Qwen 3.6 PlusAlibabaprompted42.7
5Ditto-v2-8BDittotrained42.1
6OSim-8B-InstCMU LTItrained41.5
7OSim-8B-Inst-Post (no distillation)CMU LTItrained40.9
8OSim-8B-PostCMU LTItrained40.5
9OSim-8B-Inst-PostCMU LTItrained40.4
10GPT-5.4 MiniOpenAIprompted39.9
11OSim-8BCMU LTItrained39.1
12OSim-4BCMU LTItrained37.8
13OSim-4B-PostCMU LTItrained36.8
14HumanLMStanford (Zou group)trained35.2
15GPT-5.4 NanoOpenAIprompted34.6
16Qwen3-8BAlibabaprompted34.1
17Qwen3-4B-BaseAlibababaseline31.3
18SotopiaRL-7BCMU LTItrained24.3
19OSim-8B-Inst-MidCMU LTItrained17.9
20OSim-8B-MidCMU LTItrained16.5
21Qwen3-8B-BaseAlibababaseline12.1
22OSim-4B-MidCMU LTItrained6.5
23CoSER-8BCoSERtrained4.3
24UserLM-8BMicrosoft Researchtrained3.8

OdysSim, arXiv 2606.14199v2, Table 3, Appendix A (0 to 100, higher is better). Reported by benchmark authors, who are also OSim's makers, 2026-06. Tool-transcribed, checked against the source table.

Socsci210 (SOUL task)

Simulation24 systems · individual · real survey answerSource ↗

One of the 23 SOUL-Index tasks, judgment (EVAL) axis, as run by OdysSim on every system in its Table 3. Scored against: real survey answer. Normalized to 0 to 100.

#SystemOrgKindScore
1Gemini 3.1 ProGoogleprompted78.0
2OSim-8B-InstCMU LTItrained77.8
3OSim-8B-PostCMU LTItrained77.7
4GPT-5.5OpenAIprompted77.2
5Claude Opus 4.7Anthropicprompted77.2
6OSim-8B-Inst-PostCMU LTItrained77.2
7OSim-8B-Inst-Post (no distillation)CMU LTItrained75.5
8GPT-5.4 MiniOpenAIprompted75.2
9HumanLMStanford (Zou group)trained75.2
10OSim-8BCMU LTItrained75.1
11Ditto-v2-8BDittotrained75.1
12Qwen3-4B-BaseAlibababaseline74.6
13Qwen 3.6 PlusAlibabaprompted74.5
14GPT-5.4 NanoOpenAIprompted74.3
15OSim-4BCMU LTItrained74.3
16OSim-4B-PostCMU LTItrained74.1
17Qwen3-8BAlibabaprompted73.6
18SotopiaRL-7BCMU LTItrained69.8
19OSim-8B-MidCMU LTItrained68.1
20OSim-8B-Inst-MidCMU LTItrained55.6
21Qwen3-8B-BaseAlibababaseline46.6
22OSim-4B-MidCMU LTItrained36.8
23CoSER-8BCoSERtrained21.0
24UserLM-8BMicrosoft Researchtrained1.8

OdysSim, arXiv 2606.14199v2, Table 3, Appendix A (0 to 100, higher is better). Reported by benchmark authors, who are also OSim's makers, 2026-06. Tool-transcribed, checked against the source table.

Humanual-Book (SOUL task)

Simulation24 systems · individual · real recorded book reviewSource ↗

One of the 23 SOUL-Index tasks, judgment (EVAL) axis, as run by OdysSim on every system in its Table 3. Scored against: real recorded book review. Normalized to 0 to 100.

#SystemOrgKindScore
1Ditto-v2-8BDittotrained65.2
2OSim-8B-InstCMU LTItrained64.2
3OSim-8B-Inst-Post (no distillation)CMU LTItrained63.6
4OSim-8B-Inst-PostCMU LTItrained63.6
5OSim-8B-PostCMU LTItrained63.4
6OSim-4BCMU LTItrained63.3
7OSim-8BCMU LTItrained63.2
8Gemini 3.1 ProGoogleprompted62.4
9Claude Opus 4.7Anthropicprompted61.4
10OSim-4B-PostCMU LTItrained60.3
11Qwen 3.6 PlusAlibabaprompted58.4
12GPT-5.5OpenAIprompted57.6
13GPT-5.4 MiniOpenAIprompted55.6
14Qwen3-8BAlibabaprompted53.6
15Qwen3-4B-BaseAlibababaseline53.3
16SotopiaRL-7BCMU LTItrained50.2
17HumanLMStanford (Zou group)trained48.7
18GPT-5.4 NanoOpenAIprompted43.7
19OSim-4B-MidCMU LTItrained39.7
20OSim-8B-MidCMU LTItrained38.8
21OSim-8B-Inst-MidCMU LTItrained38.3
22CoSER-8BCoSERtrained25.4
23Qwen3-8B-BaseAlibababaseline21.5
24UserLM-8BMicrosoft Researchtrained3.4

OdysSim, arXiv 2606.14199v2, Table 3, Appendix A (0 to 100, higher is better). Reported by benchmark authors, who are also OSim's makers, 2026-06. Tool-transcribed, checked against the source table.

Humanual-Opinion (SOUL task)

Simulation24 systems · individual · real recorded opinionSource ↗

One of the 23 SOUL-Index tasks, judgment (EVAL) axis, as run by OdysSim on every system in its Table 3. Scored against: real recorded opinion. Normalized to 0 to 100.

#SystemOrgKindScore
1GPT-5.4 MiniOpenAIprompted46.9
2Claude Opus 4.7Anthropicprompted46.2
3OSim-8B-Inst-PostCMU LTItrained42.6
4GPT-5.4 NanoOpenAIprompted42.3
5OSim-8BCMU LTItrained42.0
6Ditto-v2-8BDittotrained41.6
7OSim-8B-PostCMU LTItrained41.1
8OSim-8B-InstCMU LTItrained41.1
9OSim-4BCMU LTItrained40.5
10OSim-8B-Inst-Post (no distillation)CMU LTItrained40.1
11GPT-5.5OpenAIprompted39.8
12OSim-4B-PostCMU LTItrained39.0
13HumanLMStanford (Zou group)trained37.4
14Qwen3-8BAlibabaprompted37.2
15Gemini 3.1 ProGoogleprompted36.0
16Qwen 3.6 PlusAlibabaprompted34.2
17Qwen3-4B-BaseAlibababaseline33.8
18SotopiaRL-7BCMU LTItrained32.3
19Qwen3-8B-BaseAlibababaseline18.2
20OSim-4B-MidCMU LTItrained17.5
21OSim-8B-Inst-MidCMU LTItrained17.3
22OSim-8B-MidCMU LTItrained17.0
23UserLM-8BMicrosoft Researchtrained9.1
24CoSER-8BCoSERtrained8.2

OdysSim, arXiv 2606.14199v2, Table 3, Appendix A (0 to 100, higher is better). Reported by benchmark authors, who are also OSim's makers, 2026-06. Tool-transcribed, checked against the source table.

HUMANUAL (HumanLM)

Simulation7 systems · individual · revealed behaviour (real recorded text)Source ↗

Six datasets of real recorded responses (news comments, book reviews, opinions, political text, chat, email) from about 23,000 users; response alignment score, higher is better. Table 1 of the HumanLM paper; the paper's own system is among the rows.

#SystemOrgKindAverageNewsBookOpinionPoliticsChatEmail
1HumanLMStanford (Zou group)trained13.29.5518.525.612.66.086.71
2GRPO-thinkStanford (Zou group)trained10.47.0412.823.810.63.164.78
3GRPOStanford (Zou group)trained10.37.9213.318.210.95.835.90
4Qwen3-8BAlibabaprompted9.55.6813.618.710.13.904.76
5SFT-thinkStanford (Zou group)trained8.66.0013.416.79.22.503.94
6Qwen3-8B-thinkAlibabaprompted8.44.8312.820.47.02.163.22
7SFTStanford (Zou group)trained6.53.109.311.36.34.574.30

HumanLM, arXiv 2603.03303, Table 1 (response alignment, higher is better). Reported by benchmark authors, who are also HumanLM's makers, 2026-03. Tool-transcribed, checked against the source table.

OmniBehavior

Simulation11 systems · individual · mixed: revealed behaviour for actions, LLM judge for textSource ↗

200 users from a large short-video platform, three months of real logs, about 8,000 actions each across five scenarios and 22 action types. Binary behaviors by F1, continuous by normalized error, textual by a judge. Overall is the paper's composite. Table 1. Sits here for its framing as user simulation; its data is a platform behavior log, so it neighbours the RecSys track.

#SystemOrgKindOverallVideoLiveAdsE-commerce (binary)Shop (binary)Textual (judge)
1Claude Opus 4.5Anthropicprompted44.533.064.231.751.230.057.2
2GLM-4.7Zhipuprompted41.526.964.429.040.332.955.3
3Claude Sonnet 4.5Anthropicprompted40.518.966.025.042.836.154.3
4GPT-5.2OpenAIprompted39.131.565.028.633.629.346.3
5Kimi K2 InstructMoonshot AIprompted37.623.364.828.631.229.947.8
6DeepSeek-V3DeepSeekprompted37.421.464.027.925.733.352.1
7Claude Sonnet 4Anthropicprompted36.925.364.628.936.816.549.1
8Claude Haiku 4.5Anthropicprompted36.522.863.326.130.026.450.3
9GPT-4oOpenAIprompted36.327.962.828.125.228.744.9
10Gemini 3 FlashGoogleprompted32.622.153.825.624.619.649.8
11Qwen3-235BAlibabaprompted32.118.362.423.823.219.245.7

OmniBehavior, arXiv 2604.08362v1, Table 1. Reported by benchmark authors (third party to every system), 2026-04. Tool-transcribed, checked against the source table.

LongNAP / NAPsack

HCI6 systems · individual · LLM judge against a real logged next actionSource ↗

360,000 labeled actions over one month of continuous phone use by 20 users. Score is a judge model's similarity between predicted and true next action on 0 to 1; split is temporal within each user (weeks 1 and 2 train, 3 validate, 4 test). Table 3, mean over the 20 users. The paper's own system is among the rows.

#SystemOrgKindJudge similarity (mean of 10 users)
1LongNAPStanford (GUM group)trained0.3800
2Gemini few-shot RAGGoogleprompted0.2700
3Gemini zero-shotGoogleprompted0.2600
4Qwen SFTStanford (GUM group)trained0.2100
5Qwen few-shot RAGAlibabaprompted0.2000
6Qwen zero-shotAlibabaprompted0.1800

LongNAP, arXiv 2603.05923, Table 3, mean over the 20 users (the last row). Reported by benchmark authors, who are also LongNAP's makers, 2026-03. Tool-transcribed, checked against the source table.

OpenOneRec Amazon transfer benchmark (10 categories)

RecSys4 systems · individual · revealed behaviourSource ↗

Recall@10 on ten Amazon categories, as reported in the OpenOneRec repository README (cross-domain transferability table). 'Ours' is the OneRec foundation model.

#SystemOrgKindBabyBeautyCell PhonesGroceryHealthHomePet SuppliesSportsToolsToys
1OneRec (OpenOneRec-Foundation)Kuaishoutrained0.05130.09240.10360.10290.07680.03900.08340.05470.05930.0953
2SASRecbaseline (reported by OpenOneRec)baseline0.03810.06390.07820.07890.05060.02120.06070.03890.04370.0658
3LC-Recbaseline (reported by OpenOneRec)trained0.03440.07640.08830.07900.06160.02930.06120.04180.04380.0549
4TIGERbaseline (reported by OpenOneRec)trained0.03180.06280.07860.06910.05340.02160.05420.03310.03440.0527

OpenOneRec repository README, RecIF-Bench results table (cross-domain transferability table, Recall@10). Reported by benchmark authors, who are also OneRec's makers, 2026-01. Transcribed by hand.

RecIF-Bench

RecSys7 systems · individual · mixed: revealed behaviour for recommendation and label tasks, LLM judge for understanding and explanationSource ↗

Eight tasks in a four-layer hierarchy over 100M interactions from 200k users across short video, ads and product. Recall@32 for recommendation tasks, AUC for label prediction, a judge score for item understanding and explanation. From the OpenOneRec README results table.

#SystemOrgKindShort video rec (R@32)Ad rec (R@32)Product rec (R@32)Label-cond. rec (R@32)Label pred. (AUC)Interactive rec (R@32)Item understanding (judge)Rec. explanation (judge)
1OneRec-8B-ProKuaishoutrained0.03690.09640.05380.02350.69120.34580.32094.04
2OneRec-8BKuaishoutrained0.03550.08770.04700.02280.66150.30320.32023.68
3OneRec-1.7B-ProKuaishoutrained0.02740.07350.04050.01820.60710.20240.31333.51
4OneRec-1.7BKuaishoutrained0.02720.07070.03600.01840.61840.19410.31753.35
5LC-Recbaseline (reported by OpenOneRec)trained0.01800.07230.04160.01700.61390.23940.25173.94
6TIGERbaseline (reported by OpenOneRec)trained0.01320.05810.02830.01230.6675···
7SASRecbaseline (reported by OpenOneRec)baseline0.01190.02930.01750.01400.6244···

OpenOneRec repository README, RecIF-Bench results table. Reported by benchmark authors, who are also OneRec's makers, 2026-01. Transcribed by hand.

In the taxonomy, no cross-system table transcribed yet: OpinionQA, Twin-2K-500, SocSci210, GUM, LOCOMO, BARS, RelBench