Compare Models
Select models (max 5)
Muse Spark 1.2GPT 5.5
| Vals IndexGDP-weighted benchmark | 49.29%±1.10 | N/A |
| Legal Research BenchAgentic US legal research | 43.75%±3.45 | 40.38%±3.41 |
| Finance Agent (v2)Core financial analyst tasks | 60.60%±0.28 | 51.76%±0.55 |
| Tax Agent BenchAgentic US corporate tax research | 56.86%±2.32 | 60.46%±3.22 |
| MedCodeMedical billing code support | 49.35%±2.19 | 49.10%±2.19 |
| Code MigrationRewriting programs in new languages | 29.95%±4.02 | 45.16%±4.16 |
| Terminal-Bench 4.0Frontier-difficulty terminal tasks | 6.06%±1.51 | N/A |
| Vibe Code Bench v1.1Building web apps from scratch | 79.10%±3.31 | 69.85%±4.54 |
Benchmarks
Vals Index *
Muse Spark 1.2
0.00%± 1.10
(43/43)GPT 5.5
N/ALegal Research Bench *
Muse Spark 1.2
0.00%± 3.45
(73/73)GPT 5.5
0.00%± 3.41
(73/73)Finance Agent (v2) *
Muse Spark 1.2
0.00%± 0.28
(74/74)GPT 5.5
0.00%± 0.55
(74/74)Tax Agent Bench *
Muse Spark 1.2
0.00%± 2.32
(65/65)GPT 5.5
0.00%± 3.22
(65/65)MedCode *
Muse Spark 1.2
0.00%± 2.19
(104/104)GPT 5.5
0.00%± 2.19
(104/104)Code Migration *
Muse Spark 1.2
0.00%