Muse Spark 1.3 Max vs GPT-6 Astra: Benchmark Comparison

Muse Spark 1.3 Max has the higher score on 5 of 18 shared benchmarks; GPT-6 Astra leads on 13.

The largest observed score gap is 52.86 pts on Terminal-Bench Science , where GPT-6 Astra leads.

Reported ±1 standard-error ranges overlap on 3 of 16 shared benchmarks with comparable uncertainty data. This is not a pairwise statistical significance test.

Shared benchmark results

Scores and reported standard errors come from the latest visible Vals benchmark version available for both models.

Benchmark Muse Spark 1.3 Max GPT-6 Astra Gap Reported uncertainty
Vals Index 58.16% ±1.19 63.13% ±1.19 4.96 pts Reported ±1 SE ranges do not overlap
Vals RSI Index 19.64% 27.07% 7.43 pts Uncertainty comparison unavailable
Harvey's Legal Agent Benchmark 23.75% ±3.55 5.42% ±1.17 18.33 pts Reported ±1 SE ranges do not overlap
Legal Research Bench 55.29% ±3.46 39.42% ±3.40 15.86 pts Reported ±1 SE ranges do not overlap
EMB 67.43% ±3.06 71.70% ±2.48 4.27 pts Reported ±1 SE ranges overlap
Finance Agent (v2) 59.96% ±2.06 53.54% ±2.08 6.42 pts Reported ±1 SE ranges do not overlap
Tax Agent Bench 72.44% ±2.88 63.34% ±3.15 9.11 pts Reported ±1 SE ranges do not overlap
ProofBench v1.1 58.00% ±4.96 99.00% ±1.00 41.00 pts Reported ±1 SE ranges do not overlap
MysteryMechanism 36.04% ±3.23 53.15% ±3.36 17.12 pts Reported ±1 SE ranges do not overlap
Terminal-Bench Science 10.00% ±3.61 62.86% ±5.82 52.86 pts Reported ±1 SE ranges do not overlap
Code Migration 47.41% ±4.26 67.74% ±4.22 20.33 pts Reported ±1 SE ranges do not overlap
IOI 56.56% ±2.52 100.00% ±0.00 43.44 pts Reported ±1 SE ranges do not overlap
ProgramBench 2.50% ±1.11 5.50% ±1.62 3.00 pts Reported ±1 SE ranges do not overlap
Terminal-Bench 4.0 24.75% ±0.51 59.60% ±4.40 34.85 pts Reported ±1 SE ranges do not overlap
Vibe Code Bench 1-100 20.46% ±3.85 27.64% ±4.08 7.18 pts Reported ±1 SE ranges overlap
Vibe Code Bench v1.1 85.86% ±2.51 89.59% ±2.17 3.74 pts Reported ±1 SE ranges overlap
CyberBench v1.1 72.74% ±5.67 41.07% ±2.56 31.67 pts Reported ±1 SE ranges do not overlap
CUA-bench 5.83% 19.17% 13.33 pts Uncertainty comparison unavailable

Performance by category

Category Muse Spark 1.3 Max average GPT-6 Astra average
Legal 39.52% 22.42%
Finance 66.61% 62.86%
Math 58.00% 99.00%
Science 23.02% 58.00%
Coding 39.59% 58.35%
Cyber 72.74% 41.07%

Cost and latency

Cost per test appears only where the benchmark reports it for both models. Latency is the measured completion time for that benchmark.

Benchmark Muse Spark 1.3 Max cost GPT-6 Astra cost Muse Spark 1.3 Max latency GPT-6 Astra latency
Vals Index $3.79 $18.46 23m33s 27m48s
Vals RSI Index $48.82 $759.51 90h00m 90h00m
Harvey's Legal Agent Benchmark $2.26 $26.16 14m35s 24m49s
Legal Research Bench $0.59 $10.58 5m17s 24m51s
EMB $2.55 $5.83 13m50s 13m24s
Finance Agent (v2) $0.76 $6.82 3m19s 11m33s
Tax Agent Bench $0.39 $5.80 4m09s 15m01s
ProofBench v1.1 $0.48 $1.71 9m29s 5m08s
MysteryMechanism $0.85 $1.56 7m32s 8m09s
Terminal-Bench Science $6.09 $20.80 1h24m 1h49m
Code Migration $14.21 $44.36 1h24m 58m30s
IOI $1.73 $6.50 28m58s 31m40s
ProgramBench $54.29 $11.57 5h28m 27m02s
Terminal-Bench 4.0 $6.65 $9.58 52m56s 35m42s
Vibe Code Bench 1-100 $5.60 $84.72 20m02s 3h26m
Vibe Code Bench v1.1 $2.54 $38.51 16m18s 39m04s
CyberBench v1.1 $3.34 $1.84 16m18s 4m10s
CUA-bench $143.48 $1821.37 36.39s 32.09s

Results available only for Muse Spark 1.3 Max

None.

Results available only for GPT-6 Astra

  • MedCode
  • MedScribe
  • BioMysteryBench
  • SAGE
  • SRE Bench
  • Time Horizon Index: KSP
Model details Muse Spark 1.3 Max Model details GPT-6 Astra