GPT 5.4 Mini vs Inkling: Benchmark Comparison

GPT 5.4 Mini has the higher score on 7 of 18 shared benchmarks; Inkling leads on 10.

The largest observed score gap is 28.76 pts on Vibe Code Bench v1.1 , where GPT 5.4 Mini leads.

Reported ±1 standard-error ranges overlap on 6 of 18 shared benchmarks with comparable uncertainty data. This is not a pairwise statistical significance test.

Shared benchmark results

Scores and reported standard errors come from the latest visible Vals benchmark version available for both models.

Benchmark GPT 5.4 Mini Inkling Gap Reported uncertainty
Vals Index 33.17% ±1.28 28.67% ±1.10 4.50 pts Reported ±1 SE ranges do not overlap
Harvey's Legal Agent Benchmark 0.00% ±0.00 2.08% ±0.83 2.08 pts Reported ±1 SE ranges do not overlap
Legal Research Bench 12.50% ±2.30 28.36% ±3.13 15.86 pts Reported ±1 SE ranges do not overlap
EMB 45.43% ±3.65 38.36% ±3.27 7.06 pts Reported ±1 SE ranges do not overlap
Finance Agent (v2) 45.36% ±0.45 46.60% ±0.93 1.24 pts Reported ±1 SE ranges overlap
MortgageTax 63.51% ±0.91 63.99% ±0.94 0.48 pts Reported ±1 SE ranges overlap
Tax Agent Bench 36.59% ±3.10 43.52% ±2.95 6.93 pts Reported ±1 SE ranges do not overlap
TaxEval v2 71.22% ±0.90 75.31% ±0.84 4.09 pts Reported ±1 SE ranges do not overlap
GPQA Diamond 83.08% ±2.46 87.12% ±1.78 4.04 pts Reported ±1 SE ranges overlap
MMLU Pro 84.55% ±0.36 86.30% ±0.34 1.74 pts Reported ±1 SE ranges do not overlap
MMMU Pro 79.25% ±0.97 78.50% ±0.99 0.75 pts Reported ±1 SE ranges overlap
SAGE 50.81% ±3.40 36.55% ±3.22 14.26 pts Reported ±1 SE ranges do not overlap
Code Migration 12.94% ±3.64 11.79% ±3.40 1.15 pts Reported ±1 SE ranges overlap
LiveCodeBench 81.47% ±1.09 85.52% ±1.00 4.05 pts Reported ±1 SE ranges do not overlap
ProgramBench 0.00% ±0.00 0.00% ±0.00 0.00 pts Reported ±1 SE ranges overlap
SWE-bench 73.00% ±1.99 77.60% ±1.87 4.60 pts Reported ±1 SE ranges do not overlap
Terminal-Bench 4.0 2.52% ±1.01 0.51% ±0.51 2.02 pts Reported ±1 SE ranges do not overlap
Vibe Code Bench v1.1 47.97% ±5.61 19.21% ±3.66 28.76 pts Reported ±1 SE ranges do not overlap

Performance by category

Category GPT 5.4 Mini average Inkling average
Legal 6.25% 15.22%
Finance 52.42% 53.55%
Academic 82.29% 83.97%
Education 50.81% 36.55%
Coding 36.32% 32.44%

Cost and latency

Cost per test appears only where the benchmark reports it for both models. Latency is the measured completion time for that benchmark.

Benchmark GPT 5.4 Mini cost Inkling cost GPT 5.4 Mini latency Inkling latency
Vals Index $1.45 $1.46 29m49s 20m06s
Harvey's Legal Agent Benchmark $1.00 $1.25 11m21s 18m24s
Legal Research Bench $1.86 $0.49 37m36s 14m29s
EMB $2.15 $1.67 43m21s 26m22s
Finance Agent (v2) $1.20 $1.00 35m54s 18m02s
MortgageTax N/A N/A 2m06s 69.11s
Tax Agent Bench $0.81 $0.35 16m46s 13m31s
TaxEval v2 N/A N/A 44.58s 3m18s
GPQA Diamond N/A N/A 63.11s 5m09s
MMLU Pro N/A N/A 20.00s 73.26s
MMMU Pro N/A N/A 50.81s 2m04s
SAGE N/A N/A 2m21s 3m18s
Code Migration $1.95 $3.01 28m53s 29m34s
LiveCodeBench N/A N/A 3m09s 4m18s
ProgramBench N/A N/A 53m45s 45m34s
SWE-bench $0.51 $0.59 5m26s 6m34s
Terminal-Bench 4.0 $1.43 $2.51 29m41s 22m52s
Vibe Code Bench v1.1 $1.19 $1.61 34m15s 18m25s

Results available only for GPT 5.4 Mini

None.

Results available only for Inkling

  • LegalBench
  • MedCode
  • MedScribe
  • ProofBench v1.1
  • IOI
  • SkillsBench
  • Vibe Code Bench 1-100
  • CyberBench v1.1
  • Public Benefits Bench v1.1
Model details GPT 5.4 Mini Model details Inkling