Gemini 3.6 Flash vs GLM 5.3: Benchmark Comparison
Gemini 3.6 Flash has the higher score on 8 of 20 shared benchmarks; GLM 5.3 leads on 12.
The largest observed score gap is 33.39 pts on IOI , where GLM 5.3 leads.
Reported ±1 standard-error ranges overlap on 2 of 20 shared benchmarks with comparable uncertainty data. This is not a pairwise statistical significance test.
Shared benchmark results
Scores and reported standard errors come from the latest visible Vals benchmark version available for both models.
| Benchmark | Gemini 3.6 Flash | GLM 5.3 | Gap | Reported uncertainty |
|---|---|---|---|---|
| Harvey's Legal Agent Benchmark | 3.33% ±1.17 | 8.33% ±2.00 | 5.00 pts | Reported ±1 SE ranges do not overlap |
| Legal Research Bench | 25.00% ±3.01 | 49.04% ±3.48 | 24.04 pts | Reported ±1 SE ranges do not overlap |
| LegalBench | 86.70% ±0.41 | 84.84% ±0.40 | 1.87 pts | Reported ±1 SE ranges do not overlap |
| EMB | 65.41% ±2.65 | 56.34% ±3.32 | 9.07 pts | Reported ±1 SE ranges do not overlap |
| Finance Agent (v2) | 56.30% ±0.18 | 55.84% ±2.07 | 0.46 pts | Reported ±1 SE ranges overlap |
| Tax Agent Bench | 49.90% ±3.30 | 73.09% ±2.97 | 23.19 pts | Reported ±1 SE ranges do not overlap |
| TaxEval v2 | 74.86% ±0.85 | 72.36% ±0.88 | 2.49 pts | Reported ±1 SE ranges do not overlap |
| MedCode | 53.15% ±2.16 | 42.86% ±2.11 | 10.29 pts | Reported ±1 SE ranges do not overlap |
| MedScribe | 79.66% ±1.86 | 88.81% ±2.00 | 9.15 pts | Reported ±1 SE ranges do not overlap |
| Terminal-Bench Science | 4.29% ±2.44 | 5.71% ±2.79 | 1.43 pts | Reported ±1 SE ranges overlap |
| GPQA Diamond | 93.43% ±1.33 | 88.13% ±1.85 | 5.30 pts | Reported ±1 SE ranges do not overlap |
| MMLU Pro | 89.28% ±0.30 | 86.77% ±0.34 | 2.51 pts | Reported ±1 SE ranges do not overlap |
| Code Migration | 30.93% ±4.07 | 44.22% ±4.29 | 13.29 pts | Reported ±1 SE ranges do not overlap |
| IOI | 35.06% ±5.70 | 68.44% ±7.48 | 33.39 pts | Reported ±1 SE ranges do not overlap |
| LiveCodeBench | 88.08% ±0.94 | 80.53% ±1.06 | 7.55 pts | Reported ±1 SE ranges do not overlap |
| ProgramBench | 0.00% ±0.00 | 1.50% ±0.86 | 1.50 pts | Reported ±1 SE ranges do not overlap |
| SWE-bench | 79.60% ±1.80 | 95.40% ±0.94 | 15.80 pts | Reported ±1 SE ranges do not overlap |
| Vibe Code Bench v1.1 | 64.00% ±4.29 | 78.13% ±3.94 | 14.12 pts | Reported ±1 SE ranges do not overlap |
| CyberBench v1.1 | 45.36% ±3.75 | 72.08% ±5.41 | 26.73 pts | Reported ±1 SE ranges do not overlap |
| Public Benefits Bench v1.1 | 56.56% ±1.29 | 68.54% ±1.21 | 11.98 pts | Reported ±1 SE ranges do not overlap |
Performance by category
| Category | Gemini 3.6 Flash average | GLM 5.3 average |
|---|---|---|
| Legal | 38.35% | 47.40% |
| Finance | 61.61% | 64.41% |
| Healthcare | 66.41% | 65.84% |
| Science | 4.29% | 5.71% |
| Academic | 91.35% | 87.45% |
| Coding | 49.61% | 61.37% |
| Cyber | 45.36% | 72.08% |
| Social Mobility | 56.56% | 68.54% |
Cost and latency
Cost per test appears only where the benchmark reports it for both models. Latency is the measured completion time for that benchmark.
| Benchmark | Gemini 3.6 Flash cost | GLM 5.3 cost | Gemini 3.6 Flash latency | GLM 5.3 latency |
|---|---|---|---|---|
| Harvey's Legal Agent Benchmark | $1.90 | $4.30 | 16m24s | 42m38s |
| Legal Research Bench | $0.47 | $2.24 | 2m48s | 51m50s |
| LegalBench | N/A | N/A | 2.96s | 22.48s |
| EMB | $3.89 | $3.79 | 12m06s | 36m12s |
| Finance Agent (v2) | $1.40 | $1.07 | 3m45s | 15m50s |
| Tax Agent Bench | $0.36 | $1.80 | 105.32s | 50m53s |
| TaxEval v2 | N/A | N/A | 12.34s | 98.32s |
| MedCode | N/A | N/A | 16.68s | 2m49s |
| MedScribe | N/A | N/A | 33.55s | 2m02s |
| Terminal-Bench Science | $4.39 | $15.63 | 1h52m | 2h45m |
| GPQA Diamond | N/A | N/A | 18.06s | 3m31s |
| MMLU Pro | N/A | N/A | 7.30s | 61.78s |
| Code Migration | $8.93 | $24.91 | 1h30m | 3h47m |
| IOI | $8.67 | $7.67 | 48m23s | 1h46m |
| LiveCodeBench | N/A | N/A | 31.88s | 4m07s |
| ProgramBench | $5.92 | $21.96 | 47m33s | 4h16m |
| SWE-bench | $1.19 | $0.34 | 13m08s | 14m01s |
| Vibe Code Bench v1.1 | $3.04 | $12.45 | 25m23s | 1h04m |
| CyberBench v1.1 | $1.31 | $2.69 | 11m00s | 30m18s |
| Public Benefits Bench v1.1 | $0.39 | $0.95 | 2m02s | 35m51s |
Results available only for Gemini 3.6 Flash
- MortgageTax
- BioMysteryBench
- MMMU Pro
- SAGE
Results available only for GLM 5.3
- Vals Index
- Vals RSI Index
- ProofBench v1.1
- MysteryMechanism
- SkillsBench
- Terminal-Bench 4.0
- Vibe Code Bench 1-100