Gemini 3.8 Flash vs GPT-6 Astra: Benchmark Comparison

Gemini 3.8 Flash has the higher score on 5 of 23 shared benchmarks; GPT-6 Astra leads on 18.

The largest observed score gap is 71.67 pts on Time Horizon Index: KSP , where GPT-6 Astra leads.

Reported ±1 standard-error ranges overlap on 6 of 21 shared benchmarks with comparable uncertainty data. This is not a pairwise statistical significance test.

Shared benchmark results

Scores and reported standard errors come from the latest visible Vals benchmark version available for both models.

Benchmark Gemini 3.8 Flash GPT-6 Astra Gap Reported uncertainty
Vals Index 54.83% ±1.03 63.13% ±1.19 8.30 pts Reported ±1 SE ranges do not overlap
Vals RSI Index 19.93% 27.07% 7.14 pts Uncertainty comparison unavailable
Harvey's Legal Agent Benchmark 10.00% ±2.42 5.42% ±1.17 4.58 pts Reported ±1 SE ranges do not overlap
Legal Research Bench 38.94% ±3.39 39.42% ±3.40 0.48 pts Reported ±1 SE ranges overlap
EMB 72.20% ±2.42 71.70% ±2.48 0.50 pts Reported ±1 SE ranges overlap
Finance Agent (v2) 61.44% ±0.13 53.54% ±2.08 7.90 pts Reported ±1 SE ranges do not overlap
Tax Agent Bench 66.77% ±3.15 63.34% ±3.15 3.44 pts Reported ±1 SE ranges overlap
MedCode 48.13% ±2.18 48.49% ±2.13 0.35 pts Reported ±1 SE ranges overlap
MedScribe 84.50% ±1.94 87.91% ±1.94 3.41 pts Reported ±1 SE ranges overlap
ProofBench v1.1 48.00% ±5.02 99.00% ±1.00 51.00 pts Reported ±1 SE ranges do not overlap
BioMysteryBench 62.22% ±1.70 79.26% ±0.98 17.04 pts Reported ±1 SE ranges do not overlap
MysteryMechanism 36.49% ±3.24 53.15% ±3.36 16.67 pts Reported ±1 SE ranges do not overlap
Terminal-Bench Science 8.57% ±3.37 62.86% ±5.82 54.29 pts Reported ±1 SE ranges do not overlap
SAGE 35.06% ±3.36 46.37% ±3.44 11.30 pts Reported ±1 SE ranges do not overlap
Code Migration 36.55% ±4.18 67.74% ±4.22 31.19 pts Reported ±1 SE ranges do not overlap
IOI 56.94% ±2.26 100.00% ±0.00 43.06 pts Reported ±1 SE ranges do not overlap
ProgramBench 1.00% ±0.70 5.50% ±1.62 4.50 pts Reported ±1 SE ranges do not overlap
Terminal-Bench 4.0 19.19% ±2.52 59.60% ±4.40 40.40 pts Reported ±1 SE ranges do not overlap
Vibe Code Bench 1-100 18.77% ±3.31 27.64% ±4.08 8.87 pts Reported ±1 SE ranges do not overlap
Vibe Code Bench v1.1 78.65% ±3.88 89.59% ±2.17 10.94 pts Reported ±1 SE ranges do not overlap
CyberBench v1.1 43.75% ±2.21 41.07% ±2.56 2.68 pts Reported ±1 SE ranges overlap
CUA-bench 4.17% 19.17% 15.00 pts Uncertainty comparison unavailable
Time Horizon Index: KSP 18.83% ±0.00 90.50% ±0.00 71.67 pts Reported ±1 SE ranges do not overlap

Performance by category

Category Gemini 3.8 Flash average GPT-6 Astra average
Legal 24.47% 22.42%
Finance 66.80% 62.86%
Healthcare 66.32% 68.20%
Math 48.00% 99.00%
Science 35.76% 65.09%
Education 35.06% 46.37%
Coding 35.18% 58.35%
Cyber 43.75% 41.07%

Cost and latency

Cost per test appears only where the benchmark reports it for both models. Latency is the measured completion time for that benchmark.

Benchmark Gemini 3.8 Flash cost GPT-6 Astra cost Gemini 3.8 Flash latency GPT-6 Astra latency
Vals Index $5.73 $18.46 51m57s 27m48s
Vals RSI Index $332.04 $759.51 90h00m 90h00m
Harvey's Legal Agent Benchmark $3.66 $26.16 29m21s 24m49s
Legal Research Bench $1.63 $10.58 5m24s 24m51s
EMB $8.24 $5.83 12m49s 13m24s
Finance Agent (v2) $2.00 $6.82 3m21s 11m33s
Tax Agent Bench $0.86 $5.80 2m58s 15m01s
MedCode N/A N/A 42.89s 102.86s
MedScribe N/A N/A 21.14s 2m34s
ProofBench v1.1 $0.60 $1.71 6m09s 5m08s
BioMysteryBench $1.55 $1.72 5m27s 4m46s
MysteryMechanism $1.73 $1.56 5m22s 8m09s
Terminal-Bench Science $5.64 $20.80 55m11s 1h49m
SAGE N/A N/A 28.64s 41.63s
Code Migration $18.49 $44.36 2h37m 58m30s
IOI $3.98 $6.50 14m20s 31m40s
ProgramBench $10.93 $11.57 46m39s 27m02s
Terminal-Bench 4.0 $8.77 $9.58 1h48m 35m42s
Vibe Code Bench 1-100 $18.91 $84.72 1h06m 3h26m
Vibe Code Bench v1.1 $6.87 $38.51 8m39s 39m04s
CyberBench v1.1 $1.31 $1.84 9m49s 4m10s
CUA-bench $280.03 $1821.37 36.55s 32.09s
Time Horizon Index: KSP $543.14 $4206.59 N/A N/A

Results available only for Gemini 3.8 Flash

  • LegalBench
  • MortgageTax
  • TaxEval v2
  • GPQA Diamond
  • MMLU Pro
  • MMMU Pro
  • LiveCodeBench
  • SkillsBench
  • SWE-bench
  • Public Benefits Bench v1.1

Results available only for GPT-6 Astra

  • SRE Bench
Model details Gemini 3.8 Flash Model details GPT-6 Astra