Sep 22, 2026
OpenAI's GPT-6 Luna evaluated across our benchmark suite
We evaluated OpenAIβs new GPT-6 Luna across our benchmark suite.
-
GPT-6 Luna places #20 of 65 on the Vals Index (58.45%) at just $0.42 per test, the lowest cost of any model in the top 20 and roughly 18x cheaper than GPT-6 Sol (62.57%, $7.56).
-
Strongest results: EMB (68.52%, #11 of 65), ProofBench v1.1 (64.00%, #14 of 40), BioMysteryBench (61.48%, #14 of 16), Vibe Code Bench (81.65%, #15 of 103), IOI (55.56%, #16 of 32), Code Migration (42.55%, #21 of 68) and Terminal-Bench 2.1 (73.03%, #22 of 73).
-
Mid-board on SAGE (48.09%, #26 of 85), SNAP (57.65%, #28 of 40), Finance Agent v2 (49.87%, #34 of 68), MedCode (44.69%, #35 of 98), MedScribe (83.71%, #36 of 100), Legal Research Bench (30.29%, #36 of 68) and Tax Agent Bench (58.86%, #19 of 26). On several of these (Finance Agent v2, Legal Research, SNAP, Tax Agent, SAGE, MedScribe) it edges out Sol at a fraction of the cost.
-
Weak spots are the hardest agentic tasks: 2.92% on Harveyβs Legal Agent Benchmark (#35 of 69), MysteryMechanism (19.37%, #13 of 14) and ProgramBench (0.50% fully resolved, #16 of 51).
No fallback models were used; refusals and provider policy blocks are counted as failed tasks. Terminal-Bench 2.1 saw 1 refusal (0.38% of tasks), which did not change its score.
GPT-6 Luna is priced at $0.10 per million input tokens and $0.50 per million output tokens ($0.20 / $0.75 for long-context requests). The model has a 1M-token context window and 128k max output tokens. Evaluations were run with reasoning effort set to βmaxβ.
Congrats to the OpenAI team on the release!