Earlier this week, we published an independent evaluation of TypeSafe’s Jev. We ran Inception’s new model, Mercury Decide, through the same tests.
Mercury Decide works like Jev: you input a state and a list of allowed answers, and it returns one answer with a probability for every option, in a single call. Inception prices it at half of Jev’s rate, $0.021 per million input tokens, and output is free. The datasets, prompts, scoring, and statistics for this follow-up are described fully in our Jev blog post.
Key Takeaways
- On quality, Mercury Decide and Jev are indistinguishable on claim verification (0.978 and 0.975). On LegalBench, Mercury Decide is ahead of Jev: 0.869 against 0.740 accuracy, and 0.821 against 0.730 class-balanced.
- On cost, Mercury Decide is a quarter of Jev’s price: $0.006 against $0.025 per 1,000 claims, the lowest of the thirteen systems.
- On speed, Jev is faster. Mercury Decide’s median is 0.16s for one judgment and 0.40s for 64, while Jev stays near 0.1s. It is still faster than every LLM, whose medians run from 3.7s to 65s at 64 judgments.
- With a 1% error budget, Mercury Decide automated 94% of held-out claims at 0.8% error, within the budget. Jev automated 95% at 1.6% error, over it, as did five of the other eleven original systems.
Results
Claim verification
Mercury Decide, like Jev, is frontier-grade at claim verification, and its paired difference from Jev runs from -0.017 to +0.021, so the two cannot be told apart. The difference is price: Mercury Decide costs a quarter as much as Jev. Mercury Decide also costs 1/13th as much as GPT-6 Luna, which is one item behind it, and about 1/2,000th as much as GPT-6 Astra.
By category, Mercury Decide is stronger than Jev on subtle contradictions (0.988 against 0.912) and weaker on direct support (0.963 against 1.000). You can use the menu to the upper right of the pareto chart to explore results across all categories.
LegalBench
Most of Mercury Decide’s lead over Jev comes from the two contract entailment subtasks. On contract_nli_permissible_copy and contract_nli_confidentiality_of_agreement Mercury Decide scores 0.970 on both, against Jev’s 0.485 and 0.576; the best other system scores 1.00 on both.
Mercury Decide’s weakest subtask relative to the field is international_citizenship_questions (0.70, against 0.93 for the best system), where Jev scores 0.82.
Mercury Decide’s calibration is similar to Jev’s on claims (ECE 0.011 for both) and better on LegalBench (ECE 0.058 against 0.108).
How much work can be automated
We repeated the routing test from the Jev post: set a confidence threshold on 30% of the claims to meet a 1% error budget, freeze it, and apply it to the other 70%.
| Router | ||
|---|---|---|
| Claude Sonnet 5.5 | 100% | 0.4% |
| Gemini 4 Argon | 100% | 0.8% |
| Claude Opus 5.5 | 99% | 0.4% |
| GPT-6.1 Sol | 99% | 1.9% |
| GPT-6 Astra | 99% | 1.5% |
| GPT-6 Luna | 97% | 1.9% |
| GPT-5.6 Sol | 97% | 1.6% |
| DeepSeek V4.1 Flash | 96% | 0.8% |
| Jev 1.13.0 | 95% | 1.6% |
| GPT-5.6 Terra | 94% | 1.2% |
| Mercury Decide | 94% | 0.8% |
| GPT-5.4 mini | 84% | 0.9% |
| Claude Haiku 4.5 | 22% | 0.0% |
Mercury Decide automated 94% of held-out cases at 0.8% error, which was within the budget. Seven of the thirteen systems stayed within it, and the same six as before exceeded it. Mercury Decide automates fewer cases than Sonnet 5.5, Opus 5.5, and Gemini 4 Argon, which all take about 99% to 100%.
Latency and cost
We asked Mercury Decide for 1 to 64 judgments about one long document, one request at a time. Its median rises from 0.16s at one judgment to 0.40s at 64. That is faster than every LLM, whose medians run from 3.7s to 65s at 64 judgments, but it grows with the count where Jev’s barely moves. Called at the same moment on the same claims, Mercury Decide had a median latency of 1.44x Jev’s.
Mercury Decide’s billed input tokens also grow with the number of judgments, from about 1,700 at one judgment to about 106,000 at 64. At launch prices that is $0.0022 per call, against $0.00016 for Jev.
Takeaways
Mercury Decide is a highly cost-effective model that matches the frontier systems on bounded claim verification and falls behind on LegalBench, where it is still well ahead of Jev. If your workload looks like claim verification over a document, Mercury Decide costs a quarter of what Jev does for the same accuracy. If you ask many questions about the same document in one call, Jev’s latency and cost scale better.
Methodology
We called Mercury Decide through Inception’s API with one call per item, no retries on malformed responses, and the same instructions and label definitions as every other system. Mercury Decide has no version pin, and this blog post describes it as of October 8, 2026. Costs use each vendor’s listed price; for Mercury Decide, that is the launch price of $0.021 per million input tokens, applied to the input tokens it reported.