Enso, measured
Enso orchestrates 400+ models behind one API. Here is how it scores when we run it — and the field — on a single harness: three differentiated tiers, accuracy-at-cost, and every number kept with its source.
Three tiers, monotonic in quality
Ultra > Pro > Flash — a cost/quality contract, not a model alias. GPQA is Hanzo-measured; price bands are published input→output $/MTok.
Enso Ultra
FlagshipTop-tier accuracy for the hardest, highest-stakes work — reaching 98.0% GPQA-Diamond at a price below premium single models that score far lower.
Enso Pro
DefaultThe production default: strong 96.0% GPQA-Diamond for coding, review, research, and responsive agents — priced for everyday scale. 1M context.
Enso Flash
Fast answers at production scale — for chat, extraction, classification, and simple steps at a strong 92.9% GPQA-Diamond. Lowest cost per request.
Accuracy at cost
enso-ultra reaches 98.0% GPQA-Diamond — top-left is the goal (high accuracy, low cost), and it sits there at a price below premium single models that score far lower. Solid dots are Hanzo-measured; hollow dots are vendor-reported. Every dot is labelled; hover for the exact figure.
Reported vs. what we measured
Pick a benchmark, then filter by provenance. Enso numbers are all Hanzo-measured; the rest of the field shows a mix of what we measured and what vendors report. 134 models, 12 benchmarks.
| Model | GPQA-Diamond | Source | $/MTok out |
|---|---|---|---|
| enso-ultraenso | 98 | Hanzo | $20 |
| ensoenso | 96 | Hanzo | $12 |
| gemini-3.1-pro | 94.3 | Provider-reported | — |
| gpt-5.5 | 93.6 | Provider-reported | $8.25 |
| gpt-5.2-pro | 93.2 | LLM Stats | $139 |
| enso-flashenso | 92.9 | Hanzo | $6 |
| gpt-5.6-sol | 92.9 | Hanzo | $25 |
| gpt-5.2 | 91.7 | Vals AI | $11.55 |
| gpt-5.4 | 91.7 | Vals AI | $12.5 |
| kimi-k2.6 | 89.1 | Vals AI | $2.71 |
| qwen3.5-397b-a17b | 88.4 | LLM Stats | $2.04 |
| gpt-5.6-terra | 87.9 | Hanzo | $12.5 |
| opus-4.8 | 87.4 | Hanzo | $21 |
| nemotron-3-ultra-550b-a55b | 86.1 | Vals AI | $1.54 |
| opus-4.5 | 85.9 | Vals AI | — |
| gpt-5 | 85.6 | Vals AI | $8.25 |
| glm-5.2 | 85.6 | Vals AI | $3.73 |
| sonnet-4.6 | 85.6 | Vals AI | — |
| glm-5.1 | 84.5 | Vals AI | $3.63 |
| gemma-4-31b | 84.3 | LLM Stats | $0.44 |
| o3 | 84.1 | Vals AI | $6.8 |
| kimi-k2.5 | 84.1 | Vals AI | $1.69 |
| glm-5 | 83.3 | Vals AI | $2.07 |
| gpt-5.6-luna | 82.8 | Hanzo | $5 |
| nemotron-3-super | 82.7 | LLM Stats | $0.4 |
| minimax-m2.5 | 82.1 | Vals AI | $0.76 |
| sonnet-4.5 | 81.6 | Vals AI | — |
| mimo-v2.5 | 81.6 | Vals AI | $0.24 |
| fable-5 | 81.3 | Hanzo | $42 |
| deepseek-v3.2 | 80.3 | Vals AI | — |
| gpt-5-mini | 80.3 | Vals AI | $1.65 |
| deepseek-v3.2-exp | 79.9 | DeepSeek-V3.2-Exp model … | — |
| gpt-oss-120b | 78.5 | Vals AI | $0.41 |
| deepseek-4-flash | 76.9 | Hanzo | $0.2 |
| opus-4.1 | 76.3 | Vals AI | $63 |
| o3-mini | 75.5 | Vals AI | $3.74 |
| deepseek-v4-pro | 75.3 | Hanzo | $2.5 |
| o1 | 73.2 | Vals AI | $51 |
| haiku-4.5 | 72.2 | Vals AI | — |
| deepseek-r1 | 71.5 | DeepSeek-R1 model card (… | — |
Vendors report on their own harness; Hanzo measures everyone on one. Where both exist the gap is the harness talking — not the model getting better. Hover a source for its exact provenance.
Build on the tier that fits
Flash, Pro, and Ultra behind one OpenAI-compatible API. Switch by changing the model id.