Hanzo-measured · GPQA-Diamond, one common harness

Enso, measured

Enso orchestrates 400+ models behind one API. Here is how it scores when we run it — and the field — on a single harness: three differentiated tiers, accuracy-at-cost, and every number kept with its source.

Three tiers, monotonic in quality

Ultra > Pro > Flash — a cost/quality contract, not a model alias. GPQA is Hanzo-measured; price bands are published input→output $/MTok.

Enso Ultra

Flagship
enso-ultra
98%GPQA$5 → $20

Top-tier accuracy for the hardest, highest-stakes work — reaching 98.0% GPQA-Diamond at a price below premium single models that score far lower.

Enso Pro

Default
enso
96%GPQA$3 → $12

The production default: strong 96.0% GPQA-Diamond for coding, review, research, and responsive agents — priced for everyday scale. 1M context.

Enso Flash

enso-flash
92.9%GPQA$2 → $6

Fast answers at production scale — for chat, extraction, classification, and simple steps at a strong 92.9% GPQA-Diamond. Lowest cost per request.

Accuracy at cost

enso-ultra reaches 98.0% GPQA-Diamond — top-left is the goal (high accuracy, low cost), and it sits there at a price below premium single models that score far lower. Solid dots are Hanzo-measured; hollow dots are vendor-reported. Every dot is labelled; hover for the exact figure.

768084889296100$0.5$2$8$30$120GPQA %cheap ← output $/MTok → expensivegpt-5.5 — 93.6% GPQA-Diamond · $8.25/MTok (vendor-reported)gpt-5.5 93.6%gpt-5.2-pro — 93.2% GPQA-Diamond · $138.6/MTok (vendor-reported)gpt-5.2-pro 93.2%gpt-5.6-sol — 92.9% GPQA-Diamond · $25/MTokgpt-5.6-sol 92.9%kimi-k2.6 — 89.1% GPQA-Diamond · $2.71/MTok (vendor-reported)kimi-k2.6 89.1%qwen3.5-397b-a17b — 88.4% GPQA-Diamond · $2.04/MTok (vendor-reported)qwen3.5-397b-a17b 88.4%opus-4.8 — 87.4% GPQA-Diamond · $21/MTokopus-4.8 87.4%glm-5.2 — 85.6% GPQA-Diamond · $3.73/MTok (vendor-reported)glm-5.2 85.6%gemma-4-31b — 84.3% GPQA-Diamond · $0.44/MTok (vendor-reported)gemma-4-31b 84.3%fable-5 — 81.3% GPQA-Diamond · $42/MTokfable-5 81.3%enso-ultra — 98% GPQA-Diamond · $20/MTokenso-ultra 98%enso — 96% GPQA-Diamond · $12/MTokenso 96%enso-flash — 92.9% GPQA-Diamond · $6/MTokenso-flash 92.9%

Reported vs. what we measured

Pick a benchmark, then filter by provenance. Enso numbers are all Hanzo-measured; the rest of the field shows a mix of what we measured and what vendors report. 134 models, 12 benchmarks.

ModelGPQA-DiamondSource$/MTok out
enso-ultraenso98Hanzo$20
ensoenso96Hanzo$12
gemini-3.1-pro94.3Provider-reported
gpt-5.593.6Provider-reported$8.25
gpt-5.2-pro93.2LLM Stats$139
enso-flashenso92.9Hanzo$6
gpt-5.6-sol92.9Hanzo$25
gpt-5.291.7Vals AI$11.55
gpt-5.491.7Vals AI$12.5
kimi-k2.689.1Vals AI$2.71
qwen3.5-397b-a17b88.4LLM Stats$2.04
gpt-5.6-terra87.9Hanzo$12.5
opus-4.887.4Hanzo$21
nemotron-3-ultra-550b-a55b86.1Vals AI$1.54
opus-4.585.9Vals AI
gpt-585.6Vals AI$8.25
glm-5.285.6Vals AI$3.73
sonnet-4.685.6Vals AI
glm-5.184.5Vals AI$3.63
gemma-4-31b84.3LLM Stats$0.44
o384.1Vals AI$6.8
kimi-k2.584.1Vals AI$1.69
glm-583.3Vals AI$2.07
gpt-5.6-luna82.8Hanzo$5
nemotron-3-super82.7LLM Stats$0.4
minimax-m2.582.1Vals AI$0.76
sonnet-4.581.6Vals AI
mimo-v2.581.6Vals AI$0.24
fable-581.3Hanzo$42
deepseek-v3.280.3Vals AI
gpt-5-mini80.3Vals AI$1.65
deepseek-v3.2-exp79.9DeepSeek-V3.2-Exp model …
gpt-oss-120b78.5Vals AI$0.41
deepseek-4-flash76.9Hanzo$0.2
opus-4.176.3Vals AI$63
o3-mini75.5Vals AI$3.74
deepseek-v4-pro75.3Hanzo$2.5
o173.2Vals AI$51
haiku-4.572.2Vals AI
deepseek-r171.5DeepSeek-R1 model card (…

Vendors report on their own harness; Hanzo measures everyone on one. Where both exist the gap is the harness talking — not the model getting better. Hover a source for its exact provenance.

Build on the tier that fits

Flash, Pro, and Ultra behind one OpenAI-compatible API. Switch by changing the model id.