One Model to Command Them All
Frontier-level performance without single-vendor lock-in. Enso dynamically orchestrates the world’s best models to tackle complex, multi-step tasks — plug collective intelligence into your workflows through a single API.
Enso is proprietary and available only via Hanzo Cloud. The open-weights Zen family stays free to self-host.
A multi-agent system, delivered as one model
Instead of hand-designing team roles and workflows, Enso learns to assemble agents from a pool and coordinate them through efficient, non-obvious collaboration patterns — automatically, per task.
One API — every model, every modality
Text, code, vision, documents, images, audio, and video through one OpenAI- and Anthropic-compatible endpoint. Enso unifies the frontier and open models across every modality — the first hyper-modal interface. You write one integration, not ten.
Superior on complex, multi-step work
Built for coding, reasoning, research, and other quality-critical workflows. Enso delivers stronger, more reliable results on hard, multi-step tasks than any single model.
You control the agent pool
Opt specific providers or models out of Enso’s pool to meet data, privacy, compliance, or org requirements — with a full audit trail of which models ran, on your organization’s cloud.
Proprietary orchestration, on open foundations
Enso
Proprietary · Hanzo Cloud only
- Learned orchestration over the best models
- Flash · Pro · Ultra presets
- One OpenAI + Anthropic endpoint
- Managed, metered, audited on Hanzo Cloud
Zen
Open weights · run anywhere
- Open-weight frontier models
- Chat, code, and agents
- Self-host on your own hardware
- Free — or managed on Hanzo Cloud
Research-driven coordination for multi-agent intelligence
Enso is grounded in Hanzo’s research on learned model orchestration (HIP-0510): how a system can learn to assemble, route, and coordinate expert agents for each task instead of relying on hand-designed workflows.
Microsecond routing
A lightweight coordinator scores every request and dispatches it to the right model in microseconds — routing overhead you can ignore, applied to every call.
Roles, turns, and verification
Enso assigns Thinker / Worker / Verifier roles and adaptively delegates across coding, math, reasoning, and knowledge tasks — coordinating diverse model pools to outperform any single worker.
Three presets — price × performance
Ultra, Pro, and Flash are distinct cost/quality contracts, monotonic in quality (98.0 > 96.0 > 92.9 GPQA-Diamond). Pick the one that fits your workload, or switch without changing your integration — one endpoint, OpenAI- and Anthropic-compatible.
Enso Ultra
Flagshipenso-ultra
Maximum quality
Top-tier accuracy for hard, high-stakes problems — research reproduction, security analysis, and long-running autonomous work. Reaches 98.0% GPQA-Diamond at a price below premium single models like Opus and fable-5, which score far lower.
- Research & paper reproduction
- Security assessment
- Deep, long-running tasks
Enso Pro
Defaultenso · the default
Balanced — the everyday default
Strong 96.0% GPQA-Diamond with sensible latency — the ideal default for real work: coding, code review, and responsive agents. Priced for everyday scale. Opt out of specific providers to meet data and compliance constraints.
- Coding & code review
- Responsive agents
- Provider opt-out controls
Enso Flash
enso-flash
Fastest, most economical
The high-volume default — lowest latency and cost for everyday chat, classification, extraction, and simple agent steps, at a strong 92.9% GPQA-Diamond.
- High-volume, low latency
- Cheapest per request
- Great default for chat & tools
Frontier capability, measured — without single-vendor risk
Enso reaches frontier-level results by routing each request to the right model in microseconds. Real, measured numbers — not a fabricated benchmark table.
Enso delivers frontier capability without the risk of single-vendor export controls or lock-in — the router always dispatches to a currently-available model in its pool.
The savings are the product
Enso delivers frontier accuracy and bills a fraction of what always calling a top model costs. enso-ultra reaches 98.0% GPQA-Diamond at a price below premium single models that score far lower.
Measured, stated plainly: enso-ultra 98.0%, enso-pro 96.0%, enso-flash 92.9% on GPQA-Diamond — each a distinct price/quality tier, all through one API.
Accuracy at cost
The goal is the top-left: high accuracy, low cost. enso-ultra sits there — 98.0% GPQA-Diamond at a price below premium single models like Opus and fable-5, which score far lower. Pro and Flash trade accuracy for even lower cost.
Solid dots are Hanzo-measured; hollow dots are vendor-reported. enso-ultra (98.0%) leads on accuracy while pricing below premium single models — the win is accuracy-per-dollar across the three tiers.
Frontier accuracy without frontier prices
Top-tier results without paying a top-tier rate on every request. Output price per million tokens across models in the field:
fable-5 costs $42/MTok and scores 81.3% solo on our harness — a premium coordinator that is expensive and worse. The cheap models Enso coordinates run $0.44–$3.73/MTok: up to ~95× cheaper for the same coordination job.
Pay for what each request needs
Simple requests cost little; only the hardest work costs more. You pick the tier, and Enso keeps every request inside that price/quality contract.
Everyday requests cost a fraction of always paying the top rate — extra spend goes only to the hard fraction that needs it. Quality and cost are a property of the tier you choose, not a caller parameter. Per-request costs modeled from published token prices.
What it saves you
Estimate the monthly bill for your volume: always calling a top model, versus Enso routing most requests to a cheap one and escalating only the hard fraction.
Model: a top model bills the premium rate on every request; Enso serves the easy majority cheaply (~$0.002/req) and only the hard fraction costs more (~$0.43/req). Illustrative at published token prices for a typical short-answer request — your mix sets the exact number.
What teams build with Enso
Coding & code review
Enso finds the bugs a single model misses — comprehensive reviews that surface twenty issues where others flag three. Drop it into your existing coding tools unchanged.
Research & autonomy
Point Enso at a paper or a patent landscape and it works autonomously — reading, implementing, training, evaluating, and connecting sources across dozens of documents in hours, not days.
Security assessment
From one scoped instruction, Enso drives an end-to-end assessment — recon, injection and auth checks, and a clean report with evidence and retest steps — staying strictly inside scope.
Orchestration at scale
Frontier-level output with unusually strong persona and identity stability across long sessions — the property that matters most for production agent products.
Pay for intelligence, not integrations
Usage-based, per-organization billing on Hanzo Cloud. When one agent handles a task you pay the standard rate for that model; when Enso coordinates several, you’re charged a single rate based on the top-tier model involved — never stacked fees.
For production workloads that need maximum reliability. Consumption-based tokens, served at higher priority, with transparent per-request cost you can predict and export.
- Single rate — no stacked model fees
- Per-request orchestration trace
- Per-org usage & cost export
For casual, everyday hands-on use. Every tier includes Flash, Pro, and Ultra — upgrade when you need longer, heavier, or more frequent sessions.
- All three presets on every tier
- Standard · Pro · Max usage tiers
- Upgrade or downgrade anytime
Questions, answered
Where can I use Hanzo Enso?
Enso is proprietary and available ONLY through Hanzo Cloud — a single endpoint that speaks both the OpenAI and Anthropic API styles natively. Point your existing OpenAI or Anthropic client at the Hanzo base URL and call an `enso-*` model id — or use it in Claude Code or Codex via the Hanzo CLI. (The open-weights Zen family, by contrast, is free to run on Hanzo Cloud or self-host anywhere.)
What are Flash, Pro, and Ultra?
The three default Enso presets: Flash for fast, high-volume work; Pro as the balanced everyday default for coding and agents; Ultra for maximum quality on hard, high-stakes problems. All behind one API — switch by changing the model id. Zen and other models remain available too.
How is Enso different from the Zen models?
Zen is the family of OPEN-WEIGHT models co-designed by Hanzo AI and the Zoo Labs Foundation (our nonprofit) that you can self-host. Enso is Hanzo’s PROPRIETARY orchestration layer on top — a learned router that assembles and coordinates the best available models (Zen and frontier) per task. Enso runs only on Hanzo Cloud; Zen runs anywhere.
Can I control which models or providers Enso uses?
Yes. Opt specific providers or models out of Enso’s pool to satisfy data-residency, privacy, or compliance requirements. Every request records which models actually ran.
Will my data be used to train models? Can I opt out?
No customer data is used to train models. Enso runs inside your Hanzo Cloud organization with a full audit trail; opt-out and data controls are first-class.
Can I see which underlying models Enso used for each query?
Yes. Each response carries the orchestration trace — the models selected, the roles they played, and the routing decisions — visible in the console and via the API.
Is Enso generally available?
Yes. Enso is available now on Hanzo Cloud and is the default for new chats and API requests — every default request routes through the Enso router, which selects the right tier (Flash, Pro, or Ultra) per task. Zen and other models stay available for explicit selection. Enterprise and dedicated deployment are available on request.
Ready to build with Hanzo Enso?
Enable Enso for your Hanzo Cloud organization, or talk to us about enterprise and dedicated deployment.