# Evals — Hanzo Cloud

> Benchmark, grade, and regression-test models and agents. LLM-as-judge, golden sets, and drift detection with reproducible scorecards anchored on-chain.

[Cloud](https://hanzo.ai/products)/Observe

Beta

# Evals

Measure model and agent quality.

Open sourceOn-chain settlement

API

/v1/evals

Paste and ship

Paste this into any agent. It reads the skill manifest and calls Evals from there.

Agent promptapi /v1/evalsCopy

```
Read https://hanzo.ai/skill.md and use Hanzo Evals in my project. Start with: hanzo eval runs create
```

Evals — Hanzo

[Overview](https://hanzo.ai/overview)[Evals](https://hanzo.ai/cloud/evals)[Logs](https://hanzo.ai/cloud/logs)[Metrics](https://hanzo.ai/metrics)[O11y](https://hanzo.ai/o11y)[AI Metrics](https://hanzo.ai/cloud/ai-metrics)[Dashboards](https://hanzo.ai/dashboards)[Alerts](https://hanzo.ai/cloud/alerts)

- gatewayReady
- payments-apiReady
- checkout-webBuilding
- docs-siteReady

Every request enters here, so the limiter does too.

`api.hanzo.ai/v1/evals`[Open in the console](https://hanzo.ai/home)

Example project. Every name in it is one you would give your own.

Benchmark, grade, and regression-test models and agents. LLM-as-judge, golden sets, and drift detection with reproducible scorecards anchored on-chain.

LLM-as-judge & golden sets

Regression gates in CI

Drift detection

Anchored scorecards

The AI cloud — one API for every capability. Open models. On-chain.

[Talk to us](https://hanzo.ai/contact/sales)[Source](https://github.com/hanzoai/o11y)

## Start building with Evals

[Get your API key](https://platform.hanzo.ai/api-keys)

One API key. One credit balance. Every primitive.

Related capabilities

[Overview](https://hanzo.ai/overview)[LogsSearch every line](https://hanzo.ai/cloud/logs)[MetricsTime series & SLOs](https://hanzo.ai/metrics)
