Beta
Evals
Measure model and agent quality.
Open sourceOn-chain settlement
- API
- /v1/evals
Paste and ship
Paste this into any agent. It reads the skill manifest and calls Evals from there.
Every request enters here, so the limiter does too.
api.hanzo.ai/v1/evalsOpen in the consoleBenchmark, grade, and regression-test models and agents. LLM-as-judge, golden sets, and drift detection with reproducible scorecards anchored on-chain.
LLM-as-judge & golden sets
Regression gates in CI
Drift detection
Anchored scorecards
The AI cloud — one API for every capability. Open models. On-chain.
Start building with Evals
One API key. One credit balance. Every primitive.