Product
Reliability infrastructure for AI agents.
Zeian tests what agents actually do: the tools they call, the arguments they pass, and the outcomes they produce. It is early stage and under active development.
01
Task suites
Realistic tasks declared in YAML: input, allowed MCP tools, expected outcome, trajectory assertions, budgets. Suites live in your repo, next to the agent they verify.
02
Trajectory evaluation
Full traces of every run: tool calls, arguments, retries, cost, latency. Claude-powered scoring grades the path your agent took — deterministic budgets enforced locally.
03
Regression gates
Committed golden baselines per task and a diff on every run. A behavior drift fails the GitHub check and posts the trace excerpt on the PR.
A run, end to end
What Zeian is not
Not another eval scoreboard
Benchmarks grade answers. Zeian grades behavior — the sequence of calls, retries, and decisions that produced the answer.
Not a mock harness
Suites run against real MCP servers. Faults are injected deliberately, not assumed away.
Not observability
Tracing tells you what went wrong after the fact. Zeian runs before merge — regressions never reach production.
Status and roadmap
- Task suite format and local runnerin progress
- Trace collection for MCP tool callsin progress
- Claude-powered trajectory evaluationplanned
- Baselines and regression diffsplanned
- GitHub Action and PR checksplanned
- Hosted dashboardplanned
Want to follow along? The docs track what is real today — and early design partners see what is next.
