Product

Reliability infrastructure for AI agents.

Zeian tests what agents actually do: the tools they call, the arguments they pass, and the outcomes they produce. It is early stage and under active development.

01

Task suites

Realistic tasks declared in YAML: input, allowed MCP tools, expected outcome, trajectory assertions, budgets. Suites live in your repo, next to the agent they verify.

02

Trajectory evaluation

Full traces of every run: tool calls, arguments, retries, cost, latency. Claude-powered scoring grades the path your agent took — deterministic budgets enforced locally.

03

Regression gates

Committed golden baselines per task and a diff on every run. A behavior drift fails the GitHub check and posts the trace excerpt on the PR.

A run, end to end

define suite→
run tasks→
collect trace→
score trajectory→
diff baseline→
verdict + gate

What Zeian is not

Not another eval scoreboard

Benchmarks grade answers. Zeian grades behavior — the sequence of calls, retries, and decisions that produced the answer.

Not a mock harness

Suites run against real MCP servers. Faults are injected deliberately, not assumed away.

Not observability

Tracing tells you what went wrong after the fact. Zeian runs before merge — regressions never reach production.

Status and roadmap

  • Task suite format and local runnerin progress
  • Trace collection for MCP tool callsin progress
  • Claude-powered trajectory evaluationplanned
  • Baselines and regression diffsplanned
  • GitHub Action and PR checksplanned
  • Hosted dashboardplanned

Want to follow along? The docs track what is real today — and early design partners see what is next.